AI Safety Evaluation Scorecard
Score an AI safety evaluation programme on harm coverage, adversarial method, measurement quality and process, and identify which claims your current testing actually supports.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
A benchmark score is a claim about the benchmark
Public safety benchmarks cover generic categories, and a model scoring well on them has demonstrated that it handles those categories. The harms that matter in a specific product are usually specific to it: what a medical triage assistant can get catastrophically wrong is not in any general suite, and neither is what a legal research tool can. Evaluations built only from public benchmarks measure comparability against other models, which is a different and less useful thing than fitness for the deployment.
Over-refusal is a failure and one-sided evaluation rewards it
An evaluation that counts only harmful outputs makes a model that refuses everything look perfect. Refusing ordinary requests is a real product failure with real users, and the pressure to reduce one number without measuring the other reliably produces it. Measuring both is not a nicety; without it the evaluation is actively pushing the model in a direction nobody wants, and the direction is invisible in the metric being watched.
Single-turn testing misses how misuse actually works
Real misuse builds context across a conversation: establishing a premise, then a role, then asking for something that follows naturally from both. Most published jailbreaks work this way, and a test suite of individual prompts cannot detect any of them. Multi-turn testing is more expensive and it is where the failures that matter actually live.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Are public benchmarks enough?
No. They measure comparability with other models on generic categories. The harms specific to your product appear in no public suite, which is why a good benchmark score says little about fitness for a deployment.
Why measure over-refusal?
Because an evaluation counting only harmful outputs makes a model that refuses everything look perfect, and pressure to reduce one number without measuring the other reliably produces exactly that.
Why does multi-turn testing matter?
Because real misuse builds context across a conversation, and most published jailbreaks work that way. A suite of individual prompts cannot detect any of them.
How large should the evaluation set be?
Large enough to detect the difference you care about. Fifty examples cannot distinguish 82 percent from 78 percent, so comparisons at that size are noise presented as evidence.
When should the evaluation run?
Before release, and again on every prompt, model or retrieval change. All three change behaviour, and prompt changes are made frequently and rarely reviewed with the care given to code.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Tick what your evaluation programme does to score it.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.