LLM Guardrail Coverage Checker
Score the controls around a language model application across input handling, output handling, blast radius and monitoring, separating the guardrails that constrain text from the ones that constrain actions.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Guardrails on text do not constrain actions
Most guardrail effort goes into what the model is allowed to say, because that is what is visible and what demos badly when it fails. What the model can do is set somewhere else entirely, by the credentials behind its tools and by what the retrieval layer is willing to return. An application with excellent content filtering and a database credential that reads every tenant's rows has a serious problem that no amount of text filtering addresses, and the text filtering is where the review will have spent its time.
Authorisation belongs in the retrieval query
Filtering retrieved documents by permission after the model has produced an answer is too late: the content entered the context and may already have been summarised, paraphrased or reasoned about in the response. The filter has to be part of the query that fetches the documents, so that unauthorised material is never retrieved. This is a straightforward requirement that is regularly missed because retrieval is built first and multi-tenancy is added later, at which point the natural place to add the filter is the wrong one.
Filters inspect a different string from the one the model reads
Base64, homoglyph substitution, zero-width characters and unusual scripts defeat naive pattern matching while remaining perfectly legible to a model. A guardrail that inspects raw input without normalising it first is checking a string the model will not see. Normalisation before inspection is not a complete defence, but its absence makes the filter trivially avoidable by anyone who tries once.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Are content filters enough?
No. They constrain what the model says, not what it can do. What it can do comes from the credentials behind its tools and from what retrieval returns, and no text filter changes either.
Where should permission filtering happen?
Inside the retrieval query, so unauthorised documents are never fetched. Filtering after generation is too late because the content has already entered the context and may be summarised into the answer.
Why normalise input before inspecting it?
Because base64, homoglyphs and zero-width characters defeat naive matching while remaining legible to the model. A filter on raw text is checking a different string from the one the model reads.
Is rendering model output as HTML risky?
Yes. It gives anyone who can influence the context a cross-site scripting vector, and retrieved content is exactly such an influence. Escape by default.
How often should adversarial testing be repeated?
Before release and after any change to the prompt, the model or the retrieval corpus. A guardrail tested only against the attacks its author imagined has been tested against the weakest possible adversary.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Tick the controls you have in place to score coverage.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.