AI Security

LLM Guardrail Coverage Checker

Score the controls around a language model application across input handling, output handling, blast radius and monitoring, separating the guardrails that constrain text from the ones that constrain actions.

Last reviewed by the Radiatus Cloud team

Results appear here.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Guardrails on text do not constrain actions

Most guardrail effort goes into what the model is allowed to say, because that is what is visible and what demos badly when it fails. What the model can do is set somewhere else entirely, by the credentials behind its tools and by what the retrieval layer is willing to return. An application with excellent content filtering and a database credential that reads every tenant's rows has a serious problem that no amount of text filtering addresses, and the text filtering is where the review will have spent its time.

Authorisation belongs in the retrieval query

Filtering retrieved documents by permission after the model has produced an answer is too late: the content entered the context and may already have been summarised, paraphrased or reasoned about in the response. The filter has to be part of the query that fetches the documents, so that unauthorised material is never retrieved. This is a straightforward requirement that is regularly missed because retrieval is built first and multi-tenancy is added later, at which point the natural place to add the filter is the wrong one.

Filters inspect a different string from the one the model reads

Base64, homoglyph substitution, zero-width characters and unusual scripts defeat naive pattern matching while remaining perfectly legible to a model. A guardrail that inspects raw input without normalising it first is checking a string the model will not see. Normalisation before inspection is not a complete defence, but its absence makes the filter trivially avoidable by anyone who tries once.

Related tools

Frequently Asked Questions

Are content filters enough?

No. They constrain what the model says, not what it can do. What it can do comes from the credentials behind its tools and from what retrieval returns, and no text filter changes either.

Where should permission filtering happen?

Inside the retrieval query, so unauthorised documents are never fetched. Filtering after generation is too late because the content has already entered the context and may be summarised into the answer.

Why normalise input before inspecting it?

Because base64, homoglyphs and zero-width characters defeat naive matching while remaining legible to the model. A filter on raw text is checking a different string from the one the model reads.

Is rendering model output as HTML risky?

Yes. It gives anyone who can influence the context a cross-site scripting vector, and retrieved content is exactly such an influence. Escape by default.

How often should adversarial testing be repeated?

Before release and after any change to the prompt, the model or the retrieval corpus. A guardrail tested only against the attacks its author imagined has been tested against the weakest possible adversary.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Tick the controls you have in place to score coverage.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.