Prompt Canary Token Generator
Generate unique canary tokens to embed in a system prompt, retrieved documents or tool definitions, so that any leak into model output or logs is unambiguously attributable to its source.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
A canary turns an ambiguous leak into an attributable one
When a system prompt appears in the wild, the first question is where it came from: was it extracted from this deployment, reconstructed by someone probing similar behaviour, or copied from a competitor with a similar prompt? Without a marker there is no way to tell. A canary is a string with no meaning and no chance of arising naturally, placed in exactly one location, so that its appearance anywhere else identifies both that a leak happened and which component leaked.
Placement is what makes it useful
One canary in the system prompt tells you the prompt leaked. Separate canaries in the system prompt, in each tool definition and in retrieved documents tell you which of them leaked, which is the difference between knowing there is a problem and knowing where to look. Canaries in retrieved content are particularly valuable because they detect the case where a model has been induced to repeat back material from another tenant's documents, which is otherwise very difficult to notice.
They detect, they do not prevent
A canary is a smoke detector, not a lock. It is worth stating plainly because canaries are sometimes deployed in place of the controls that would actually stop the leak. They also only work if something is watching: a canary with no monitoring on outputs, logs and public search is a string that will be in the leak nobody spotted. And a sufficiently careful extraction can paraphrase around a canary, so absence of a hit is not evidence of no leak.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
What makes a good canary?
High entropy, no natural meaning, and a recognisable fixed prefix so a scanner can find it without knowing every value. It must appear in exactly one place, or it cannot attribute anything.
Where should canaries go?
One per component: the system prompt, each tool definition, and retrieved document sets. Separate values are what let you tell which component leaked rather than just that something did.
Will the model repeat the canary in normal use?
It should not, since it has no meaning and nothing prompts it. If it appears in ordinary output, the model is quoting its instructions more freely than intended, which is itself the finding.
Does a canary prevent prompt extraction?
No. It is a detector, not a control. It also fails against a careful extraction that paraphrases rather than copies, so no hit is not evidence of no leak.
Where should I monitor for hits?
Model outputs, application logs, error reports, support tickets, paste sites and public search. A canary nobody watches for is a string in a leak nobody noticed.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Choose how many canaries you need and where they will be placed.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.