AI Security

Prompt Injection Simulator

A prompt injection simulator classifies an attack prompt by technique and shows how each defence layer would handle it. It does not fake a model reply; the output is the analysis of which injection family the prompt belongs to and whether your system prompt holds anything worth stealing.

Last reviewed by the Radiatus Cloud team



Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Techniques it recognises

The simulator matches the attack against nine families: direct instruction override, role-play or persona jailbreak, prompt-extraction requests, translation and transformation laundering, fiction framing, encoded payloads, delimiter and role-tag escapes, indirect injection buried in content, and tool or action abuse. Each match comes with why it works and what stops it. Eight sample attacks are built in so you can see how each family is scored.

The defence matrix

For a given prompt the tool shows how five layers respond: an input keyword filter, real role separation with escaped input, an output filter comparing the reply against the system prompt, keeping secrets out of the prompt entirely, and requiring confirmation before tools take actions. The point it makes repeatedly is that input filters catch only verbatim attacks, and the durable defences are architectural, not string matching.

The lesson it enforces

If the system prompt you paste contains a secret (a code, key or password), the tool warns that any successful injection leaks it, and that the real fix is to keep secrets out of the prompt so extraction costs nothing. Bing Chat's instructions were extracted days after launch in 2023 by asking it to ignore previous instructions; the same has happened to many products since. Assume the prompt is public.

What it cannot tell you

It classifies the attack family; it does not run your model, so it cannot say whether your specific deployment falls for a given prompt. Use it to understand the technique and to see which defence layer is relevant, then test the prompt against your real system with a range of paraphrases.

Related tools

  • AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
  • AI Prompt Sanitizer — Scan and sanitize prompts for injection attacks before sending to LLMs. Detect hidden instructions and unsafe patterns.
  • AI Hallucination Detector — Paste AI-generated text to highlight uncertainty markers like 'possibly' and 'reportedly', get a score, and see which sentences to fact-check first.
  • Prompt to JSON Schema — Turn a plain-language description into a JSON Schema for constraining model output, with the keywords that actually matter.

Frequently Asked Questions

Does it send prompts to a language model?

No. It matches the attack prompt against pattern families in your browser and reports the classification and the relevant defences. There is no model call and no fake AI response.

Can it prove my system is safe?

No. It identifies which injection technique a prompt uses and which defences apply. Whether a specific model resists it depends on that model; run the prompt against your deployment, with paraphrases, to know.

What is indirect prompt injection?

An attack where the malicious instructions arrive inside content the model is asked to process, such as a web page, email or document, rather than typed by the user. It is the hardest kind to defend, and the tool flags language typical of it.

Why does keeping secrets out of the prompt matter so much?

Because prompt extraction attacks eventually succeed against most models. If the prompt holds a key or code, extraction is a breach; if it holds nothing sensitive, extraction is harmless. Moving secrets to server-side tools removes the prize.

What defence actually stops instruction override?

No single one fully. Model alignment, a clear instruction hierarchy, and treating user input as data via real chat roles reduce it; output filtering catches leakage. Verbatim input filters are the weakest layer because any paraphrase evades them.

Privacy & Security

Processed locally.

Data: None
Client-side-Side
Active
v1.0

About This Tool

This tool runs entirely in your browser. No data is sent to any server, ensuring complete privacy. Simply use the interface above to get started — no registration or login required.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.