AI Security

AI Incident Severity Classifier

Classify an AI incident by harm, reach, reversibility, detection and regulatory exposure to produce a consistent severity, the response time it implies, and which notification clocks may already be running.

Last reviewed by the Radiatus Cloud team

Severity appears here.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

AI incidents do not fit a standard severity scale

Conventional severity scales are built around availability and data loss: the system is down, or records were exposed. AI failures are frequently neither. A model that quietly gives systematically worse answers to one group of users, or that has been recommending an unsafe action for three weeks, registers as green on every operational dashboard. The severity has to come from the harm and its reach rather than from the state of the infrastructure, because the infrastructure is usually fine.

Reversibility matters more than magnitude

A wrong answer shown on a screen and corrected is different in kind from a wrong answer that was acted on, and different again from one that has been used to make a decision about someone which they have not been told about. The question that separates a serious incident from a minor one is not how wrong the output was but whether the consequence can still be undone. That is why an incident affecting ten people whose loan applications were declined outranks one affecting ten thousand who saw a badly formatted response.

Detection lag is part of the incident

The time between an AI failure starting and anyone noticing is usually the largest factor in its total impact, and it is the one most often left out of the write-up. A three-week silent degradation is a much larger incident than the same degradation caught in an hour, and it also indicates that the monitoring did not cover the failure mode, which is a second finding worth recording separately from the first.

Related tools

Frequently Asked Questions

Why not use the normal severity scale?

Because conventional scales key on availability and data loss, and AI failures are often neither. A model giving systematically worse answers to one group registers as green on every operational dashboard.

Why is reversibility weighted so heavily?

Because a wrong answer that was shown and corrected differs in kind from one that was acted on. The question separating a serious incident from a minor one is whether the consequence can still be undone.

Does a small number of affected people mean low severity?

No. Ten people whose applications were wrongly declined outranks ten thousand who saw a formatting error, because the harm and its reversibility differ, not the count.

Why record detection lag?

Because it is usually the largest factor in total impact and indicates that monitoring did not cover the failure mode, which is a second finding worth recording separately.

Does this decide whether I must notify a regulator?

No. It flags where notification duties are likely to be engaged so you consult the right people quickly. Whether a duty applies is a legal question and the clocks are short.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Describe the incident to get a severity and response expectation.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.