AI Security

RAG Retrieval Quality Evaluator

Paste your queries with their relevant documents and the ranked results your retriever returned, and get precision@k, recall@k, MRR, nDCG and hit rate computed per query and overall.

Last reviewed by the Radiatus Cloud team

Metrics appear here.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Most RAG failures are retrieval failures

When a retrieval-augmented system answers badly, the instinct is to change the prompt or the model. In practice the answer is usually wrong because the right passage was never retrieved, and no prompt recovers from that. Measuring retrieval separately from generation is the single most useful thing you can do to a RAG pipeline, because it tells you which half to work on, and the measurement needs only a labelled set of queries rather than a model.

Recall matters more than precision here

In classical search, precision matters because a user scans the results. In retrieval-augmented generation the model reads whatever it is given, so a missing relevant document is fatal while an extra irrelevant one is usually only noise. That asymmetry means recall@k is the metric to optimise first, and it is why increasing k often improves answers even though it lowers precision. The point at which extra context stops helping is an empirical question, not a principle.

Rank matters, which is why nDCG is worth computing

Precision and recall treat the top result and the tenth as equivalent. They are not: models attend unevenly across a long context and material in the middle is demonstrably less influential. nDCG discounts each hit by the logarithm of its position and normalises against the best possible ordering, so it separates a retriever that finds the right document from one that finds it and puts it first.

Related tools

Frequently Asked Questions

Why measure retrieval separately from generation?

Because a bad answer is usually caused by the right passage never being retrieved, and no prompt change recovers from that. Measuring the two separately tells you which half to work on.

Should I optimise precision or recall?

Recall first. The model reads whatever it is given, so a missing relevant document is fatal while an extra irrelevant one is mostly noise. That is why raising k often improves answers while lowering precision.

What does nDCG add over precision?

Position sensitivity. Precision treats the first and tenth result as equivalent, but models attend unevenly across a long context, so a retriever that ranks the right document first is better than one that merely includes it.

What is MRR?

Mean reciprocal rank: the average of 1 divided by the position of the first relevant result. It answers how far down the list the user or model has to go before finding something useful.

How many labelled queries do I need?

Enough that a change larger than the noise is visible. Thirty to fifty is a workable floor for detecting sizeable differences; a handful will move around too much to guide decisions.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Paste labelled queries and retrieved results to measure retrieval quality.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.