AI Security

ROUGE Score Calculator

Calculate ROUGE-N recall, precision and F1 between a candidate and reference text, a standard metric for evaluating summaries and generated text.

Last reviewed by the Radiatus Cloud team

Calculate ROUGE-N recall, precision and F1 for generated text.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate ROUGE scores

ROUGE is a family of metrics for evaluating automatic summaries and generated text by comparing them against a reference. This calculator computes ROUGE-N, which measures the overlap of n-grams, contiguous runs of n words, between a candidate and a reference. It reports recall, the fraction of reference n-grams found in the candidate; precision, the fraction of candidate n-grams found in the reference; and their harmonic mean, the F1 score. ROUGE-1 uses single words and ROUGE-2 uses word pairs.

Recall has traditionally been the focus of ROUGE, since a good summary should cover the key content of the reference.

Evaluating generated text

ROUGE is the standard automatic metric in text summarisation and is widely used to track the quality of generated text against human references, complementing human judgement. Because it is based on surface n-gram overlap, it rewards matching wording and can miss paraphrases that convey the same meaning in different words, so it is best used alongside other measures. Higher ROUGE scores generally indicate closer agreement with the reference.

The bigram variant, ROUGE-2, is stricter than ROUGE-1 because it requires matching word order within pairs. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

What is ROUGE-N?

It measures the overlap of n-grams, runs of n consecutive words, between a candidate text and a reference, reporting recall, precision and F1.

What is the difference between recall and precision here?

Recall is the share of reference n-grams the candidate covers, while precision is the share of candidate n-grams that appear in the reference.

Why is ROUGE focused on recall?

A good summary should capture the key content of the reference, so recall, the coverage of reference n-grams, has traditionally been emphasised.

What are the limits of ROUGE?

It rewards surface word overlap and can miss valid paraphrases, so it is best used alongside human judgement and other metrics.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the candidate and reference text, and choose the n-gram size.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.