AI Security

BLEU Score Calculator

Calculate the BLEU score between a candidate and reference sentence using modified n-gram precision and a brevity penalty, up to BLEU-4.

Last reviewed by the Radiatus Cloud team

Calculate the BLEU score between a candidate and reference sentence.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate the BLEU score

BLEU, which stands for bilingual evaluation understudy, is the most widely used automatic metric for machine translation. This calculator computes it from a candidate and a reference sentence using modified n-gram precision up to the chosen order, combined as a geometric mean, and multiplied by a brevity penalty that discourages overly short candidates. BLEU-4, using n-grams up to length four, is the standard variant reported in translation research.

Modified precision clips the count of each n-gram to the number of times it appears in the reference, which prevents a candidate from inflating its score by repeating a correct word.

Interpreting BLEU

A BLEU score of one, or one hundred percent, means a perfect match with the reference, which is rare; typical machine translations score well below that. The brevity penalty is important because precision alone would reward very short candidates that only emit words certain to be correct. Because BLEU rewards exact n-gram matches, it can undervalue valid translations that use different wording, so it is best read alongside human evaluation.

If any n-gram order has zero matches, the geometric mean becomes zero, which is why short candidates often score zero on BLEU-4 without smoothing. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

What does BLEU measure?

It measures how closely a candidate translation matches a reference using modified n-gram precision and a brevity penalty, from zero to one.

What is modified n-gram precision?

It counts matching n-grams but clips each to how often it appears in the reference, preventing score inflation from repeating correct words.

What is the brevity penalty?

A factor below one applied when the candidate is shorter than the reference, discouraging short outputs that would otherwise score high on precision alone.

Why might BLEU-4 be zero?

If there are no matching four-grams, the geometric mean of the precisions becomes zero. Short candidates often hit this without smoothing.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter the candidate and reference text to compute BLEU.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.