BLEU Score Calculator
Calculate the BLEU score between a candidate and reference sentence using modified n-gram precision and a brevity penalty, up to BLEU-4.
Last reviewed by the Radiatus Cloud team
Calculate the BLEU score between a candidate and reference sentence.
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Calculate the BLEU score
BLEU, which stands for bilingual evaluation understudy, is the most widely used automatic metric for machine translation. This calculator computes it from a candidate and a reference sentence using modified n-gram precision up to the chosen order, combined as a geometric mean, and multiplied by a brevity penalty that discourages overly short candidates. BLEU-4, using n-grams up to length four, is the standard variant reported in translation research.
Modified precision clips the count of each n-gram to the number of times it appears in the reference, which prevents a candidate from inflating its score by repeating a correct word.
Interpreting BLEU
A BLEU score of one, or one hundred percent, means a perfect match with the reference, which is rare; typical machine translations score well below that. The brevity penalty is important because precision alone would reward very short candidates that only emit words certain to be correct. Because BLEU rewards exact n-gram matches, it can undervalue valid translations that use different wording, so it is best read alongside human evaluation.
If any n-gram order has zero matches, the geometric mean becomes zero, which is why short candidates often score zero on BLEU-4 without smoothing. All calculation happens locally in your browser.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
What does BLEU measure?
It measures how closely a candidate translation matches a reference using modified n-gram precision and a brevity penalty, from zero to one.
What is modified n-gram precision?
It counts matching n-grams but clips each to how often it appears in the reference, preventing score inflation from repeating correct words.
What is the brevity penalty?
A factor below one applied when the candidate is shorter than the reference, discouraging short outputs that would otherwise score high on precision alone.
Why might BLEU-4 be zero?
If there are no matching four-grams, the geometric mean of the precisions becomes zero. Short candidates often hit this without smoothing.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Enter the candidate and reference text to compute BLEU.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.