AI Security

F1 Score Calculator

Calculate the F1 score from precision and recall, or from true positives, false positives and false negatives, a key classification metric.

Last reviewed by the Radiatus Cloud team

Calculate the F1 score from precision and recall.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate the F1 score

The F1 score is the harmonic mean of precision and recall, giving a single number that balances the two. Precision is the fraction of positive predictions that are correct, and recall is the fraction of actual positives that are found; the F1 score rewards models that do well on both. This calculator computes it from precision and recall values, along with the F0.5 and F2 variants that weight precision or recall more heavily.

Because it uses the harmonic mean, the F1 score is low whenever either precision or recall is low, so it punishes imbalance between the two.

Why F1 matters

In classification tasks, especially with imbalanced data, accuracy can be misleading, so the F1 score is widely used to summarise performance in a way that accounts for both false positives and false negatives. The weighted variants let you favour precision, with F0.5, when false positives are costly, or recall, with F2, when missing a positive is worse. The choice depends on the real-world consequences of each type of error.

An F1 of one means perfect precision and recall, while zero means the model missed everything or was never correct. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

What is the F1 score?

It is the harmonic mean of precision and recall, a single metric that balances the two and is low if either is low.

Why use the harmonic mean?

The harmonic mean is dominated by the smaller value, so the F1 score only stays high when both precision and recall are high.

What are F0.5 and F2?

They are weighted versions: F0.5 emphasises precision, useful when false positives are costly, and F2 emphasises recall, useful when missing positives is worse.

When should I prefer F1 over accuracy?

On imbalanced datasets, where a model can score high accuracy by ignoring the rare class. F1 reflects performance on the positive class.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter precision and recall, or the raw counts.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.