AI Security

Text Cosine Similarity Calculator

Calculate the cosine similarity between two texts using term-frequency vectors, a core measure of semantic closeness in text analysis and search.

Last reviewed by the Radiatus Cloud team

Calculate the cosine similarity between two texts using word-frequency vectors.

Securing AI in production?

We build guardrails, governance & compliance for AI systems.

Talk to an AI advisor

Calculate cosine similarity of texts

Cosine similarity measures the angle between two vectors, and in text analysis those vectors are the word-frequency counts of two documents. This calculator builds a term-frequency vector for each text and computes the cosine of the angle between them, which ranges from zero, meaning no shared words, to one, meaning identical word distributions. It is the classic measure of how similar two pieces of text are by content.

Because it compares the direction of the vectors rather than their length, cosine similarity is unaffected by document length, so a short and a long text about the same topic can still score highly.

From word counts to semantic search

Cosine similarity on term frequencies is the foundation of classic information retrieval and search ranking, and the same cosine measure applied to embedding vectors powers modern semantic search and retrieval-augmented generation. Understanding the term-frequency version builds intuition for how similarity scoring works before embeddings are involved. The dot product and shared-term count shown here reveal how the score is formed.

For meaning beyond exact word matches, embedding-based cosine similarity captures synonyms and context that this term-frequency version cannot. All calculation happens locally in your browser.

Related tools

Frequently Asked Questions

What is cosine similarity?

It is the cosine of the angle between two vectors. For text, the vectors are word-frequency counts, and the result ranges from zero to one.

Why is it unaffected by text length?

It compares the direction of the frequency vectors rather than their magnitude, so proportionally similar texts score highly regardless of length.

How does this relate to embeddings?

The same cosine measure is applied to embedding vectors in semantic search. This version uses word counts, which only match exact words.

What does a score of one mean?

A cosine of one means the two texts have identical word-frequency distributions, pointing in exactly the same direction in the vector space.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Enter two texts to compute their cosine similarity.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.