Embedding Similarity Calculator
Paste embedding vectors and compare them with cosine similarity, dot product, Euclidean and Manhattan distance and angular distance, with a full similarity matrix and norm diagnostics.
Last reviewed by the Radiatus Cloud team
Securing AI in production?
We build guardrails, governance & compliance for AI systems.
Cosine and dot product agree only when vectors are normalised
Many embedding models return unit-length vectors, and where they do, cosine similarity and the dot product are the same number and the choice does not matter. Where they do not, the dot product rewards magnitude, so a long vector scores highly against everything and dominates a nearest-neighbour search regardless of direction. Mixing embeddings from two models, or from two versions of one model, is the usual way an index ends up with vectors of different scales and a retrieval system with an inexplicable favourite document.
Cosine similarity has no absolute meaning
A cosine of 0.8 is not a statement about how related two texts are. Its meaning depends entirely on the model, because different models occupy different regions of the space: some produce cosines clustered between 0.7 and 0.95 for everything, so 0.8 there means unrelated. The only sound way to set a threshold is to compute similarities across a labelled sample from your own data and your own model, then choose the value that separates the classes. A threshold borrowed from a blog post about a different model is a guess.
Euclidean distance and cosine rank differently unless normalised
For unit vectors the two produce identical rankings, because squared Euclidean distance is exactly two minus twice the cosine. Off the unit sphere they diverge, and a vector database configured for one metric while the embeddings suit the other returns plausible-looking but wrong neighbours. Checking the norms before choosing a metric takes a moment and settles the question.
Related tools
- AI Prompt Leakage Analyzer — Paste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
- LLM Data Exposure Checker — Check if text contains data likely to be memorized or exposed by LLMs.
- AI Usage Policy Generator — Generate an acceptable use policy for AI tools in your company.
- Model Hallucination Estimator — Estimate risk of hallucinations based on task type and temperature.
Frequently Asked Questions
Is cosine the same as the dot product?
Only when both vectors have unit length. Otherwise the dot product rewards magnitude, so a long vector scores highly against everything and dominates a nearest-neighbour search regardless of direction.
What cosine value means "similar"?
There is no universal answer. Different models occupy different regions of the space, and some cluster everything between 0.7 and 0.95, where 0.8 means unrelated. Calibrate against a labelled sample from your own model.
Does Euclidean distance rank differently from cosine?
Not for unit vectors, where squared Euclidean distance is exactly two minus twice the cosine. Off the unit sphere they diverge, which is how a vector database configured for the wrong metric returns plausible but wrong neighbours.
Why check the norms?
Because they tell you whether your embeddings are normalised, which decides whether the metric choice matters at all. Mixed norms usually mean vectors from more than one model or model version are in the same index.
Can I compare embeddings from different models?
No. Different models produce different spaces, and a similarity between vectors from two models is arithmetic without meaning. Re-embed everything with one model when you change models.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Paste two or more vectors, one per line, to compare them.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.
Related Tools
AI Prompt Leakage Analyzer
AI SecurityPaste a system prompt and a hostile user input to see whether the prompt holds secrets and whether the input carries injection patterns. Local, instant.
LLM Data Exposure Checker
AI SecurityCheck if text contains data likely to be memorized or exposed by LLMs.
AI Usage Policy Generator
AI SecurityGenerate an acceptable use policy for AI tools in your company.