Cosine similarity measures how closely two vectors point in the same direction. It is one of the most common comparison methods in vector search, especially for text embeddings.
The name comes from geometry. Two vectors create an angle, and the cosine of that angle becomes the similarity score. The vectors do not need to have the same length.
Try it yourself: Explore cosine similarity lets you select word vectors, inspect every step of the calculation, and change vector length without changing the score.
Direction instead of straight-line distance
Imagine two arrows that start at the same origin:
- Arrows pointing in the same direction have cosine similarity near 1.
- Arrows at a right angle have cosine similarity 0.
- Arrows pointing in opposite directions have cosine similarity near -1.
This differs from Euclidean distance, which measures the straight-line gap between vector endpoints. Two arrows can point in exactly the same direction but end far apart because one is much longer.
A = [1, 1]
B = [10, 10]
cosine similarity = 1
The two vectors have different magnitudes, but their directions are identical. For a broader comparison with Euclidean, Manhattan, and dot product measures, read vector distance and similarity metrics.
Cosine ignores positive rescaling. Multiplying [1, 1] by 10 keeps the same direction, so the score stays 1. Multiplying by a negative number flips the arrow to the opposite direction, so the sign of the score flips too.
The cosine similarity formula
Cosine similarity divides the dot product by the product of both vector lengths:
cosine_similarity(A, B) = dot(A, B) / (magnitude(A) * magnitude(B))
Each part has a job:
- The dot product increases when matching dimensions have the same sign and large values.
- The magnitudes remove the effect of overall vector length.
- The final ratio measures directional alignment.
The calculation requires two nonzero vectors. A zero vector has no direction, so cosine similarity is undefined when either magnitude is zero. Software should handle that case explicitly rather than silently returning a meaningful-looking score.
Real software should also clamp tiny floating-point drift. A calculation might produce 1.0000000002 or -1.0000000001 because of rounding, even though valid cosine scores live between -1 and 1.
Worked example
Consider two small teaching vectors:
puppy = [0.90, 0.50, -0.70, 0.90]
dog = [0.85, 0.45, -0.45, 0.65]
First calculate the dot product by multiplying matching dimensions:
dot
= (0.90 * 0.85)
+ (0.50 * 0.45)
+ (-0.70 * -0.45)
+ (0.90 * 0.65)
= 1.890
Then calculate each magnitude:
magnitude(puppy) = sqrt(0.90^2 + 0.50^2 + -0.70^2 + 0.90^2)
= 1.536
magnitude(dog) = sqrt(0.85^2 + 0.45^2 + -0.45^2 + 0.65^2)
= 1.245
Finally divide:
cosine_similarity = 1.890 / (1.536 * 1.245)
= 0.988
The value is close to 1, so these vectors point in almost the same direction.
These numbers are hand-authored for teaching. Real embeddings are learned by a model, and their individual dimensions usually do not have human-readable meanings.
Positive, zero, and negative scores
The mathematical range is from -1 to 1:
- 1 means identical direction.
- 0 means no directional alignment.
- -1 means exactly opposite direction.
Do not treat these values as universal confidence percentages. A score of 0.8 is not automatically “80% semantically identical.” Useful thresholds depend on the embedding model, the dataset, and the task. A threshold that works for one model may fail for another.
Some text embedding models produce mostly nonnegative similarity scores for ordinary inputs, but that is an observed model behavior, not a rule of cosine similarity. Negative scores are mathematically valid. Evaluate score distributions with real examples before deciding what counts as a match.
Cosine similarity and cosine distance
Search systems often need a distance where smaller means closer. A common conversion is:
cosine_distance = 1 - cosine_similarity
This gives:
- Similarity 1 -> distance 0
- Similarity 0 -> distance 1
- Similarity -1 -> distance 2
Cosine distance is useful for search APIs, but 1 - cosine_similarity is not a true metric in the strict mathematical sense. Libraries and databases do not always use terminology consistently. Some expose cosine similarity, some expose 1 - similarity, and others return a value labeled “score” with their own ordering. Check whether higher or lower is better before interpreting results.
Why cosine similarity is common in vector search
In semantic search, a query and every stored document are represented as vectors. The system compares the query vector with each candidate and returns the highest cosine similarities.
Cosine similarity is useful when direction captures the pattern of meaning and magnitude is not intended to affect ranking. It also has practical advantages:
- It works well with many text embedding models.
- It compares vectors with a simple dot product after normalization.
- It is supported by common vector databases and nearest-neighbor indexes.
- Its directional interpretation is easy to reason about.
However, the embedding model’s recommendation is more important than a general rule. Some models are trained and evaluated for dot product or Euclidean distance instead. The metric should match the model documentation and the way the vector index is configured.
Normalization makes cosine cheaper
A normalized vector has magnitude 1. Dividing a vector by its magnitude changes its length but not its direction.
When both vectors are normalized:
magnitude(A) = 1
magnitude(B) = 1
cosine_similarity(A, B) = dot(A, B)
This means a system can normalize vectors once and use the dot product for efficient cosine ranking. It also means dot product and cosine similarity produce the same ordering for unit vectors.
For a tiny example:
A = [3, 4] magnitude = 5
B = [6, 8] magnitude = 10
normalize A -> [0.6, 0.8]
normalize B -> [0.6, 0.8]
dot([0.6, 0.8], [0.6, 0.8])
= 0.36 + 0.64
= 1.00
Cosine similarity is the dot product after both vectors have been normalized to length 1.
Normalization should match the model and database configuration. Do not normalize automatically if the model intentionally uses vector magnitude as part of its score.
What cosine similarity cannot tell you
Cosine similarity compares numbers. It does not understand language by itself.
If an embedding model puts unrelated inputs in similar directions, cosine similarity will faithfully return a high score. A toy vector can make “wolf” and “storm” look similar because both were assigned natural, dangerous, cold, and serious features. The calculation is correct even if that relationship is not useful for the product.
Other limitations include:
- A high score does not prove factual equivalence.
- Scores from different embedding models are not directly comparable.
- Thresholds that work for one dataset may fail for another.
- Long or mixed-topic chunks may hide the specific meaning a query needs.
- Approximate nearest-neighbor indexes may trade a small amount of recall for speed.
Good vector search therefore requires useful embeddings, sensible chunking, metadata filters, and evaluation with real queries. The metric is only one part of the retrieval system.
The key idea
Cosine similarity measures the angle between two nonzero vectors. It rewards matching direction and removes the effect of overall length. That makes it a common choice for semantic vector search, but its results are only as meaningful as the vectors it compares.