← Back to concepts
9 min read

Cosine similarity

Cosine similarity measures how closely two vectors point in the same direction. It is one of the most common comparison methods in vector search, especially for text embeddings.

The name comes from geometry. Two vectors create an angle, and the cosine of that angle becomes the similarity score. The vectors do not need to have the same length.

Try it yourself: Explore cosine similarity lets you select word vectors, inspect every step of the calculation, and change vector length without changing the score.

Direction instead of straight-line distance

Imagine two arrows that start at the same origin:

  • Arrows pointing in the same direction have cosine similarity near 1.
  • Arrows at a right angle have cosine similarity 0.
  • Arrows pointing in opposite directions have cosine similarity near -1.

This differs from Euclidean distance, which measures the straight-line gap between vector endpoints. Two arrows can point in exactly the same direction but end far apart because one is much longer.

A = [1, 1]
B = [10, 10]

cosine similarity = 1

The two vectors have different magnitudes, but their directions are identical. For a broader comparison with Euclidean, Manhattan, and dot product measures, read vector distance and similarity metrics.

Cosine ignores positive rescaling. Multiplying [1, 1] by 10 keeps the same direction, so the score stays 1. Multiplying by a negative number flips the arrow to the opposite direction, so the sign of the score flips too.

Cosine similarity compares direction instead of endpoint gap Two vectors start at the same origin and point along the same diagonal direction. Vector A is short with endpoint one comma one. Vector B is longer with endpoint three comma three. Their endpoints are separated by Euclidean distance, but the angle between their directions is zero degrees, so cosine similarity is one. Same direction, different length Scaled teaching example: A = [1, 1], B = [3, 3]. x y A endpoint [1, 1] B endpoint [3, 3] Cosine view angle = 0 degrees similarity = 1.00 length does not change the score
Cosine similarity ignores how far the endpoints are from the origin and focuses on the angle between directions.

The cosine similarity formula

Cosine similarity divides the dot product by the product of both vector lengths:

cosine_similarity(A, B) = dot(A, B) / (magnitude(A) * magnitude(B))

Each part has a job:

  • The dot product increases when matching dimensions have the same sign and large values.
  • The magnitudes remove the effect of overall vector length.
  • The final ratio measures directional alignment.

The calculation requires two nonzero vectors. A zero vector has no direction, so cosine similarity is undefined when either magnitude is zero. Software should handle that case explicitly rather than silently returning a meaningful-looking score.

Real software should also clamp tiny floating-point drift. A calculation might produce 1.0000000002 or -1.0000000001 because of rounding, even though valid cosine scores live between -1 and 1.

Worked example

Consider two small teaching vectors:

puppy = [0.90, 0.50, -0.70, 0.90]
dog   = [0.85, 0.45, -0.45, 0.65]

First calculate the dot product by multiplying matching dimensions:

dot
= (0.90 * 0.85)
  + (0.50 * 0.45)
  + (-0.70 * -0.45)
  + (0.90 * 0.65)
= 1.890

Then calculate each magnitude:

magnitude(puppy) = sqrt(0.90^2 + 0.50^2 + -0.70^2 + 0.90^2)
                 = 1.536

magnitude(dog)   = sqrt(0.85^2 + 0.45^2 + -0.45^2 + 0.65^2)
                 = 1.245

Finally divide:

cosine_similarity = 1.890 / (1.536 * 1.245)
                  = 0.988

The value is close to 1, so these vectors point in almost the same direction.

These numbers are hand-authored for teaching. Real embeddings are learned by a model, and their individual dimensions usually do not have human-readable meanings.

The worked example as dot product divided by magnitudes The puppy and dog teaching vectors are compared in three steps. Matching dimensions are multiplied and added to get a dot product of 1.890. The puppy vector has magnitude 1.536 and the dog vector has magnitude 1.245. Dividing 1.890 by 1.536 times 1.245 gives cosine similarity 0.988. Worked example pieces 1. Dot product 0.90 * 0.85 = 0.765 0.50 * 0.45 = 0.225 -0.70 * -0.45 = 0.315 0.90 * 0.65 = 0.585 dot = 1.890 2. Lengths |puppy| = 1.536 |dog| = 1.245 1.536 * 1.245 3. Divide 1.890 / (1.536 * 1.245) 0.988 almost the same direction Numbers are rounded to three decimals, matching the article's teaching example.
The formula removes vector length by dividing the alignment score by both magnitudes.

Positive, zero, and negative scores

The mathematical range is from -1 to 1:

  • 1 means identical direction.
  • 0 means no directional alignment.
  • -1 means exactly opposite direction.
Cosine similarity scores across three angles Three panels show unit vectors from the same origin. In the first, both arrows point right and cosine similarity is 1. In the second, one arrow points right and one points up, making a ninety degree angle with cosine similarity 0. In the third, the arrows point in opposite horizontal directions and cosine similarity is minus 1. Direction determines the score Same direction cos = 1 Right angle cos = 0 Opposite direction cos = -1
The score is about angle: aligned directions score high, perpendicular directions score zero, and opposite directions score negative.

Do not treat these values as universal confidence percentages. A score of 0.8 is not automatically “80% semantically identical.” Useful thresholds depend on the embedding model, the dataset, and the task. A threshold that works for one model may fail for another.

Some text embedding models produce mostly nonnegative similarity scores for ordinary inputs, but that is an observed model behavior, not a rule of cosine similarity. Negative scores are mathematically valid. Evaluate score distributions with real examples before deciding what counts as a match.

Cosine similarity and cosine distance

Search systems often need a distance where smaller means closer. A common conversion is:

cosine_distance = 1 - cosine_similarity

This gives:

  • Similarity 1 -> distance 0
  • Similarity 0 -> distance 1
  • Similarity -1 -> distance 2

Cosine distance is useful for search APIs, but 1 - cosine_similarity is not a true metric in the strict mathematical sense. Libraries and databases do not always use terminology consistently. Some expose cosine similarity, some expose 1 - similarity, and others return a value labeled “score” with their own ordering. Check whether higher or lower is better before interpreting results.

In semantic search, a query and every stored document are represented as vectors. The system compares the query vector with each candidate and returns the highest cosine similarities.

Cosine similarity is useful when direction captures the pattern of meaning and magnitude is not intended to affect ranking. It also has practical advantages:

  • It works well with many text embedding models.
  • It compares vectors with a simple dot product after normalization.
  • It is supported by common vector databases and nearest-neighbor indexes.
  • Its directional interpretation is easy to reason about.

However, the embedding model’s recommendation is more important than a general rule. Some models are trained and evaluated for dot product or Euclidean distance instead. The metric should match the model documentation and the way the vector index is configured.

Normalization makes cosine cheaper

A normalized vector has magnitude 1. Dividing a vector by its magnitude changes its length but not its direction.

When both vectors are normalized:

magnitude(A) = 1
magnitude(B) = 1

cosine_similarity(A, B) = dot(A, B)

This means a system can normalize vectors once and use the dot product for efficient cosine ranking. It also means dot product and cosine similarity produce the same ordering for unit vectors.

For a tiny example:

A = [3, 4]          magnitude = 5
B = [6, 8]          magnitude = 10

normalize A -> [0.6, 0.8]
normalize B -> [0.6, 0.8]

dot([0.6, 0.8], [0.6, 0.8])
= 0.36 + 0.64
= 1.00

Cosine similarity is the dot product after both vectors have been normalized to length 1.

Normalization keeps direction and sets length to one The left side shows a longer vector A with magnitude three and a shorter vector B with magnitude two. Both are scaled onto the same unit circle on the right. Their directions stay the same, their magnitudes become one, and cosine similarity can be computed as a dot product between unit vectors. Normalize once, compare cheaply Original vectors |A| = 3 |B| = 2 divide by length Unit vectors |A'| = 1 |B'| = 1 cosine(A, B) = dot(A', B')
Normalization preserves direction, so cosine ranking can use a dot product once every vector has unit length.

Normalization should match the model and database configuration. Do not normalize automatically if the model intentionally uses vector magnitude as part of its score.

What cosine similarity cannot tell you

Cosine similarity compares numbers. It does not understand language by itself.

If an embedding model puts unrelated inputs in similar directions, cosine similarity will faithfully return a high score. A toy vector can make “wolf” and “storm” look similar because both were assigned natural, dangerous, cold, and serious features. The calculation is correct even if that relationship is not useful for the product.

Other limitations include:

  • A high score does not prove factual equivalence.
  • Scores from different embedding models are not directly comparable.
  • Thresholds that work for one dataset may fail for another.
  • Long or mixed-topic chunks may hide the specific meaning a query needs.
  • Approximate nearest-neighbor indexes may trade a small amount of recall for speed.

Good vector search therefore requires useful embeddings, sensible chunking, metadata filters, and evaluation with real queries. The metric is only one part of the retrieval system.

The key idea

Cosine similarity measures the angle between two nonzero vectors. It rewards matching direction and removes the effect of overall length. That makes it a common choice for semantic vector search, but its results are only as meaningful as the vectors it compares.