← Back to concepts
9 min read

Vector representations for AI systems

Embeddings are useful because they turn messy inputs — sentences, documents, images, audio, or code — into vector representations. A vector representation is a fixed list of numbers that a system can compare, store, and search.

The important idea is not that the numbers are meaningful by themselves. The meaning comes from their position relative to other vectors produced by the same model.

This article focuses on retrieval-system engineering: how chunks become vectors, how indexes find candidates, how metadata filters narrow the search, and how ranking choices affect what reaches the model.

From content to coordinates

An embedding model maps an input into a point in a high-dimensional space. Instead of two or three coordinates, a production embedding might have hundreds or thousands of dimensions.

You do not inspect each dimension directly. The vector is an internal representation learned from data. What matters is the pattern across all dimensions:

  • Similar inputs should produce nearby vectors.
  • Unrelated inputs should produce distant vectors.
  • The same model should represent all items in a comparable way.

For example, a question about “resetting a password” should land closer to a help article about account recovery than to a pricing page, even if the exact words differ.

Similarity is a distance calculation

Once content is represented as vectors, search becomes a comparison problem. A system embeds the user’s query, compares that query vector to stored document vectors, and returns the closest matches.

Common similarity measures include:

  • Cosine similarity, which compares direction and is common for text embeddings.
  • Dot product, which can be efficient when vectors are unit-normalized or model-specific scoring expects it.
  • Euclidean distance, which compares straight-line distance between points.

The metric must match the embedding model and vector database configuration. A strong embedding model can perform poorly if vectors are indexed with the wrong similarity measure.

Dot product and cosine similarity give the same ordering only when every stored vector and the query vector are normalized to length 1. If lengths carry signal, the two can rank results differently. The broader metric tradeoffs are covered in vector distance metrics.

Retrieval is a pipeline

A production retrieval system rarely does “embed query, return nearest vector” and stop. It usually combines several steps:

  1. Turn documents into chunks with metadata.
  2. Embed each chunk and store it in an index.
  3. At query time, retrieve candidates with vector search, keyword search, or both.
  4. Filter candidates by metadata such as tenant, language, product, or date.
  5. Rerank the best candidates with a stronger model or business rule.
  6. Send the top chunks into the LLM as context.
A retrieval pipeline with vector search, keyword search, filters, and rerankingA user query is sent to both vector search and keyword search. Their candidates are combined, metadata filters remove invalid chunks, a reranker orders the survivors, and the top chunks are passed to the LLM as context.Query-time retrievalUser query"reset password"Vector searchmeaning matchKeyword searchexact termsCandidate poolmerge + dedupeMetadatafiltersReranktop kTop chunksbecome LLM contextHybrid search improves recall; reranking improves precision. Exact design depends on the product.
Retrieval quality comes from the whole pipeline, not just the embedding model.

Why chunks matter

Most AI applications do not embed an entire knowledge base as one vector. They split content into chunks, embed each chunk, and search across those chunks.

Chunking controls what the vector represents:

  • A chunk that is too small may lose context.
  • A chunk that is too large may mix unrelated ideas.
  • A chunk without useful metadata may be hard to explain or cite.
  • Overlap can keep an answer that spans a boundary from being split apart, but too much overlap wastes storage and tokens.
  • Headings, document titles, URLs, and timestamps often belong in metadata or in the chunk text so retrieved evidence is easy to explain.

Good chunks are usually centered on one coherent idea: a section, policy, answer, API endpoint, or troubleshooting step. The embedding should represent something specific enough to retrieve and broad enough to be useful.

How chunk size changes what a vector represents The same help page has three sections: Password reset, Billing, and Data export. A chunk that covers only one line becomes a vector for a single step with too little context. A chunk that covers the whole Password reset section becomes a vector for one complete idea. A chunk that covers all three sections becomes a blurred mix of three topics that matches many queries weakly. One help page, three ways to chunk it Too small Password reset Billing Data export Vector represents: one step, missing context Too little context to answer Coherent Password reset Billing Data export Vector represents: one complete idea Specific and self-contained Too large Password reset Billing Data export Vector represents: a blurred mix of 3 topics Matches many queries weakly
The dashed boundary is the chunk: whatever falls inside it is compressed into a single vector, so the boundary decides what can be retrieved.

Indexes trade speed for recall

A small system can compare the query vector with every stored vector. This is exact search, but it gets slow as the collection grows.

Most production systems use an approximate nearest neighbor index, or ANN index. “Approximate” means the index is designed to find very close candidates quickly, not prove it checked every possible vector. A common idea is an HNSW graph: each vector is connected to nearby vectors, and search walks the graph toward better matches instead of scanning the whole collection.

The tradeoff is speed versus recall. Higher recall means the index is more likely to include the truly relevant chunk in its candidate set. Faster settings may miss some good matches. Retrieval systems usually tune this with real queries, not just benchmark numbers.

Vectors are model-specific

Vectors from different embedding models should not be mixed in the same index unless the system explicitly supports that. Each model learns its own coordinate system, so vectors from different models live in different spaces and cannot be compared directly.

Changing the embedding model usually means re-embedding stored content. Otherwise, new query vectors may be compared against old document vectors that were created in a different space.

Why vectors from different models cannot be compared Two panels use the same stored document vectors from model A: a reset article near the top left and a pricing page near the bottom right. In the left panel, the query is also embedded by model A and lands next to the reset article, which is the correct match. In the right panel, the same query is embedded by model B, whose numbers follow a different coordinate system. Read in model A's space, the query lands next to the pricing page, which is the wrong match. Same model: comparable Documents and query both from model A dim 1 dim 2 Reset article Pricing page Query [0.38, 0.72] Nearest: Reset article (correct) Mixed models: not comparable Documents from model A, query from model B dim 1 dim 2 Reset article Pricing page Query [0.74, 0.36] Nearest: Pricing page (wrong) Illustrative 2D example. Each model learns its own coordinate system, so the same numbers mean different things.
A query vector only has meaning inside the space of the model that produced it, so switching models means re-embedding stored content too.

This is why production systems track:

  • Which embedding model created each vector.
  • The vector dimension count.
  • The chunking strategy used.
  • When the source content and vector were last updated.

Metadata filters, hybrid search, and reranking

Vector representation choices affect quality, cost, and reliability:

  • Dimension size influences storage, memory, and search speed.
  • Domain fit determines whether the model understands your vocabulary.
  • Refresh strategy decides how quickly changed content appears in search.
  • Filtering metadata helps combine semantic similarity with business rules such as tenant, language, date, or product area.
  • Evaluation data shows whether retrieval improves for real user questions, not just demo prompts.

Metadata filters are useful, but they can also remove the answer. A filter like language = en is safe if the metadata is complete. A filter like product = billing can fail if the answer lives in a shared account article labeled differently.

Hybrid search combines keyword search with vector search. Keyword search is good at exact names, error codes, IDs, and rare terms. Vector search is good at paraphrases and related meaning. A reranker can then read the top candidates and reorder them with a more expensive model before the final chunks are sent to the LLM.

Combining metadata filters with similarity ranking A query for resetting a password carries the metadata tenant acme and language en. Step one applies a metadata filter to five stored chunks: a reset steps chunk from another tenant and a German reset FAQ are excluded, while reset steps, login help, and pricing chunks for acme in English are kept. Step two ranks the kept chunks by similarity, with illustrative scores of 0.88 for reset steps, 0.71 for login help, and 0.19 for pricing. Reset steps for acme in English is the top match. QUERY "reset my password" metadata: tenant = acme language = en 1. METADATA FILTER Reset steps | acme | en kept Reset steps | globex | en other tenant Reset FAQ | acme | de wrong language Login help | acme | en kept Pricing | acme | en kept 2. SIMILARITY RANKING Reset steps 0.88 Login help 0.71 Pricing 0.19 Top match: Reset steps (acme, en) Scores are illustrative.
Metadata filters enforce business rules alongside similarity, so a semantically close chunk from the wrong tenant or language is not returned.

Treat embeddings as a search index with model behavior, not as a magic understanding layer. The vectors are powerful, but they only help if the content, chunking, metadata, model, and ranking strategy work together.

Evaluating retrieval

A simple starting metric is recall@k: for each test query, did the correct chunk appear in the top k retrieved results?

Three labeled test queries, k = 3:

Q1 correct chunk appears at rank 1 -> hit
Q2 correct chunk appears at rank 4 -> miss for recall@3
Q3 correct chunk appears at rank 2 -> hit

recall@3 = 2 hits / 3 queries = 0.67

Recall does not measure whether the final answer is well written. It measures whether the retriever gave the generator a fair chance. Pair it with answer-quality checks, citation checks, and latency and cost measurements.

Common failure modes

  • Bad chunk boundaries: the answer is split across chunks, so no single vector represents it well.
  • Stale index: source content changed, but old vectors are still being searched.
  • Mixed embedding versions: queries and documents were embedded with different models or dimensions.
  • Filters remove the answer: metadata rules are too strict or source metadata is incomplete.
  • Keyword-only blind spots: exact search misses paraphrases.
  • Vector-only blind spots: semantic search misses exact IDs, names, or error codes.

The key idea

Vector representations make semantic comparison practical. They let an AI system ask “which items mean something similar to this request?” instead of only “which strings contain the same words?”

That shift powers retrieval-augmented generation, recommendations, duplicate detection, clustering, and many other AI product features.