Semantic Search
/AI/4 min read
Matching by meaning solves the vocabulary problem and creates a new one: one vector per chunk averages everything together, so a single distinguishing term contributes about 1/√n of the direction. At 200 tokens that is 7%.
Keyword search matches words. Ask about "fixing my car" and a page titled "automobile repair guide" does not appear, because no word in the query appears in the document.
Semantic search fixes this by comparing meaning. Every document is turned into a vector by a model trained so that texts about similar things land near each other; the query is turned into a vector the same way; the nearest ones come back.
That works, and the way it works has a consequence that is rarely stated.
One vector for the whole chunk
A chunk of text becomes a single fixed-length vector. However long the chunk, however many distinct things it mentions, the output is one point.
So the vector is a kind of average. And an average of many things is dominated by the bulk of them.
Quantifying it with vectors in 768 dimensions, averaged over 200 trials — a chunk of n tokens, then querying with exactly one of the terms in it:
| chunk length | match on one exact term | 1/√n | match on a topical query |
|---|---|---|---|
| 20 | 0.2230 | 0.2236 | 0.630 |
| 50 | 0.1415 | 0.1414 | 0.400 |
| 100 | 0.0962 | 0.1000 | 0.283 |
| 200 | 0.0709 | 0.0707 | 0.197 |
| 400 | 0.0512 | 0.0500 | 0.142 |
The measured values track 1/√n almost exactly, which is what the geometry predicts. In a 200-token chunk, a single distinguishing term accounts for about 7% of the chunk's direction.
That is the structural weakness. A document containing one crucial identifier — a part number, an error code, a policy name — has that identifier diluted into insignificance by the surrounding 199 tokens of ordinary prose. A query that is that identifier scores 0.07, while some unrelated chunk that happens to be broadly about the same topic scores 0.20 and wins.
Keyword search does not have this problem at all, and for the opposite reason: it scores a term by how rare it is in the collection, so a term appearing once in one document is the strongest possible signal rather than the weakest.
What this means for chunk size
The table above is the whole chunking debate in one place.
Short chunks keep specifics retrievable — halving a chunk from 200 tokens to 100 raises the single-term signal by √2 — and lose surrounding context, so a retrieved fragment may not carry enough to be useful on its own.
Long chunks preserve context and average specifics away.
There is no setting that is good at both, because the dilution is geometric rather than a tuning artefact. Which is why real systems stop trying: they index at more than one granularity, or they keep a keyword index alongside and merge the results, or they retrieve small and then expand to the surrounding region before sending it to a model.
The parts that are straightforward
The measure. Cosine similarity — the angle between two vectors, from −1 to 1, ignoring length. Both vectors must come from the same model at the same version; a distance between vectors from two different models is a meaningless number, not a weak one.
The storage. Documents are embedded once, ahead of time, and kept in a vector store. Only the query is embedded at request time, which is what makes the search fast enough to be interactive.
The scale. Comparing against every stored vector is exact and too slow past a certain size, so production systems use approximate search: examine a promising neighbourhood rather than everything, accept an occasional miss, and get orders of magnitude more speed for it.
What to take away
Semantic search solves vocabulary mismatch and introduces specificity loss. Both come from the same design: one vector standing in for a whole passage.
So the useful question when it disappoints is which of the two failures you are looking at. If the right document was never retrieved and it contained an exact term from the query, the embedding did not fail to understand — it averaged that term down to 7% and something blander outranked it. That is a job for a keyword index, not a better model.