Semantic Search vs Keyword Search: What's the Difference?
Keyword search matches literal terms; semantic search matches meaning via vector embeddings. How each works, and why most production systems use both.
Keyword search finds documents that contain the literal words in a query, ranked by how often and where those words appear. Semantic search finds documents whose meaning is close to the query’s meaning, even if they share no words at all, by comparing numerical vector representations instead of text. A search for “car won’t start in cold weather” can match a document about “battery failure at low temperatures” under semantic search, but would return nothing under pure keyword matching — neither phrase shares a single word with the other.
How keyword search works
Keyword search — also called lexical search — is built on decades-old information retrieval math. The dominant algorithm, BM25, scores a document by how many query terms it contains, weighted by how rare each term is across the whole corpus (a word like “the” counts for little; a word like “arrhythmia” counts for a lot) and adjusted for document length so long documents don’t win purely by having more words.
This is fast, cheap to compute, exact, and highly interpretable — you can always explain why a document matched. Its weakness is vocabulary mismatch: it has no concept that “car” and “vehicle” or “purchase” and “buy” mean roughly the same thing. If the query and the document don’t share terms, keyword search finds nothing, no matter how relevant the content actually is.
How semantic search works
Semantic search runs the query and every candidate document through an embedding model, which converts text into a dense vector of numbers positioned in a high-dimensional space such that semantically similar text ends up close together. Finding relevant results becomes a geometry problem: embed the query, then search a vector database for the nearest document vectors, typically by cosine similarity.
This catches synonyms, paraphrases, and conceptual matches that share no vocabulary with the query. It also degrades more gracefully with typos and loose phrasing, since the embedding captures overall meaning rather than exact tokens. Its weaknesses are the mirror image of keyword search’s strengths: it can miss an exact required term (a part number, an error code, a proper noun the model hasn’t learned to weight heavily), it’s more expensive to compute and index, and it’s harder to explain why any particular result ranked where it did.
Where each one wins
| Keyword search | Semantic search | |
|---|---|---|
| Best at | Exact terms, codes, names, jargon | Synonyms, paraphrases, conceptual matches |
| Handles typos/loose phrasing | Poorly | Well |
| Compute cost | Low | Higher — requires embedding every document and query |
| Explainability | High — matched terms are visible | Low — similarity score, not a reason |
| Fails when | Query and document share no vocabulary | The exact term matters more than the concept |
| Classic algorithm | BM25 / TF-IDF | Cosine similarity over embeddings |
Neither one is strictly better; they fail in opposite, largely non-overlapping ways. A user searching a product catalog for an exact model number wants keyword precision. A user asking a support tool “why does my app crash on startup” benefits from semantic matching against documents that never use the word “crash.”
Hybrid search: using both
Most production search systems built in the last few years don’t pick one — they run both retrieval methods in parallel and merge the results, an approach usually called hybrid search. A common implementation runs BM25 and vector similarity independently, then combines the two ranked lists with a fusion algorithm like Reciprocal Rank Fusion, which rewards documents that rank well in either list without requiring the two scoring systems to share a scale.
This is also the standard retrieval pattern underneath retrieval-augmented generation: a RAG pipeline that only does vector search will silently miss documents that contain an exact required keyword but happen to embed slightly off-target, so many RAG implementations layer keyword search back in specifically to catch that failure mode. It’s common to follow hybrid retrieval with a reranker that re-scores the merged candidate list with a more expensive model, since neither BM25 nor raw cosine similarity is a particularly precise final relevance signal on its own.
Choosing an approach
- Structured data with exact identifiers — SKUs, error codes, legal citations — leans keyword-first; semantic matching adds little and can actively hurt precision.
- Natural-language questions over unstructured documents — support tickets, internal wikis, research papers — leans semantic-first, since users rarely phrase a query the way the answer is worded.
- Anything user-facing and general-purpose, like a documentation search bar or a RAG-backed assistant, benefits from hybrid retrieval by default, because you can’t predict in advance whether a given query will hinge on an exact term or a paraphrased concept.
The takeaway
Keyword search matches words; semantic search matches meaning. They fail in different, mostly non-overlapping ways — keyword search on vocabulary mismatch, semantic search on exact-term precision — which is why hybrid search, running both and fusing the results, has become the default for anything more demanding than a simple exact-match lookup. When building or evaluating a search feature, the question isn’t which one to use, but whether your queries lean toward exact terms, open-ended meaning, or — like most real traffic — both.
Tagged
Keep reading
Chisato · · 5 min read Constrained Decoding: How LLMs Output Guaranteed JSON
Constrained decoding masks invalid tokens at each step so an LLM can only emit output matching a grammar, regex, or JSON Schema. How it works and its limits.
Chisato · · 5 min read LLM Logprobs Explained: Token Probabilities in Practice
Logprobs are the log probabilities an LLM assigns to each token it generates. What they mean, how to read them, and practical uses like classification.
Chisato · · 4 min read Precision vs Recall, Explained
Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.