Articles

Hybrid Search: Combining BM25 and Vector Search

Hybrid search blends keyword-based BM25 ranking with semantic vector search, fixing the blind spots each method has on its own.

Chisato Chisato · · 4 min read
Abstract network of connected nodes representing an AI model

Hybrid search is a retrieval approach that runs a query through two different ranking methods — a keyword-based algorithm like BM25 and a semantic vector search over embeddings — and merges the two result sets into one ranked list. Neither method alone handles every kind of query well, and hybrid search exists specifically to cover the gap between them.

What BM25 actually does

BM25 (Best Matching 25) is a keyword-ranking algorithm descended from the older TF-IDF family. It scores a document against a query based on term frequency (how often the query’s words appear), inverse document frequency (how rare those words are across the whole collection, so common words count for less), and a length-normalization factor that stops long documents from winning purely by containing more words. It’s fast, needs no training, and is exact: if a document contains the literal string a user searched for, BM25 will find it.

That exactness is also its limit. BM25 has no notion of meaning — it can’t match a query for “car” against a document that only says “automobile,” and it can’t tell that “Python” the language and “python” the snake are different concepts unless the surrounding words disambiguate them for it.

What vector search adds

Semantic vector search takes the opposite approach. An embedding model converts both the query and each document into a dense numeric vector such that texts with similar meaning end up close together in that vector space, regardless of exact wording. A search for “affordable laptop” can then retrieve a document about a “budget notebook” even though the two share no keywords. See what vector embeddings are and what a vector database is for the mechanics behind storing and querying these vectors efficiently.

Vector search’s weakness mirrors BM25’s strength: it’s fuzzy by design, which makes it worse at exact matches — model numbers, error codes, proper nouns, SKUs, acronyms — where a user wants the literal string, not something conceptually nearby.

Where each one fails alone

Query typeBM25 aloneVector search alone
Exact model number or error codeStrongOften misses or retrieves near-neighbors
Rare acronym or proper nounStrongWeak unless the term appears in training data context
Paraphrased or conceptual queryWeak — needs literal overlapStrong
Cross-lingual or synonym-heavy queryWeakStrong
Very short, keyword-style queryStrongCan be ambiguous

How the merge actually works

Running both searches is the easy part; combining two differently-scaled ranking signals into one ordered list is where hybrid search earns its complexity. BM25 scores and cosine-similarity scores live on different numeric scales and don’t average meaningfully, so most hybrid implementations use reciprocal rank fusion (RRF) instead of the raw scores. RRF ignores the actual score values and works from each result’s rank position in its own list:

RRF_score(doc) = 1 / (k + rank_bm25(doc)) + 1 / (k + rank_vector(doc))

where k is a small constant (commonly 60) that dampens the effect of any single top-ranked result dominating the merge. A document that ranks well in both lists rises to the top; a document that ranks well in only one still gets credit, just less of it. Because RRF only needs rank order, it works regardless of how differently the two underlying systems compute their scores.

Some systems instead retrieve a candidate set from both methods and pass the combined pool through a dedicated reranker — a model trained specifically to score query-document relevance — which tends to outperform RRF at the cost of extra latency.

Where hybrid search fits in a retrieval pipeline

For retrieval-augmented generation, hybrid search typically sits upstream of chunk selection: candidates come back from both BM25 and vector search over your chunked documents, get merged, and the top results are what actually get passed into the model’s context. See agentic RAG vs traditional RAG for how that retrieval step fits into a larger pipeline, and what a context window is for why narrowing down to the right chunks — rather than throwing everything at the model — matters in the first place.

Pure vector search is usually enough for open-ended, conversational queries where meaning matters more than exact wording — a support chatbot answering “how do I reset my password,” for instance. Pure BM25 still wins for catalog and log search where users type exact identifiers. Hybrid search earns its added complexity when your query traffic is a genuine mix of both — product search, documentation search, and enterprise knowledge bases are the classic cases, since a single search box has to handle everything from “error code E4021” to “why does my order keep failing.”

The takeaway

BM25 finds exact keyword matches and is blind to meaning; vector search finds semantic matches and is weak on exact terms. Hybrid search runs both and merges the results, typically with reciprocal rank fusion or a reranker, to cover the queries either method would miss on its own. It costs more to run than either method alone, so it’s worth adopting when your actual query traffic — not just your test queries — genuinely spans both exact and conceptual search.

Chisato Chisato · · 5 min read

IVF vs HNSW: Vector Index Algorithms Compared

IVF clusters vectors into partitions to narrow a search; HNSW builds a navigable graph. Both trade recall for speed differently at scale.

#AI #Databases #LLMs
Chisato Chisato · · 4 min read

What Is a Knowledge Graph?

A knowledge graph stores facts as entities and labeled relationships instead of rows or documents, letting queries traverse connections directly.

#AI #Databases #LLMs