Keyword and Hybrid Search
Goal
Compare lexical and vector retrieval, explain why each can succeed where the other fails, and combine normalized scores into a simple hybrid ranking.
Suppose you search for the product code XR-4172. Exact characters matter a lot: confusing it with XR-4100 is not "almost right." Now suppose you search for "device dies too fast" while the manual says "short battery duration." Exact words matter less; related meaning matters more.
Those two cases explain why retrieval systems often combine lexical and semantic evidence. Keyword-style search is strong when exact terms, codes, and names carry the signal. Vector search can bridge paraphrases. Hybrid search tries to keep both strengths, but it must combine the two evidence sources carefully rather than pretending their raw score scales mean the same thing.
Lexical search rewards shared terms
A simple lexical score might count or weight shared words. For:
query: XR-4172 reset
chunk A: XR-4172 factory reset procedure
chunk B: reset instructions for XR-4100
A lexical method can strongly favor A because the exact identifier matches. More advanced lexical systems such as BM25 consider term frequency and document frequency. The durable idea is that literal token overlap carries useful evidence.
Vector search can bridge paraphrases
Query:
How long can the device run before charging?
Chunk:
Typical battery duration is ten hours.
The exact word overlap may be weak. A semantic embedding can still place the texts near each other. So the two retrievers have complementary failure modes.
Hybrid scores need comparable scales
Suppose:
lexical score = 12.0
vector cosine = 0.82
Adding them directly gives lexical score much more numerical influence because the scales differ. A simple educational hybrid can normalize each score into a comparable range:
hybrid = α * lexical_normalized
+ (1-α) * vector_normalized
The exact normalization and α are system choices that need evaluation. Do not treat 0.5 as universally optimal.
Candidate fusion is another option
Instead of merging raw scores, a system can:
- retrieve top-k lexical;
- retrieve top-k vector;
- union candidates;
- rerank the union.
This avoids pretending two scoring systems have naturally comparable raw values.
Reciprocal rank fusion combines ranks instead of raw scores
A common alternative is reciprocal rank fusion (RRF). RRF does not add BM25 and cosine values directly. It combines each retriever's rank position:
RRF(document) = Σ 1 / (k + rank_in_each_retriever)
A document that appears near the top of several retrievers accumulates a larger fused score. Because the inputs are ranks, lexical and vector scores do not need to share the same numerical scale.
The constant k controls how strongly the very top ranks are emphasized. It is a system choice, not a universal constant, so evaluate it on held-out queries. RRF also cannot rescue a document that is missing from every candidate list. The next Lesson will separate candidate generation from reranking more explicitly.
Hybrid systems still need diagnostics
For every result, record enough to answer:
lexical rank/score
vector rank/score
hybrid/fusion rank
final result
If one method consistently dominates, you need to know whether that is intended or a scaling bug.
Exact identifiers are a special case
Queries often contain values such as:
XR-417
incident-2026-0912
policy_17B
A semantic retriever may understand the surrounding topic while weakening the exact identifier signal. Lexical retrieval is often strong here because exact terms are evidence in their own right. A hybrid system can preserve that lexical signal while still using embeddings for paraphrases.
Hybrid weights must be evaluated, not guessed
After score normalization, a weighted combination might be:
hybrid = 0.4 * lexical + 0.6 * semantic
The numbers are not universal constants. Different corpora may need different trade-offs, and some queries may be better handled by rank fusion instead of score addition. Evaluate hybrid choices on fixed queries with relevant-document labels. Otherwise a weight can look reasonable while quietly reducing recall on exact-name or paraphrase-heavy slices.
Predict
Complete the hybrid Lab
The Lab has three candidate chunks with a lexical (keyword) score and a vector (meaning) score each. alpha sets how much weight the lexical side gets.
- Click Run. Compare
lexical-only ranking: ['exact_code', 'distractor', 'semantic_paraphrase']withvector-only ranking: ['semantic_paraphrase', 'exact_code', 'distractor']. The two methods disagree about what is best. - The hybrid line is wrong for now, because
minmaxreturns all zeros and the checks fail. The two score lists use different scales (lexical goes up to12.0, vector only to0.92), so they must be put on the same 0-to-1 scale before being combined. - Complete the TODO in
minmax: map the smallest score to0, the largest to1, and everything else in between. If all scores are equal, return0.0for each. - Click Run again. With
alpha=0.5, you should seehybrid ranking: ['exact_code', 'semantic_paraphrase', 'distractor']. - Change only
alpha=0.5toalpha=0.2, giving more weight to the vector side. Before running, predict the new winner. - Click Run.
semantic_paraphrasenow ranks first (0.8versus0.587). The underlying scores never changed—only the weighting did.
Loading lab…
Quick Check
Explain it back
Give one query that lexical search should handle well and one that semantic search should handle well. Then describe one hybrid strategy and what parameter you would evaluate.
Key Takeaways
- Lexical and semantic retrieval have complementary strengths.
- Exact identifiers often favor lexical methods.
- Paraphrases often favor semantic vectors.
- Raw score scales may not be directly comparable.
- Hybrid weighting, candidate fusion, and RRF need held-out evaluation.
Next Lesson
Next, retrieve a broad candidate set first and use a second scoring step to rerank those candidates.
References
Completion is stored locally on this device.