Skip to main content
L9.6

Keyword and Hybrid Search

Goal

Compare lexical and vector retrieval, explain why each can succeed where the other fails, and combine normalized scores into a simple hybrid ranking.

Suppose you search for the product code XR-4172. Exact characters matter a lot: confusing it with XR-4100 is not "almost right." Now suppose you search for "device dies too fast" while the manual says "short battery duration." Exact words matter less; related meaning matters more.

Those two cases explain why retrieval systems often combine lexical and semantic evidence. Keyword-style search is strong when exact terms, codes, and names carry the signal. Vector search can bridge paraphrases. Hybrid search tries to keep both strengths, but it must combine the two evidence sources carefully rather than pretending their raw score scales mean the same thing.

Lexical search rewards shared terms​

A simple lexical score might count or weight shared words. For:

query: XR-4172 reset
chunk A: XR-4172 factory reset procedure
chunk B: reset instructions for XR-4100

A lexical method can strongly favor A because the exact identifier matches. More advanced lexical systems such as BM25 consider term frequency and document frequency. The durable idea is that literal token overlap carries useful evidence.

Vector search can bridge paraphrases​

Query:

How long can the device run before charging?

Chunk:

Typical battery duration is ten hours.

The exact word overlap may be weak. A semantic embedding can still place the texts near each other. So the two retrievers have complementary failure modes.

Hybrid scores need comparable scales​

Suppose:

lexical score = 12.0
vector cosine = 0.82

Adding them directly gives lexical score much more numerical influence because the scales differ. A simple educational hybrid can normalize each score into a comparable range:

hybrid = α * lexical_normalized
+ (1-α) * vector_normalized

The exact normalization and α are system choices that need evaluation. Do not treat 0.5 as universally optimal.

Candidate fusion is another option​

Instead of merging raw scores, a system can:

  • retrieve top-k lexical;
  • retrieve top-k vector;
  • union candidates;
  • rerank the union.

This avoids pretending two scoring systems have naturally comparable raw values.

Reciprocal rank fusion combines ranks instead of raw scores​

A common alternative is reciprocal rank fusion (RRF). RRF does not add BM25 and cosine values directly. It combines each retriever's rank position:

RRF(document) = Σ 1 / (k + rank_in_each_retriever)

A document that appears near the top of several retrievers accumulates a larger fused score. Because the inputs are ranks, lexical and vector scores do not need to share the same numerical scale.

The constant k controls how strongly the very top ranks are emphasized. It is a system choice, not a universal constant, so evaluate it on held-out queries. RRF also cannot rescue a document that is missing from every candidate list. The next Lesson will separate candidate generation from reranking more explicitly.

Hybrid systems still need diagnostics​

For every result, record enough to answer:

lexical rank/score
vector rank/score
hybrid/fusion rank
final result

If one method consistently dominates, you need to know whether that is intended or a scaling bug.

Exact identifiers are a special case​

Queries often contain values such as:

XR-417
incident-2026-0912
policy_17B

A semantic retriever may understand the surrounding topic while weakening the exact identifier signal. Lexical retrieval is often strong here because exact terms are evidence in their own right. A hybrid system can preserve that lexical signal while still using embeddings for paraphrases.

Hybrid weights must be evaluated, not guessed​

After score normalization, a weighted combination might be:

hybrid = 0.4 * lexical + 0.6 * semantic

The numbers are not universal constants. Different corpora may need different trade-offs, and some queries may be better handled by rank fusion instead of score addition. Evaluate hybrid choices on fixed queries with relevant-document labels. Otherwise a weight can look reasonable while quietly reducing recall on exact-name or paraphrase-heavy slices.

Predict

Which query is most likely to benefit from lexical matching?

Complete the hybrid Lab​

The Lab has three candidate chunks with a lexical (keyword) score and a vector (meaning) score each. alpha sets how much weight the lexical side gets.

  1. Click Run. Compare lexical-only ranking: ['exact_code', 'distractor', 'semantic_paraphrase'] with vector-only ranking: ['semantic_paraphrase', 'exact_code', 'distractor']. The two methods disagree about what is best.
  2. The hybrid line is wrong for now, because minmax returns all zeros and the checks fail. The two score lists use different scales (lexical goes up to 12.0, vector only to 0.92), so they must be put on the same 0-to-1 scale before being combined.
  3. Complete the TODO in minmax: map the smallest score to 0, the largest to 1, and everything else in between. If all scores are equal, return 0.0 for each.
  4. Click Run again. With alpha=0.5, you should see hybrid ranking: ['exact_code', 'semantic_paraphrase', 'distractor'].
  5. Change only alpha=0.5 to alpha=0.2, giving more weight to the vector side. Before running, predict the new winner.
  6. Click Run. semantic_paraphrase now ranks first (0.8 versus 0.587). The underlying scores never changed—only the weighting did.

Loading lab…

Quick Check

1. What is a strength of lexical retrieval?
2. Why not simply add raw BM25-like scores to cosine scores?
3. What is candidate fusion?

0 of 3 questions answered.

Explain it back​

Give one query that lexical search should handle well and one that semantic search should handle well. Then describe one hybrid strategy and what parameter you would evaluate.

Key Takeaways

  • Lexical and semantic retrieval have complementary strengths.
  • Exact identifiers often favor lexical methods.
  • Paraphrases often favor semantic vectors.
  • Raw score scales may not be directly comparable.
  • Hybrid weighting, candidate fusion, and RRF need held-out evaluation.

Next Lesson

Next, retrieve a broad candidate set first and use a second scoring step to rerank those candidates.

References

Lesson actions

Completion is stored locally on this device.

View progress