Hybrid Search and Reranking for Production RAG: An ADR
Most retrieval-augmented generation (RAG) systems that disappoint in production fail at the same place: the retriever hands the language model the wrong five chunks. Teams respond by swapping the generator, rewriting prompts, or buying a bigger model, when the cheaper fix sits upstream. A pure vector index is excellent at paraphrase and poor at exact tokens, so a query containing an error code, a part number, or a statute section quietly returns semantically similar but wrong passages.
This record documents a decision many teams reach independently: adopt hybrid search reranking RAG, meaning lexical BM25 plus dense vector retrieval, merged by rank fusion, followed by a cross-encoder reranker over a small candidate set. It is written as an architecture decision record (ADR) so you can copy the structure into your own repository.
You will leave with the reasoning behind each stage, a fusion method you can implement in twenty lines, a latency budget with clearly labelled illustrative numbers, and a rule for when to skip reranking entirely.
What this covers: the decision context, the options considered, a decision matrix, the reference architecture with four diagrams, fusion and reranking mechanics, latency and cost, failure modes, and a rollout checklist.
Context and Background
An ADR starts with forces, not solutions. Here the forces are these. Corpora are heterogeneous: product manuals, tickets, contracts, and code all share one index. Queries are a mixture of natural-language questions and exact-match lookups. The generator can only use a handful of chunks, typically three to ten, before cost, latency, and distraction effects dominate. And the retriever must meet a latency budget, because retrieval time is added to a generation call that already takes seconds.
Dense retrieval, which embeds queries and chunks into vectors and ranks by cosine or inner-product similarity, became the default after embedding models improved. It handles synonyms and paraphrase well. Its weakness is documented: the BEIR benchmark by Thakur and colleagues evaluated retrieval models zero-shot across 18 datasets and found that BM25, the classic lexical scoring function, remained a robust baseline, while re-ranking and late-interaction models achieved the best average performance at a high computational cost (BEIR, arXiv 2104.08663). That single finding motivates the whole ADR: no single retriever dominates, and the strongest rankers are too expensive to run over a whole corpus.
BM25 scores documents by term frequency, inverse document frequency, and length normalization. It needs no training, handles rare tokens precisely, and fails when the query and the document use different words for the same idea. Dense and lexical retrievers therefore fail on different queries, which is exactly the condition under which combining them pays off.
If you are still deciding where the vectors live, our comparison of pgvector against a dedicated vector database covers that adjacent decision, and it interacts with this one: hybrid support differs sharply between engines. For how to measure whether a retrieval change helped, see RAG evaluation metrics with RAGAS and faithfulness.
Anthropic’s published Contextual Retrieval experiments give a concrete, sourced illustration of stacking these stages. On their evaluation, contextual embeddings alone cut the top-20-chunk retrieval failure rate by 35 percent, from 5.7 percent to 3.7 percent. Adding contextual BM25 brought the reduction to 49 percent, a 2.9 percent failure rate. Adding a reranker reached 67 percent, a 1.9 percent failure rate (Anthropic, Introducing Contextual Retrieval). Those figures come from Anthropic’s own datasets and models, so treat them as evidence of direction and rough magnitude, not as a forecast for your corpus.
Decision Record: Hybrid Retrieval with a Reranking Stage
Direct answer. Run BM25 and dense retrieval in parallel, each returning roughly 100 candidates. Merge them with reciprocal rank fusion, which needs no score calibration. Pass the top 30 to 50 fused candidates to a cross-encoder reranker, and send only the top three to ten reranked chunks to the language model. Skip the reranker only when measured precision is already adequate or latency forbids it.
Status, context, and decision
Status: proposed, with the decision to be validated against your own evaluation set before rollout.
Context: a general-purpose RAG service over mixed documents, a p95 end-to-end budget of a few seconds including generation, and queries that include both semantic questions and identifier lookups.
Decision: adopt a three-stage retrieval pipeline. Stage one is candidate generation by two independent retrievers for recall. Stage two is rank fusion to produce a single ordered list. Stage three is cross-encoder reranking for precision. Each stage has a single job, and each can be measured and disabled independently.

Figure 1: Reference pipeline for hybrid search reranking RAG. Two retrievers feed rank fusion; a cross-encoder narrows fifty candidates to five.
The diagram reads left to right. A query is optionally rewritten, then fans out to a lexical index and a vector index. Both return ranked identifiers. The fusion step merges them into one list, the top candidates go to the reranker, and only the reranked top chunks reach the generator. Nothing in this pipeline is novel; the value is in the staging, because each stage trades a different resource.
Why two retrievers instead of one
Recall is the ceiling for everything downstream. If the right chunk is not in the candidate set, no reranker or prompt can recover it. Dense retrieval and BM25 have uncorrelated misses, so the union of their top-100 lists covers more relevant documents than either list alone. This is the same logic as ensemble methods: independent errors cancel.
The practical illustration is the identifier query. A user asks about “error E4012 on the pump controller.” A dense model has likely never seen that token as meaningful and maps it near other error discussions. BM25 treats E4012 as a rare term with high inverse document frequency and ranks the page that contains it first. Conversely, a user asks “why does the pump shut down when pressure drops,” and the manual says “low-pressure cutout trips the motor.” BM25 shares almost no vocabulary with that passage; the dense model bridges it.
This is why the phrase BM25 vs dense retrieval is a slightly misleading framing. The engineering question is not which wins but how to combine them cheaply, which is a fusion problem.
Why fusion before reranking
You could skip fusion and send both lists to the reranker. That doubles duplicates and inflates reranker cost, because the cross-encoder is the most expensive component and scales linearly with candidate count. Fusion deduplicates, orders, and truncates, so the reranker sees about fifty unique passages instead of two hundred overlapping ones. Fusion is also the stage that needs no model inference at all, which makes it effectively free.
Why a cross-encoder at the end
A bi-encoder, the dense retrieval model, embeds the query and each document independently, so document vectors can be precomputed and searched with an approximate nearest neighbor (ANN) index in milliseconds. A cross-encoder reads the query and one document together in a single forward pass and outputs a relevance score. Because attention spans both texts, it captures token-level interaction that a single pooled vector cannot. The cost is that nothing can be precomputed: every query-document pair needs a fresh inference. That is why cross-encoders rank a few dozen candidates, never a corpus.
A third family sits between the two. Late-interaction models such as ColBERT keep one vector per token and score with a maximum-similarity operation, retaining much of the cross-encoder’s fidelity while allowing precomputed document representations. Qdrant exposes this as multivector queries in its Query API, usable as a rescoring stage after a cheaper first pass (Qdrant hybrid queries documentation). The price is storage: one vector per token multiplies index size.
Options Considered and the Decision Matrix
An honest ADR lists the alternatives it rejected and why. Five options were evaluated for the retrieval layer of a mixed-corpus RAG service.
Option A: dense only. One embedding model, one ANN index. Simplest to operate, lowest latency, and often good enough for clean prose corpora with paraphrased queries. It fails on identifiers, rare names, and out-of-domain jargon the embedding model never learned.
Option B: BM25 only. Cheap, explainable, and strong on exact terms. It fails on vocabulary mismatch and conversational queries.
Option C: hybrid with fusion, no reranker. Recall improves over either retriever alone, latency stays at roughly the slower of the two retrievers, and there is no model-serving dependency beyond the embedder. Precision at the top of the list is limited by the quality of the fused ordering.
Option D: hybrid with fusion and a cross-encoder reranker. The highest precision of the options, at the cost of a model-serving component, extra latency, and per-query inference spend.
Option E: dense retrieval plus a reranker, no lexical leg. Recovers some precision but cannot recover recall lost at stage one, so exact-token queries still fail.
| Criterion | A Dense only | B BM25 only | C Hybrid, no rerank | D Hybrid + rerank | E Dense + rerank |
|---|---|---|---|---|---|
| Recall on mixed queries | Medium | Medium | High | High | Medium |
| Exact identifier lookups | Weak | Strong | Strong | Strong | Weak |
| Paraphrase and synonyms | Strong | Weak | Strong | Strong | Strong |
| Precision in top 5 | Medium | Medium | Medium | High | High |
| Added latency (relative) | Lowest | Lowest | Low | Highest | High |
| Operational components | Vector index | Text index | Both, plus fusion | Both, fusion, reranker | Vector index, reranker |
| Explainability | Low | High | Medium | Medium | Low |
The ratings are qualitative judgments from the mechanics described above, not measurements. The matrix shows that D dominates on quality and loses only on cost and complexity, which is why the decision reads as “D by default, C when latency or budget binds, A when the corpus is simple and proven.” The remainder of this record explains how to make that call with data instead of intuition.
Fusion Mechanics: Reciprocal Rank Fusion in Detail
Dense scores and BM25 scores live on incompatible scales. A cosine similarity falls between minus one and one, usually clustered in a narrow band. A BM25 score is unbounded, depends on corpus statistics, and shifts when you add documents. Adding or averaging the raw numbers is therefore unreliable. You can normalize each list (min-max or z-score) and then take a weighted sum, but the normalization itself is fragile when a list has outliers or few results.
Reciprocal rank fusion (RRF) sidesteps the problem by ignoring scores entirely. It was introduced by Cormack, Clarke, and Buettcher at SIGIR 2009, who showed it outperformed Condorcet fusion and individual rank-learning methods on their test collections. The formula for a document d across a set of ranked lists R is:
RRF(d) = sum over lists r of 1 / (k + rank_r(d))
A document that appears at rank 1 in one list and rank 3 in another scores 1/(k+1) + 1/(k+3). A document missing from a list contributes nothing for that list. The constant k, set to 60 in the original paper, damps the advantage of the very top ranks: with k = 60, rank 1 contributes 1/61 and rank 2 contributes 1/62, a gap of under two percent, so a document that two retrievers both place in their top ten beats one that a single retriever places first.

Figure 2: Reciprocal rank fusion. Each retriever contributes one reciprocal term per document, the terms are summed, and the sorted result is the fused list.
The figure shows the whole algorithm: score per list, sum per document, sort. There is no training and no score calibration. This is why RRF is the default fusion in much of the tooling. Qdrant’s Query API supports RRF and also a distribution-based score fusion that normalizes using three-sigma extremes. Its documentation describes RRF as the default, exposes a configurable k, and allows per-prefetch weights, for example weighting the dense leg three times the lexical leg. Note that its documented default for k differs from the 60 in the original paper, so set it explicitly if you want reproducible behavior across engines.
A twenty-line implementation
The following Python fuses any number of ranked id lists. It is deliberately engine-neutral so you can test fusion offline against labelled queries.
from collections import defaultdict
def rrf(ranked_lists, k=60, weights=None, top_n=50):
"""Fuse ranked lists of document ids with reciprocal rank fusion."""
weights = weights or [1.0] * len(ranked_lists)
scores = defaultdict(float)
for lst, w in zip(ranked_lists, weights):
for rank, doc_id in enumerate(lst, start=1):
scores[doc_id] += w / (k + rank)
fused = sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
return [doc_id for doc_id, _ in fused[:top_n]]
bm25_hits = ["d7", "d2", "d9", "d4"]
dense_hits = ["d2", "d5", "d7", "d1"]
print(rrf([bm25_hits, dense_hits], top_n=5))
Run on the toy lists, d2 and d7 rise to the top because both retrievers agree on them, while d9 and d5, each found by only one retriever, trail. That agreement signal is the entire mechanism.
When weighted score fusion beats RRF
RRF discards score magnitude, which is also a weakness. If the dense retriever returns one near-perfect match with a similarity far above the rest, RRF treats it like any other rank-one result. When both retrievers produce well-calibrated, comparable score distributions, a normalized weighted sum can use that confidence. Distribution-based fusion exists for this reason. In practice RRF is the safer default because calibration drifts as the corpus and the embedding model change. Choose score fusion only after an offline comparison on your own queries shows a gain you can reproduce.
A second tuning axis is list depth. Fusing the top 100 from each retriever gives RRF enough overlap to find agreement; fusing only the top 10 makes the result nearly a union. Depth costs little at the index but raises the candidate count for the reranker, so cut the fused list before reranking.
Where Hybrid Lives: pgvector, Qdrant, and Friends
The slug of this post compares pgvector and Qdrant because the hybrid decision is entangled with the engine decision. The two take different approaches, and the right choice depends on whether you want one system or the best system.
PostgreSQL with the pgvector extension keeps vectors beside your relational rows, so a hybrid query can join, filter, and rank in one transaction. Postgres full-text search ranks with ts_rank, which is not BM25: it lacks BM25’s saturation and document-length normalization in the same form. Teams that need true BM25 inside Postgres typically add an extension that implements it, or accept ts_rank as an approximation. The fusion step is then a SQL query that joins two ranked subqueries and sums reciprocal ranks, which is simple and fully transactional. The trade-off is that you own the tuning of two index types inside one database and its resource limits.
Qdrant treats hybrid as a first-class query shape. Its Query API accepts prefetch sub-queries, such as one dense and one sparse, then applies a fusion or rescoring query over the merged candidates, and supports nested prefetches for staged retrieval, for example a quantized first pass followed by full-precision rescoring. Sparse vectors there can carry BM25-style or learned sparse weights such as SPLADE. For a deeper engine comparison, see pgvector vs Qdrant vs LanceDB.
Search engines such as Elasticsearch and OpenSearch also ship RRF-style hybrid retrieval. Check each product’s current documentation for the default constant and the candidate window size, because both change behavior and differ between products.

Figure 3: Query sequence. The application issues both retrievals in parallel, fuses locally, calls the reranker once with a batch, and passes the top passages to the generator.
The sequence matters for latency. The two retrievals run concurrently, so their cost is the maximum of the two, not the sum. Fusion happens in the application or in the database, takes microseconds to a few milliseconds, and the reranker is called once with a batch of passages. The only serial chain is retrieval, then fusion, then reranking, then generation.
Deeper Analysis: Reranker Behavior, Latency, and Cost
What the reranker actually changes
The cross-encoder rescoring step does not add documents; it reorders a fixed set. Its entire contribution is moving relevant passages from positions 10 to 50 up into the top five, and pushing near-miss passages down. That has two measurable consequences. Precision at five improves, and the generator receives less distracting context. Recall at fifty does not change at all, which is the clearest diagnostic you have: if recall at fifty is low, fix retrieval; if recall at fifty is high and precision at five is low, add or tune the reranker.
Reranker candidates come in three deployment shapes. Hosted rerank APIs, offered by several vendors, remove serving work and charge per search or per document batch. Open-weight cross-encoders, such as the BGE reranker family, run on your own GPU or CPU and give you control over data residency. Large language models can also be prompted to rank passages, which is flexible but slower and costlier per candidate. Pricing and model names for hosted rerankers change frequently, so verify current figures on the vendor page instead of relying on a blog post, including this one.
A latency budget, labelled illustrative
The following budget is an illustration of how the stages add up, not a benchmark. Every figure is an assumption you should replace with your own measurements.
| Stage | Illustrative p50 | Illustrative p95 | Notes |
|---|---|---|---|
| Query embedding | 10 to 30 ms | 60 ms | Hosted API adds network time |
| BM25 retrieval, top 100 | 5 to 20 ms | 50 ms | Parallel with dense |
| Dense ANN, top 100 | 5 to 30 ms | 80 ms | Depends on filters and index |
| Fusion and cut | 1 to 3 ms | 5 ms | In process |
| Cross-encoder, 50 passages | 80 to 250 ms | 400 ms | Hardware and passage length dependent |
| LLM generation | 1,000 to 4,000 ms | 8,000 ms | Dominates everything |
Two observations follow from the structure, whatever the real numbers are. First, retrieval proper is small next to generation, so adding a reranker typically increases total latency by a modest fraction, not a multiple. Second, reranker latency scales with candidate count and passage length, so the cheapest optimizations are cutting candidates from 100 to 30 and truncating passages to the portion a relevance judgment needs. Measure the elbow on your evaluation set: precision gains usually flatten well before candidate counts reach one hundred.
Production RAG latency also has a tail. Reranker batches that exceed the model’s maximum sequence length get truncated or split, which can silently change quality. Hosted rerank endpoints add network variance. Set a timeout on the rerank call and define a fallback: if it exceeds, say, twice its p95, return the fused order without reranking. A degraded answer beats a failed one.
Cost arithmetic you can do yourself
Reranking cost is per query: candidates times the per-document rate for a hosted service, or GPU-seconds for self-hosting. Suppose, purely as an illustration, that a hosted reranker charged a flat price per thousand searches and your service handled ten million queries a month. The monthly bill is simply ten thousand times that price. The point of the exercise is the comparison against the alternative: the same spend on a larger generator, or on longer context windows, often buys less answer quality than a precise top five. The reranker also reduces generator cost, because sending five chunks instead of twenty cuts input tokens by roughly three quarters at a fixed chunk size.
That second effect can make the reranker net-neutral or net-positive on cost for high-volume services, which is the opposite of the intuition that it only adds spend. Compute it with your own chunk sizes and generator pricing before assuming either direction.
Enrichment before indexing
Anthropic’s Contextual Retrieval work shows that what you index matters as much as how you merge. Prepending a short, model-generated description of where a chunk sits in its document before embedding it, and before building the BM25 index, helped both retrievers in their experiments. They reported a preprocessing cost of about $1.02 per million document tokens using prompt caching, under their stated assumptions. The mechanism is intuitive: a chunk that says “revenue grew 3 percent over the previous quarter” is ambiguous until the company and period are attached. This is orthogonal to fusion and reranking, and the gains compound, as the 35, 49, and 67 percent progression above shows.
Query-side improvements that interact with hybrid
Hybrid retrieval treats the query as given, but the query is itself a lever. Query rewriting expands abbreviations, resolves pronouns from chat history, and can produce several sub-queries whose results are fused with the same RRF function. If your system retrieves iteratively or lets an agent decide what to fetch next, the retrieval stage described here becomes a tool the agent calls repeatedly; see agentic RAG architecture patterns. When retrieval quality is judged and retried automatically, corrective RAG and self-RAG patterns add a gate after the reranker, which is the natural place to check whether the top chunks actually support an answer.
Metadata filters and where they go
Filters such as tenant, document type, or date range should be applied inside both retrievers, not after fusion. A post-filter applied to a fused top 50 can leave three results, or none, when most candidates belong to other tenants. Pre-filtering in the index costs some ANN efficiency, depending on the engine and filter selectivity, but it preserves candidate depth. Multi-tenant systems should also confirm that the reranker never sees passages the caller is not authorized to read, since the reranker call is a second place where data can leak.
Consequences: Trade-offs, Gotchas, and What Goes Wrong
An ADR is incomplete without consequences, good and bad. The positive consequences are higher recall on mixed queries, better top-of-list precision, and independent, individually measurable stages. The negative ones are real.
More moving parts. You now operate a lexical index, a vector index, a fusion step, and a model-serving dependency. Each has its own failure modes, version drift, and capacity plan. If a single Postgres instance handles your scale, the operational simplicity of keeping both legs there can outweigh a better engine elsewhere.
Index drift between legs. The lexical and vector indexes must reflect the same document set. A document that is embedded but not text-indexed, or the reverse, makes fusion systematically favor one retriever. Build a reconciliation job that compares document ids across both indexes daily.
Tokenization and language mismatches. BM25 depends on the analyzer. Stemming, stop-word lists, and tokenization of identifiers like E4012 or v2.3-beta decide whether exact matches work at all. Multilingual corpora need per-language analyzers, and a multilingual embedding model does not fix a monolingual lexical index.
Chunking decides the ceiling. Hybrid retrieval and reranking operate on chunks. If chunks split a table from its header or a clause from its definition, no ranker can reassemble them. Evaluate chunking before blaming retrieval.
Reranker input limits. Cross-encoders have a maximum sequence length. A long chunk plus a query may exceed it and be truncated without error. Check the documented limit for your model and size chunks accordingly.
Position bias downstream. Even with a perfect top five, generators can underuse passages in the middle of long contexts. Putting the highest-scored chunk first and keeping the context short are cheap mitigations; the research on this effect is well known, but its magnitude varies by model, so test yours.

Figure 4: Decision flow. Add the lexical leg when exact terms matter, and add the reranker only when top-of-list precision remains the bottleneck.
When to skip reranking
Reranking is not free, and the flow above encodes the cases where skipping is correct. Skip it when your corpus is homogeneous clean prose and dense recall at twenty is already high. Skip it when your end-to-end latency budget is tight, such as voice assistants or autocomplete-style experiences where a few hundred milliseconds is the whole budget. Skip it when your evaluation shows the fused top five already contains the answer in nearly every query. And skip it initially if you cannot yet measure retrieval quality, because you cannot tell a reranker that helps from one that merely reorders. Build the evaluation set first.
The reverse also holds. Add a reranker when long-tail queries show answers sitting at positions 8 to 30, when users complain about plausible-but-wrong citations, or when the generator is given large contexts to compensate for weak ranking. In those cases the reranker usually pays for itself by letting you shrink the context.
Evaluating the decision
Measure each stage separately, because aggregate answer quality hides which stage helped. For retrieval, track recall at 20 or 50, mean reciprocal rank, and normalized discounted cumulative gain at ten on a labelled query set drawn from real traffic. For generation, track faithfulness and answer relevance using a framework such as those covered in our RAG evaluation article. An A/B comparison of options C and D, holding the generator constant, directly measures whether reranking earns its cost. Public benchmarks such as BEIR and MTEB are good for shortlisting models, but your corpus is the benchmark that matters: BEIR itself was built to show that models ranked on one domain can rank differently on another.
Practical Recommendations
Start with the simplest pipeline that can be measured, then add stages when data justifies them. In order:
First, build a labelled evaluation set of at least a few hundred real queries with known relevant chunks, including exact-identifier queries and paraphrased questions. Without this, every later choice is opinion.
Second, run dense-only as the baseline and record recall at 20 and precision at 5. Third, add BM25 and RRF with k set to 60 and depth 100 per retriever, and compare. In most mixed corpora this step gives the largest single improvement per unit of effort, and it needs no model serving.
Fourth, add the cross-encoder over the top 30 to 50 fused candidates, and compare precision at 5, end-to-end latency, and generator token spend. Keep it only if the gain survives a held-out query set. Fifth, instrument each stage with latency histograms and a timeout-with-fallback on the reranker.
Choose the engine by operational fit. If your data and permissions already live in Postgres and volume is moderate, pgvector plus full-text search plus SQL fusion keeps the stack small. If hybrid, sparse vectors, and multi-stage rescoring are central and you want them as native query primitives, a dedicated engine such as Qdrant is the cleaner fit.
Rollout checklist
- Labelled evaluation set with identifier and paraphrase queries
- Dense-only baseline recorded
- BM25 leg added with analyzers verified for your languages
- RRF fusion with explicit
kand per-leg depth - Filters applied inside both retrievers
- Reranker on 30 to 50 candidates with timeout and fallback
- Reconciliation job comparing both indexes
- Per-stage latency and quality dashboards
- Re-evaluation scheduled whenever the embedding model or chunking changes
Frequently Asked Questions
What is hybrid search in RAG?
Hybrid search runs a lexical retriever such as BM25 and a dense vector retriever on the same query, then merges their ranked results into one list. Lexical search is strong on exact terms, identifiers, and rare words. Dense search is strong on paraphrase and meaning. Because their mistakes differ, the merged list typically has higher recall than either alone. The merge is usually done with reciprocal rank fusion, and the result is often reranked before it reaches the language model.
What is reciprocal rank fusion and why use k equal to 60?
Reciprocal rank fusion scores each document by summing 1 divided by k plus its rank across every list it appears in. It ignores raw scores, so BM25 and cosine similarity need no calibration. The constant k of 60 comes from the original 2009 paper by Cormack and colleagues and damps the dominance of rank one. Some engines default to other values, so set k explicitly and tune it only against your own evaluation queries.
Is a cross-encoder reranker worth the added latency?
Often yes, when precision at the top of the list is the bottleneck, but not always. A cross-encoder reads the query and passage together, which captures interactions that a single embedding cannot, and it lets you send fewer chunks to the generator. It adds one model call per query. Measure precision gains and total latency on your own data, and skip it when dense or hybrid recall and precision are already high.
How many candidates should I pass to the reranker?
Thirty to fifty fused candidates is a common starting range, but treat it as a hypothesis. Reranker latency grows roughly linearly with candidates and passage length, while quality gains flatten as you add lower-ranked items. Sweep candidate counts, such as 20, 30, 50, and 100, on a labelled query set and choose the knee of the curve. Also confirm that candidates plus the query fit within the reranker’s maximum input length.
Can I do hybrid search entirely in PostgreSQL?
Yes, for moderate scale. pgvector provides vector search, Postgres full-text search provides lexical ranking, and a SQL query can fuse the two ranked subqueries with reciprocal rank fusion. Note that the built-in ts_rank is not BM25, so teams wanting true BM25 typically add an extension. The benefit is transactional consistency and one system to run. The limit is that very large or latency-critical workloads may fit a dedicated engine better.
Does reranking replace the need for good chunking and embeddings?
No. A reranker only reorders candidates the retrievers found, so it cannot recover a relevant passage that was never retrieved or one that chunking split apart. Recall comes from chunking, embeddings, the lexical leg, and enrichment such as contextual prefixes. Precision comes from fusion and reranking. Diagnose which one is failing by comparing recall at fifty with precision at five before spending effort on either.
Further Reading
- Agentic RAG architecture patterns for iterative retrieval where this pipeline becomes an agent tool
- pgvector vs dedicated vector database in 2026 for the engine decision behind hybrid support
- RAG evaluation metrics: RAGAS and faithfulness for measuring retrieval and generation separately
- Corrective RAG and self-RAG patterns for gating retrieved context after reranking
- BEIR: a heterogeneous benchmark for zero-shot retrieval evaluation
- Anthropic: Introducing Contextual Retrieval
- Qdrant hybrid queries documentation
By Riju — about
