How to Rerank RAG Results Without Hiding Retrieval Failures
Rerank RAG results in two stages: retrieve a broad candidate set with fast lexical, dense, or hybrid search, then apply a more precise model to compare each candidate with the actual question. After scoring, build an evidence set rather than taking the first few passages blindly: remove duplicates, cover distinct claims, preserve source authority and permissions, and keep a trace from every selected passage back to its original location.
By Sahil Maheshwari
The short answer: reranking is not retrieval repair
A first-stage retriever is designed to search a large collection cheaply. It may use BM25, embeddings, or both and return dozens of plausible candidates. A reranker can spend more computation on that smaller pool because it no longer has to compare the question with every passage in the index. Its job is to put the most useful evidence near the top.
The classic Passage Re-ranking with BERT paper demonstrated the cross-encoder pattern: evaluate a query and candidate passage together, then assign a relevance score. Joint attention can notice relationships that a single embedding and nearest-neighbour comparison may miss. The price is that every query–passage pair needs a model call, so this belongs after fast retrieval, not across the full corpus.
Reranking has a hard boundary. If the candidate pool does not contain the necessary evidence, no scoring model can recover it. A reranker also cannot fix a document that was parsed incorrectly, a permission filter that removed the right source, or an index that has not received the latest version. Treat an empty or weak candidate set as a retrieval diagnosis, not an invitation to promote the least bad passage.
A reliable pipeline therefore records four separate outputs: candidates found, reranking scores, evidence selected, and context finally sent to the generator. Keeping those stages visible makes it possible to tell whether a bad answer began with recall, ranking, set construction, or generation.
Reranking can improve the order of retrieved evidence. It cannot rank evidence that never entered the candidate set.
Define what ‘relevant’ means for the answer
Semantic similarity is only one part of useful evidence. A passage can share the question's vocabulary while describing the wrong product, jurisdiction, time period, or entity. Another can answer the right question but come from a superseded draft. Before choosing a reranker, define the conditions that make evidence eligible and useful for the workflow.
Apply hard constraints before scoring: tenant and user permissions, document status, required language, applicable date, or explicitly selected workspace. Then give the reranker the metadata that changes interpretation, such as document title, section heading, source type, publication date, and nearby structural context. Do not expect it to infer scope that ingestion removed.
Separate query relevance from authority and freshness. A model score can say that a paragraph addresses the question; it does not prove that the paragraph is current, approved, or legally controlling. Keep authority and version rules explicit, and preserve the original component scores. A single blended number is convenient for sorting but difficult to audit when an answer cites the wrong source.
Use examples from real questions to define labels. For ‘What changed in the renewal terms?’, a passage containing the latest terms is relevant, but so is the earlier version needed for comparison. For ‘What did the committee decide?’, background discussion is not interchangeable with the recorded decision. The useful evidence set depends on the work the answer must perform.
Retrieve for recall before optimising the order
Start with enough candidates to give the reranker a real choice. The right pool size depends on corpus size, chunk granularity, query type, latency budget, and the first-stage retriever. Too few candidates cap recall. Too many add cost and may give a stronger model more opportunities to be distracted by plausible but irrelevant text.
Measure candidate recall directly: for each evaluated question, did the pool contain every passage needed to support the answer? If not, tune lexical and semantic retrieval, query expansion, metadata filters, chunking, or index freshness before tuning the reranker. Keep the candidate cutoff and reranker cutoff as separate parameters so one cannot conceal the other.
The BEIR benchmark compared lexical, sparse, dense, late-interaction, and reranking systems across varied retrieval tasks. Its results are a useful warning against declaring one method universally best: generalisation and computational cost differ across domains. Evaluate on your own question and document distribution, including unfamiliar topics rather than only examples used during tuning.
Preserve the first-stage scores and source of each candidate. A passage retrieved by an exact identifier match carries a different signal from one found only by semantic similarity. These features can help ranking, but they are also diagnostic evidence when a query unexpectedly fails.
Choose a reranker by cost, control, and evidence needs
A cross-encoder scores the question and each passage together. It is a strong default when the candidate pool is modest and precise semantic comparison matters. Batch the pairs, cap passage length deliberately, and include headings or document context without letting boilerplate dominate the input. Long passages may need structure-aware windows so the relevant sentence is not truncated.
Late-interaction models offer another point on the efficiency curve. ColBERT encodes queries and documents separately but retains token-level representations for a finer-grained interaction at scoring time. This enables precomputed document representations while preserving more detail than one vector per passage. It can serve as a reranker or as the retrieval architecture itself, but it adds index and operational complexity.
An instruction-tuned language model can also rank contexts. RankRAG trains one model for both context ranking and answer generation and reports strong results on its evaluated benchmarks. That is evidence that ranking and generation can share a model, not proof that every general-purpose chat model is a calibrated reranker. Prompted listwise ranking can be sensitive to order, output format, and candidate count.
Choose the smallest model that passes the evaluation. Cache document-side representations where the architecture allows it, batch scoring, and set a latency budget per stage. A reranker that is marginally more accurate but too slow for interactive use may reduce system quality in practice by encouraging smaller candidate pools or timeouts.
- Cross-encoder: precise pairwise scoring, simple mental model, higher per-candidate cost.
- Late interaction: reusable document representations, token-level matching, more index complexity.
- LLM ranking: flexible instructions and joint workflows, but higher cost and more output variability.
Select an evidence set, not merely the top passages
The highest-scoring passages often repeat one another. A report summary, introduction, and conclusion may express the same point, while the lower-ranked appendix contains the qualifying detail. Sending three near-duplicates consumes context without increasing support. After relevance scoring, select passages as a set with a defined evidence budget.
Deduplicate exact and near-identical chunks, then diversify where the question needs coverage. The original Maximal Marginal Relevance paper balances query relevance against similarity to items already selected. The principle remains useful: each new passage should add information, not simply confirm that the first passage used common wording.
Diversity is not a goal by itself. A query asking for one exact clause may need one authoritative passage plus its surrounding section. A comparison needs evidence for both sides. A multi-source question may require one passage for each hop. Encode these coverage requirements explicitly where possible, and cap how many passages one document or duplicated source can contribute.
Expand selected chunks only after ranking. Attach the heading, preceding definition, table continuation, figure caption, or neighbouring message needed to interpret the evidence. Keep this structural expansion separate from relevance scoring so the system can cite the matched span while still giving the generator enough context.
The best context is a small set that covers the answer's distinct claims—not a leaderboard of passages that all say the same thing.
Use thresholds, abstention, and context order deliberately
Top-k always returns k items, even when every item is weak. Add an acceptance threshold or sufficiency check after reranking. Calibrate it on held-out questions, including unanswerable queries and near-miss passages. If no candidate clears the bar, broaden retrieval once or return what evidence is missing rather than manufacturing an answer from rank one.
Do not treat raw scores from different queries as automatically comparable. Cross-encoder logits, late-interaction scores, and prompted ratings have different scales, and those scales can shift by domain. If a workflow needs a global threshold, test calibration by query type and corpus segment. Otherwise use relative margins, explicit evidence checks, or a small classifier trained for the decision.
Ranking does not end when passages are selected. The Lost in the Middle study found that model performance can change with the position of relevant information in long contexts, with evidence in the middle often used less reliably in its tested settings. Keep the context compact, group related passages, label sources clearly, and test several orders rather than assuming the generator will treat every token equally.
Citations must follow the source objects through every reorder. Never cite by array position after deduplication or context expansion. Use stable document, page, section, and span identifiers so the final answer can point to the exact evidence the reranker selected.
Evaluate the whole ranking path, not one score
Create a test set with required evidence labels, not only expected answers. Include exact identifiers, paraphrases, ambiguous terms, outdated versions, strong distractors, duplicate passages, multi-part questions, permission exclusions, and questions with no adequate source. These cases reveal whether the reranker recognises evidence or merely rewards topical language.
Measure candidate recall before reranking; MRR or nDCG after reranking; and evidence coverage, redundancy, and citation precision after set selection. Then measure answer correctness, faithfulness, and abstention. A higher ranking metric is useful only if the evidence that survives actually supports better answers within the same context, latency, and cost budget.
Run ablations. Compare the first-stage order with reranked order, relevance-only selection with diversified selection, and each metadata feature with it removed. Inspect failures by query type and source type. Averages can hide a reranker that improves general questions while demoting tables, short clauses, or domain terminology that matters most to the workflow.
Skip reranking when deterministic lookup already solves the task, the corpus is tiny enough for direct inspection, or the first-stage ordering meets the measured requirement. Add it when candidate recall is healthy but the best evidence is routinely buried, repeated, or mixed with plausible distractors.
If a real retrieval question returns the right document but repeatedly selects the wrong evidence, share the anonymised query and candidate shape at sahil@granveo.com. That boundary is usually enough to locate whether the fault sits in retrieval, reranking, evidence selection, or context assembly.
Sources and further reading
- 1Passage Re-ranking with BERT
Nogueira and Cho, 2019 — Establishes the cross-encoder pattern of scoring each query and candidate passage jointly for reranking.
- 2ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Khattab and Zaharia, SIGIR 2020 — Introduces late interaction for fine-grained query–document matching with reusable document representations.
- 3RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
Yu et al., NeurIPS 2024 — Studies one instruction-tuned model used for both context ranking and answer generation.
- 4BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models
Thakur et al., NeurIPS 2021 — Compares retrieval and reranking architectures across diverse domains and highlights generalisation and cost trade-offs.
- 5The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries
Carbonell and Goldstein, SIGIR 1998 — Introduces maximal marginal relevance for balancing query relevance with novelty and reducing redundancy.
- 6Lost in the Middle: How Language Models Use Long Contexts
Liu et al., TACL 2024 — Shows that relevant-information position can affect how reliably language models use long contexts.
Continue the conversation
Where does context get lost in your work?
I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.