All field notes
Retrieval and context10 min read

How to Build RAG for Questions That Need Multiple Sources

Build multi-hop RAG around an explicit evidence chain. Decompose the question into dependent steps, retrieve the first supported fact, carry verified entities and constraints into the next search, and preserve the source for every link. If one required hop has no adequate evidence, narrow the answer or say what is missing instead of letting the model bridge the gap from memory.

By Sahil Maheshwari

A complex question split into connected retrieval steps that join evidence from several sources into one traceable answerSOURCESConnected contextrelationships stay visibleSYNTHESIS

Why one retrieval pass fails on multi-source questions

A normal RAG pipeline embeds the user’s question, retrieves similar passages, and asks a model to answer from them. That works when one passage contains most of the answer. It is much less reliable when the question names one thing but the decisive source names another thing that can only be discovered along the way.

Consider a hypothetical insurance question: “Which endorsement replaced the exclusion discussed in the broker memo, and when did it take effect?” The memo may identify a policy form. The policy history may identify the superseding endorsement. The endorsement itself supplies the effective date. No single passage needs to resemble the complete question. The second search becomes possible only after the first source reveals the form number; the third depends on the endorsement found in the second.

This is a multi-hop question. The useful retrieval query changes as evidence arrives. IRCoT formalises that idea by interleaving retrieval with reasoning: derived information guides the next retrieval step, and newly retrieved material guides subsequent reasoning. Its results across four multi-step benchmarks support the pattern, though benchmark gains do not guarantee the same performance on a private corpus.

Treat retrieval as a process with state, not a single similarity search. Keep the original question, resolved facts and their source spans, unresolved variables, search constraints, and the next subquestion. That record also makes the final answer reviewable.

For multi-hop RAG, the unit of work is not a chunk. It is a claim-to-source chain whose missing links stay visible.

Recognise the shape of the question first

Not every question that mentions several nouns needs multi-hop retrieval. “What is the cancellation period in policy X?” may be a direct lookup. A question becomes multi-hop when answering one part supplies a value needed to find or interpret another part. Classifying that dependency before retrieval helps the system choose an appropriate plan.

Bridge questions follow a chain: find the company acquired by an organisation, then find that company’s current product. Comparison questions retrieve parallel evidence about two or more entities before applying the same criterion. Intersection questions identify an entity that satisfies several independently sourced conditions. Fan-out questions gather many related facts, such as every project meeting a stated rule, then aggregate them.

The benchmark literature makes these distinctions concrete. HotpotQA includes questions that require multiple Wikipedia documents, explicit supporting facts, and comparison reasoning. FanOutQA focuses on questions that must be decomposed into many searches and then aggregated. These shapes need different execution: a bridge is sequential, a comparison can retrieve branches in parallel, and fan-out needs a completeness rule so the system knows whether it found enough items.

Start with direct retrieval when the question may be simple. Escalate only when the answer requires unresolved entities, cross-document comparison, or unsupported aggregation. This keeps latency and cost proportional to the question.

  • Direct: one source can support the answer without discovering another entity first.
  • Bridge: an answer from hop one becomes the search key for hop two.
  • Comparison: several evidence branches answer the same subquestion before synthesis.
  • Intersection: the answer must satisfy conditions established by different sources.
  • Fan-out: many related results must be found and combined under a stopping rule.

Decompose the question without losing its constraints

Represent the plan as a small dependency graph rather than a loose list of generated search queries. Each node should state a subquestion, its required inputs, its expected output type, and the evidence needed to accept that output. An edge means one node cannot run until another supplies a value. Independent branches can run together; dependent branches must wait.

Keep the original wording beside the plan. Decomposition can quietly drop a date range, jurisdiction, negation, or comparison criterion. Check that every constraint appears in a subquestion or global filter, and that each requested output has a supported result.

Do not let the planner invent the bridge entity. If the first step asks which policy form the memo discusses, the accepted form number must come from a retrieved span, not from the model’s prior knowledge. Store the exact value, a normalised identifier if one exists, the source, and the confidence rule that accepted it. The next query can use the identifier, aliases, document type, and time boundary together.

Decomposition itself needs evaluation. MuSiQue builds two- to four-hop questions and applies filters intended to reduce disconnected reasoning shortcuts. That design points to a useful production test: hide one prerequisite source and confirm that the system cannot still produce the complete answer with unjustified confidence.

A good plan tells you not only what to retrieve next, but which earlier evidence makes that retrieval valid.

Retrieve iteratively and carry only verified facts forward

At each hop, construct the query from the current subquestion plus accepted facts. Use stable entity IDs where available, but include names and aliases. Apply permissions, tenant boundaries, effective dates, document status, and other hard constraints on every hop.

Retrieve several candidates, rerank them against the subquestion, and require evidence for any value entering state. A passage may mention the right entity without supporting the needed relationship. Check whether its source span entails that relationship under the question’s time and scope. Keep disagreements visible.

Search breadth should follow uncertainty. A precise identifier can narrow the next query; an ambiguous name may require entity resolution across several sources. Preserve alternatives until evidence rules them out, or an early error will contaminate every later hop.

The MultiHop-RAG benchmark uses a news collection, multi-hop queries, answers, and supporting evidence to test this problem directly. Its authors found that evaluated RAG systems still struggled with multi-hop retrieval and answering. The benchmark is valuable because it scores supporting evidence, but its English news setting is not a substitute for tests on your documents, identifiers, and access rules.

  • Query with the unresolved subquestion and evidence-backed bridge values.
  • Rerank for the required relationship, not for general topical similarity.
  • Attach source spans and scope to every fact admitted into working state.
  • Retain competing entities or claims until evidence resolves the ambiguity.
  • Stop after a hop budget or when new retrieval no longer closes a required gap.

Choose flat, iterative, or graph retrieval deliberately

You do not need a knowledge graph for every multi-source question. Iterative search is often enough when bridge values are explicit, names are stable, and questions need only a few hops. Its failures are also easier to inspect.

Graph retrieval becomes attractive when the same entities and relationships recur across many questions: people, organisations, policies, claims, versions, citations, or dependencies. A graph can traverse typed relationships, preserve aliases, and expand a neighbourhood around a verified entity before reranking the underlying source passages. The graph should point back to evidence; a relationship extracted by a model is not automatically a fact.

HippoRAG is one concrete approach. It builds a knowledge graph from extracted entities and relations, then uses Personalized PageRank to retrieve connected evidence for multi-hop questions. The paper shows why graph structure can join clues that are weak semantic matches in isolation. It does not prove that every corpus benefits from graph construction, nor does it remove extraction errors, update work, or permission concerns.

A hybrid design can combine metadata filters, lexical and vector search, graph expansion, and reranking. Add a layer only when corpus-specific evaluation shows which failure it fixes. Complexity without a named failure mode is another place for context to disappear.

Make the evidence chain part of the answer

Give synthesis structured evidence, not a pile of chunks. Group passages by subquestion and include each accepted intermediate claim, source span, source identity, date or version, and unresolved conflict. Answer only the parts whose dependency path is complete.

Citations should follow claims. If source A identifies the policy form and source B says which endorsement superseded it, cite both near the combined claim. For comparison or fan-out answers, show the evidence for each branch before presenting the aggregate conclusion.

Define failure behaviour before deployment. When a required hop is missing, the system can answer the supported portion, name the missing source or relationship, ask a clarifying question, or abstain. It should not fill the gap from model memory and present the result as retrieved. A visible partial chain is more useful than a complete-looking answer that cannot be audited.

Keep an inspectable trace: planned hops, queries, candidates, accepted facts, cited sources, and the stop reason. It distinguishes a planning failure from a retrieval miss, extraction error, or unsupported synthesis.

If the system cannot show which source supports each hop, it has not finished the multi-source task.

Evaluate multi-hop RAG one dependency at a time

Final-answer accuracy is necessary but insufficient. A model can guess the right answer from prior knowledge or reach it through the wrong documents. Record whether the plan contained the required hops, whether retrieval found evidence for each hop, whether the accepted intermediate facts were correct, and whether the final claims were entailed by that evidence.

Build an evaluation set from real question shapes in your workspace. Include direct questions so the router is not rewarded for over-planning. Add bridge, comparison, intersection, and fan-out cases; ambiguous aliases; superseded versions; contradictory sources; permission boundaries; and questions with one deliberately missing link. Label the smallest sufficient evidence chain rather than every vaguely relevant passage.

Test for shortcuts. Remove one supporting source. Swap the bridge entity for a similar one. Add a topically similar distractor. Change a date or jurisdiction. A reliable system should change its retrieval and answer for the right reason, and it should decline to complete a chain when a prerequisite disappears. MuSiQue’s emphasis on connected reasoning and HotpotQA’s sentence-level supporting facts are useful precedents for this style of evaluation.

Track hop-level evidence recall, chain completion, intermediate-claim accuracy, citation precision, unsupported-claim rate, correct abstention, latency, and retrieval cost. Slice results by question shape and number of hops. Averages can conceal a system that handles two-hop bridges well but fails every fan-out question.

Public benchmarks are starting points, not deployment certificates. The strongest test set contains the questions users actually ask and the sources they are genuinely allowed to use.

If a real question in your workspace needs several documents and your system keeps losing a link between them, share the question shape and the sources involved at sahil@granveo.com. The broken evidence chain is the most useful place to begin.

  • Planning: did the system identify every dependency and preserve every constraint?
  • Retrieval: did each hop return the smallest sufficient supporting evidence?
  • State: were only evidence-backed entities and claims carried into later hops?
  • Synthesis: does every final claim map to a complete chain of source spans?
  • Failure: did the system expose missing, conflicting, or inaccessible evidence?

Sources and further reading

  1. 1
    Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

    Trivedi et al., ACL 2023 — Introduces IRCoT, which alternates retrieval and reasoning so intermediate information can guide later searches.

  2. 2
    HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

    Yang et al., EMNLP 2018 — Provides multi-document questions, comparison cases, and sentence-level supporting facts for explainable evaluation.

  3. 3
    MuSiQue: Multihop Questions via Single-hop Question Composition

    Trivedi et al., TACL 2022 — Builds connected two- to four-hop questions with filters intended to reduce disconnected reasoning shortcuts.

  4. 4
    MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

    Tang and Yang, 2024 — Evaluates retrieval and answering for multi-hop questions over an English news corpus with supporting evidence.

  5. 5
    HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models

    Gutierrez et al., NeurIPS 2024 — Demonstrates a graph-based multi-hop retrieval approach using extracted relations and Personalized PageRank.

  6. 6
    FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models

    Zhu et al., ACL 2024 — Tests fan-out questions that require decomposing, retrieving, and aggregating evidence from multiple documents.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.