All field notes
Retrieval and context10 min read

How to Build Multilingual RAG Without Losing Meaning

Build multilingual RAG as a language-aware retrieval system, not an English pipeline with translation bolted on. Keep original text and language metadata, route each query through same-language and cross-language search when needed, preserve names and domain terms, rerank the combined evidence, and generate in the user’s language while citing the source in its original form.

By Sahil Maheshwari

A multilingual question branching into original-language, translated, and cross-lingual retrieval paths that reconnect at cited source evidenceSOURCESConnected contextrelationships stay visibleSYNTHESIS

The short answer: define the language path first

A multilingual RAG system has to answer three separate questions: what language the user is using, what languages may contain the evidence, and what language the answer should use. Those answers are not always the same. A Hindi question may need an English circular and a bilingual email, then require a Hindi answer with the decisive English sentence visible in the citation.

Start by writing down the language path for each supported workflow. Same-language retrieval searches Hindi with Hindi or French with French. Cross-language retrieval searches a corpus in one language using a query in another. Mixed retrieval searches documents, queries, or both that contain more than one language. Each path creates different failure modes, so a single ‘multilingual’ switch is not an architecture.

The research reflects that distinction. MIRACL evaluates same-language retrieval across 18 languages, while XOR QA was designed for questions whose useful evidence may exist in another language. A model that scores well on the first task is not automatically reliable on the second.

For a first release, support a small, explicit language matrix. Record query language, document language, answer language, script, translation route, and any fallback used. If the system cannot confidently identify the path or no supported route reaches the evidence, it should say so instead of returning the nearest English passage.

Language is part of retrieval intent. Treat it as a routing decision, not a display preference.

Preserve original text beside every derived version

Keep the source text exactly as received. Store its detected and declared language, script, locale, document version, section, and access rules. If you create a translation, transliteration, summary, or normalised copy for search, save it as a derived representation linked to the original chunk. Never overwrite the evidence with its translation.

This separation matters because translation can change legal force, technical meaning, politeness, tense, or uncertainty. The English word ‘coverage’ may map to several terms depending on whether a document discusses insurance scope, geographic reach, or media reporting. A translated search field is useful; it is not the authoritative source.

Chunk along the structure of the original document before translating. Keep headings, list relationships, tables, footnotes, and cross-references together where their meaning depends on one another. Translation should not silently merge two clauses or move a qualifier into another chunk. Record the translation model or service, version, time, and source and target languages so a retrieval result can be reproduced.

Also keep stable entity and term IDs. A company name may appear in Devanagari, Latin script, an acronym, and a historical spelling. Link those surface forms without forcing them into one display string. The same principle applies to product names, regulatory terms, place names, and personal names.

  • Original chunk: authoritative wording, location, version, permissions, language, and script.
  • Derived text: translation, transliteration, normalisation method, model version, and timestamp.
  • Terminology: canonical concept ID with approved terms, aliases, acronyms, and prohibited substitutions.
  • Provenance: an explicit link from every derived representation back to the source span.

Choose retrieval routes by evidence distribution

There are three practical retrieval strategies. Translate the query into each corpus language and search language-specific indexes. Embed queries and documents with a multilingual model in one shared vector space. Or combine both routes with lexical search and merge their candidates. The right choice depends less on the number of languages than on where the evidence actually lives.

Query translation makes exact lexical matching available in the target language and lets you keep established monolingual indexes. It can work well when terminology is controlled, but one translation may erase a useful ambiguity. Generate a small set of approved variants for important terms rather than a free-form cloud of paraphrases.

Multilingual embeddings avoid an explicit translation step and can retrieve semantically related text across languages. Coverage claims still need local testing. MMTEB spans hundreds of tasks and a very broad range of languages, yet its authors also show that model leadership varies across benchmark slices. ‘Supports 100 languages’ does not mean equal retrieval quality for every domain, script, or language pair.

Use hybrid retrieval when identifiers and specialised vocabulary matter. In Mr. TyDi, dense retrieval underperformed BM25 in the reported zero-shot experiments, while sparse and dense signals worked usefully together. That benchmark covers monolingual retrieval, not your private corpus, but it is a strong warning against discarding lexical search merely because the interface is multilingual.

A sound default is to search the original query, approved translations, and transliterations through both lexical and multilingual dense indexes, then deduplicate by source chunk. Apply language, date, document type, and permission filters before reranking. Keep the scores and route labels so evaluation can reveal which path actually found the evidence.

Do not translate the entire knowledge base just to make retrieval look monolingual. Preserve originals and add search representations where they earn their cost.

Handle code-switching, transliteration, and terminology

Real questions rarely arrive as clean textbook sentences. A user may ask, ‘renewal ka premium revised quote se match karta hai?’ The sentence mixes English and Hindi, uses Latin script for Hindi, and contains insurance terms that should not be translated mechanically. Whole-query language detection will label it poorly or miss the evidence-bearing words.

Detect language and script at the span level when mixed queries are common. Preserve the original string, then produce controlled variants: native-script transliteration where useful, normalised spellings, acronym expansions, and approved domain equivalents. Search all variants, but keep them tied to the same query so duplicated results do not dominate ranking.

Create the terminology layer with subject-matter experts. It should distinguish exact identifiers from translatable concepts and note when a term must remain in the source language. Include deprecated names and regional variants. Review the logs for unmatched terms; users will reveal vocabulary that a generic language model never sees.

Code-switching deserves its own test set. The GLUECoS benchmark found that multilingual models could improve when trained on code-switched data across its tasks. It is not a RAG benchmark, but it supports a practical conclusion: multilingual competence measured on clean monolingual text does not settle performance on mixed-language input.

Rerank across languages without erasing provenance

Candidate generation should favour recall; reranking should decide which evidence deserves prompt space. Give the reranker the original query, its approved variants, each candidate’s original text, and a translation only when the model needs one. Include language and document metadata, but never let a preferred language override relevance or authority.

Calibrate reranking by language pair. Scores produced for Hindi-to-English retrieval may not be comparable with scores for Hindi-to-Hindi retrieval. Instead of taking the global top five, reserve candidates from each viable route, normalise or learn route-specific scores, then rerank the pooled set. Log when a translated representation wins over the original-language route.

Construct context in evidence units, not a pile of translations. Show the generator the authoritative source span, a clearly labelled translation if required, its surrounding qualifier, and the citation target. If two languages contain independent versions of a policy, do not assume they are equivalent. Compare dates, jurisdictions, authors, and version relationships before merging their claims.

The BGE-M3 paper demonstrates one model design that combines multilingual, dense, sparse, and multi-vector retrieval. That is useful evidence that these signals can share a system, not a reason to depend on one embedding model. Keep interfaces modular enough to replace the retriever or reranker without rebuilding provenance and evaluation.

Answer in the user’s language, cite the source language

Tell the generator the required answer language explicitly. Also tell it which passages are authoritative, which are translations, and which terms must remain unchanged. Ask it to preserve numbers, dates, names, currency, units, modality, and negation. A fluent answer that turns ‘may’ into ‘will’ is not a successful translation.

Citations should open the original passage. When readers may not know that language, display a labelled translation beside it, not instead of it. For high-stakes material, let a reviewer compare both. If the answer depends on interpreting a translated phrase, say that plainly rather than presenting the interpretation as verbatim evidence.

The 13-language study Retrieval-augmented generation in multilingual settings found that multilingual retrievers and generators still needed task-specific prompting to produce answers in the user’s language. It also reported code-switching, fluency errors, irrelevant retrieval, and incorrect reading of supplied documents among the remaining problems. Those are separate failure classes and should stay separate in monitoring.

Do not translate citations after generation from memory. Carry source spans through retrieval, reranking, and generation as immutable evidence objects. The answer can be rewritten; the cited text, document ID, language, version, and location should not be reconstructed by the model.

A translation helps the reader. The original source proves the claim. A trustworthy interface keeps both roles visible.

Evaluate every language route you promise

Build evaluation sets from real work, stratified by query language, evidence language, answer language, script, domain, and route. Include same-language, cross-language, translated, transliterated, and code-switched questions. Add cases with names that have multiple spellings, ambiguous terms, untranslated identifiers, conflicting translations, and no answer in any supported language.

Measure retrieval recall before answer quality. Track whether the authoritative chunk appeared, which route found it, reranker accuracy, citation correctness, answer-language compliance, preservation of named entities and numbers, and appropriate abstention. Evaluate translation accuracy for the terms that affect decisions, not only overall sentence fluency.

Use native-speaking reviewers for the languages and domains that matter. Automatic metrics can miss a wrong honorific, a softened prohibition, or a transliteration that points to the wrong person. Public benchmarks test general capability; they cannot reproduce your terminology, permissions, document versions, or distribution of user questions.

Roll out one language pair and one workflow at a time. Compare multilingual retrieval against a strong lexical and translation baseline. If nearly all authoritative evidence is in one language and users accept answers in that language, a careful query-translation layer may be simpler than a fully cross-lingual stack. Complexity is justified only when it improves access without obscuring meaning.

If a real mixed-language question repeatedly finds a plausible translation instead of the controlling source, share the language path and document pattern at sahil@granveo.com. That example is more useful than a generic multilingual demo because it exposes the exact route where meaning was lost.

  • Retrieval: authoritative evidence found for every supported query–corpus language pair.
  • Ranking: relevant original sources outrank fluent but weaker translations.
  • Generation: requested language, preserved terms, numbers, modality, and uncertainty.
  • Grounding: citations resolve to original spans with translations clearly labelled.
  • Safety: permissions, unsupported languages, ambiguity, and no-answer cases remain visible.

Sources and further reading

  1. 1
    Retrieval-augmented generation in multilingual settings

    Chirkova et al., KnowLLM 2024 — Evaluates an end-to-end multilingual RAG baseline across 13 languages and documents generation, evaluation, and code-switching issues.

  2. 2
    Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

    Zhang et al., TACL 2023 — Introduces a native-speaker-judged monolingual retrieval dataset spanning 18 typologically diverse languages.

  3. 3
    Mr. TyDi: A Multi-lingual Benchmark for Dense Retrieval

    Zhang et al., MRL 2021 — Compares multilingual dense and lexical retrieval signals across eleven languages and motivates sparse–dense hybrids.

  4. 4
    XOR QA: Cross-lingual Open-Retrieval Question Answering

    Asai et al., NAACL 2021 — Defines cross-lingual open retrieval for questions whose relevant evidence may not exist in the query language.

  5. 5
    BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings

    Chen et al., 2024 — Presents a multilingual model supporting dense, sparse, and multi-vector retrieval within one framework.

  6. 6
    MMTEB: Massive Multilingual Text Embedding Benchmark

    Enevoldsen et al., ICLR 2025 — Expands embedding evaluation across hundreds of tasks, many domains, and a broad set of languages.

  7. 7
    GLUECoS: An Evaluation Benchmark for Code-Switched NLP

    Khanuja et al., ACL 2020 — Evaluates multilingual models on English–Hindi and English–Spanish code-switched language tasks.

Continue the conversation

Where does context get lost in your work?

I am speaking with researchers, founders, and operators about the handoffs, evidence, and decisions that are hardest to keep connected.