Diagnosing Conflicts in RAG Pipelines
A RAG pipeline that retrieves the right document can still produce the wrong answer, because retrieval and generation fail independently and most teams only instrument one of the two. When the retrieved context contradicts what the model already “believes” from pretraining, or two retrieved passages disagree with each other, the system has to resolve a conflict — and most production RAG stacks have no explicit mechanism for that resolution at all. They just generate, and whichever signal the model happens to weight more heavily wins, silently.
The three places a RAG pipeline actually fails
It’s tempting to treat “the RAG answer was wrong” as one failure mode, but it’s useful to separate it into three, because each needs a different fix.
Retrieval failure: the right document was never surfaced. This is the failure mode most teams instrument, because it’s the easiest to detect — you can check whether the gold document appears in the top-k results independent of anything the generator does.
Generation failure with correct retrieval: the right document was retrieved, but the model didn’t use it faithfully — it hallucinated around it, or ignored it in favor of parametric knowledge. This is harder to detect because it requires comparing the generated answer against the retrieved context specifically, not just against a gold answer.
Conflict failure: the retrieved context and the model’s parametric knowledge disagree, or two retrieved passages disagree with each other, and the system resolves the disagreement in a way nobody explicitly designed. This is the failure mode most production systems have zero visibility into, because most evaluation setups only check the final answer against a gold label — they never ask why the model picked the answer it picked when its inputs were in tension.
Why conflict resolution is invisible by default
A standard RAG evaluation pipeline looks at (query, retrieved context, generated answer, gold answer) and scores whether the generated answer matches the gold answer, sometimes with an LLM-as-judge computing something like a faithfulness or context-adherence score. That setup has a structural blind spot: it can tell you the final answer was wrong, but it can’t tell you whether the model contradicted the retrieved context, contradicted its own parametric knowledge, or was fed genuinely contradictory retrieved passages in the first place. Those are three different bugs with three different fixes, and a single aggregate faithfulness score collapses all of them into one number.
This matters more than it might seem, because the fix for “the model ignored correct retrieved context” (better prompting, stronger instruction to prioritize context) is close to the opposite of the fix for “the retrieved context was itself wrong or stale” (better retrieval, source filtering, freshness checks) — and applying the wrong one doesn’t just fail to help, it can actively make the other failure mode worse. Tuning a model to trust retrieved context more aggressively, when the underlying problem is retrieval surfacing stale or low-quality passages, means the system will now confidently repeat whatever bad information it retrieves.
A diagnostic checklist before you touch the model
Before changing a prompt, a retrieval threshold, or a reranker, it’s worth walking through a fixed sequence of checks on a sample of failures, because the order matters — later checks are only meaningful once earlier ones have ruled out simpler explanations:
- Was the correct document retrieved at all? If not, this is a retrieval problem — check embedding quality, chunk boundaries, and query reformulation before looking anywhere else.
- Did the retrieved passages agree with each other? If two retrieved chunks state contradictory facts, the generator was asked to adjudicate a conflict the retrieval layer should have flagged or filtered.
- Did the generated answer match the retrieved context, independent of whether it matched the gold answer? A model can be perfectly faithful to bad context (retrieval’s fault) or unfaithful to good context (generation’s fault) — conflating these produces the wrong fix.
- Where the model deviated from context, was it toward its own parametric knowledge, or toward a different retrieved passage? These point to different mitigations: the former needs stronger context-grounding instructions or a lower generation temperature; the latter needs deduplication or source-authority weighting at retrieval time.
Comparison: symptom versus root cause
| Observed symptom | Plausible root cause | Wrong fix people often try first |
|---|---|---|
| Wrong final answer, right document retrieved | Generation ignored context | Retuning the retriever (won’t help) |
| Wrong final answer, wrong document retrieved | Retrieval failure | Prompting the generator to “be more careful” (won’t help) |
| Answer contradicts itself across similar queries | Conflicting retrieved passages, no dedup | Raising generation temperature down (treats symptom, not cause) |
| Confident answer that’s stale | Retrieval has no freshness signal | Better prompting for “say if unsure” (model isn’t unsure — it’s wrong) |
A worked example
Suppose a support-bot RAG system is asked “What’s the refund window for this product?” and answers “30 days” when the correct, current policy is 14 days. Walking the checklist: the retriever surfaced two documents — an old policy page (30 days, still indexed) and the current policy page (14 days). Both were retrieved; the generator picked the older one. This is not a generation-faithfulness problem in the sense of “the model made something up” — it faithfully reported one of its two retrieved sources. It’s a retrieval-conflict problem: two contradictory sources reached the generator with no signal about which one was authoritative or current, and no mechanism forced the generator to notice the contradiction rather than pick one and move on. The fix is upstream of the model: retire the stale document from the index, or attach recency/authority metadata that a reranker or the generator can use to prefer the current source when passages conflict.
Limitations
None of this replaces having a retrieval quality metric and a faithfulness metric in the first place — the checklist assumes you can already tell, per example, whether the right document was retrieved and whether the answer matches the retrieved context. Building that instrumentation is itself nontrivial: “faithfulness” scores computed by an LLM judge have their own failure modes, including ceiling effects that make good and great systems look identical once accuracy is already high. Diagnosing conflicts is a useful second layer on top of solid retrieval and faithfulness measurement, not a substitute for it.
FAQ
Is this the same problem as hallucination?
Related but not identical. Hallucination usually refers to the model generating content unsupported by any input. Conflict failure specifically involves the model having relevant input — from retrieval, or from its own pretraining — that disagrees with other input it also has, and resolving that disagreement implicitly. A model can be conflict-prone without hallucinating in the classic sense: every fact it states was in one of its two contradictory sources.
Can reranking alone fix conflict failures?
Reranking helps with retrieval-quality conflicts (surfacing the more authoritative or current document higher), but it doesn’t help when the generator still receives multiple passages and has to decide how to handle disagreement between what made the cut. Deduplication and explicit conflict-flagging at the retrieval layer address a different part of the problem than reranking does.
How does this relate to the natural-language front end of a search system?
A natural-language query system decides what a user is asking for before retrieval happens at all — if that stage misinterprets intent, everything downstream, including conflict resolution, is working from the wrong starting point. Diagnosing conflicts assumes intent was captured correctly; if it wasn’t, that’s a fourth failure mode worth checking before the three above.
Does a bigger, more capable model make this problem go away?
Not by default. A stronger model is generally better at noticing when two passages disagree if explicitly asked to check, but nothing about scale alone causes a model to surface a conflict rather than silently pick a side — that behavior has to be specifically evaluated for and, where it’s missing, explicitly prompted or architected for.
Bottom line
Most RAG evaluation setups measure whether the final answer was right, which tells you that something failed but not where. Splitting failures into retrieval, generation-faithfulness, and conflict-resolution — and checking them in that order — turns “the RAG system is wrong sometimes” into a specific, fixable bug report, and stops teams from tuning the wrong layer of the stack in response to a symptom produced by a different one.