Source Fidelity and Citation Reliability in Grounded LLM Outputs
Four distinct failure modes explain why citations still hallucinate despite retrieval.

Retrieval-augmented generation does not stop a large language model from hallucinating. It changes what the hallucination looks like: instead of fabricating an answer from nothing, the model ignores the context it was given, or cites a passage that never said what the output claims it said. A pipeline can pass every dashboard check, latency, retrieval hit rate, answer relevance, and still hand a user a confidently wrong statement wrapped in a real-looking citation.
The organizational cost of this gap lands in three different places at once. Developers get a "wrong answer" ticket with no indication of which stage of the pipeline produced the error. Compliance staff inherit a user-facing claim with no trace back to the document that supposedly supports it. Product teams watch trust erode even when the citations are technically present, because a citation that doesn't hold up under inspection is worse than no citation.
Legal research makes the stakes concrete. Studies of retrieval-augmented systems on legal datasets have reported hallucination rates between 58 and 80 percent for general-purpose large language models on legal tasks, and hallucinated citations have continued to appear in actual court filings as recently as April 2026. That a problem this well documented keeps reappearing in professional, reviewed legal work says something about how hard grounding and citation fidelity are to solve with generic fixes, even when the people using the tool are trained to check their sources.
The four failure modes that grounding conflates into one
Calling all of this "hallucination" flattens four distinct problems into one number, and that number tells a team nothing about which one it's facing. A single aggregate score can be high while any one of the four is actively broken, and because the score doesn't say which, teams end up applying the same fix, usually a prompt change or a bigger context window, to failures that have nothing in common with each other.
Factual failure is the closed-book case: the model states something false about the world, independent of anything it retrieved. This is the error retrieval was supposed to prevent, and it's also the one retrieval cannot fix on its own, because the mistake has nothing to do with what was fetched.
Faithfulness failure happens downstream of a successful retrieval. The right documents come back, but the generated answer drifts from them, contradicts them, or stretches past what they actually support. The retrieval stage did its job. The generation stage didn't. DS@GT ARC's submission to CLEF 2026 found exactly this pattern in frontier models: GPT-5.5 Thinking and Claude Opus 4.7 Thinking scored well on ROUGE, on BERT-based metrics, and on LLM-judged answer relevance, while RAGAs-based faithfulness evaluation showed those same models frequently failed to ground their answers in the material they cited. Fluent, relevant-sounding text and a faithful account of the source turned out to be two separate properties, and a model can have one without the other.
Citation failure is narrower and easier to miss. The cited document and the retrieved chunk are both real. But the specific span the model points to does not contain the claim attributed to it. FutureAGI treats this as a distinct chunk-attribution problem, separate from groundedness scoring altogether, because a system can be well-grounded in the sense that it drew from the right material and still misattribute a claim to the wrong line within that material.
Each of these four modes needs its own evaluator: groundedness checks for grounding failures, RAGAs-style faithfulness scoring for faithfulness failures, chunk attribution for citation failures, and factual accuracy checks for the closed-book case. Running one aggregate metric across all four collapses four separate signals into a number that can't tell a team where to look.
How retrieval architecture determines which failure modes appear first
Which of these four modes a system is most likely to hit is not random. It follows from choices made at the retrieval stage, well before the generator ever sees a prompt, so fixing grounding at the generation layer alone addresses only part of the problem.
Research on legal retrieval-augmented systems defines and measures a retrieval-stage failure called Document-Level Retrieval Mismatch, or DRM, where the retriever pulls information from an entirely wrong source document. DRM is a structural cause of both grounding and citation failures downstream, and it happens independent of how good the generator is. Dahl et al.'s 2024 finding of 58 to 80 percent hallucination rates for general-purpose models on legal tasks comes from exactly this kind of environment: large collections of documents that share vocabulary, formatting, and structure while differing in identity. Dense vector search, built to match meaning rather than identity, struggles to tell these documents apart; that weakness is why DRM is concentrated so heavily in legal corpora specifically.
The fix that research validates is Summary-Augmented Chunking, or SAC: each chunk gets enriched with a synthetic summary of the whole document it came from, so the chunk carries document-level context that standard chunking throws away. A chunk pulled from page 40 of a filing can otherwise look identical to a chunk pulled from page 40 of an unrelated filing with similar phrasing; the summary gives the retriever a signal that distinguishes the two. A generic summarization approach outperformed summarization tuned by domain experts, so the fix is an architectural correction rather than something that needs legal-specific tuning.
HyDE, or Hypothetical Document Embeddings, shows the same upstream-to-downstream chain from a different angle. HyDE improves recall on vague or underspecified queries by generating a hypothetical answer and searching for documents that resemble it. But if that hypothetical answer itself contains a hallucinated detail, the retriever goes looking for chunks that match the hallucination rather than the truth, and whatever comes back feeds a faithfulness failure into generation. The decision to use HyDE is made at the retrieval stage, and its consequence appears at the generation stage.
None of this is visible if retrieval quality gets collapsed into one number. Context Precision, Context Recall, and Mean Reciprocal Rank measure different things, and a retriever can score well on one while failing the other two. A single retrieval score hides what a single hallucination score hides, just one stage earlier in the pipeline.
Clinical and scientific QA show the failure modes most clearly
General-purpose benchmarks tend to blur these four modes together because nothing in the evaluation forces answer and evidence apart. Domains where answer-evidence alignment gets formally scored make the four modes visible, because each one carries a distinct, traceable cost when a patient or a researcher is on the other end of the answer.
The HealthNLP_Retrievers system, built for the ArchEHR-QA 2026 shared task, takes on clinical electronic health record question answering, where patient-authored questions have to be answered using only evidence identified in that patient's own record, with every claim in the response linked to a specific supporting sentence. That's a many-to-many alignment requirement built into the task itself, not a scoring rubric applied after the fact. The system runs a four-stage cascaded pipeline: query reformulation, evidence scoring, grounded response generation, and answer-evidence alignment. Each stage maps to one of the four failure modes. Reformulation keeps the question properly scoped, which addresses factual failure. Evidence scoring addresses grounding. Constrained generation addresses faithfulness. Alignment addresses citation fidelity. The ArchEHR-QA 2026 task itself separates these as distinct subtasks, question interpretation, evidence identification, answer generation, answer-evidence alignment, which is the same division the four-mode argument makes on independent grounds.
DS@GT ARC's CLEF 2026 work on scientific question answering arrives at the same division from the opposite direction. Instead of engineering grounding and faithfulness in from the start, its pipeline corrects for their absence after the fact. Corrective RAG, or CRAG, filters retrieved chunks before generation to address grounding failures. CiteFix prunes citations that don't hold up after generation, to address citation failures. These are two separate mechanisms, aimed at two separate failure modes, run in sequence.
The CLEF 2026 result already described is the clearest evidence that these failure modes don't travel together: frontier models maximized fluency and answer relevance while failing RAGAs faithfulness evaluation. A team monitoring only ROUGE or BERT-based scores would have seen a system performing well and would have had no way to detect that its citations weren't holding up. The clinical system shows what it looks like to design against all four modes from the start. The scientific QA system shows what it looks like to catch two of them after the model has already produced an answer. Read together, they show the same four-mode structure holding across both ends of the pipeline.
Detection strategies that match each failure mode to its own signal
Because each failure mode has a different cause, each one needs its own detection signal placed at the point where that mode actually occurs. A monitoring setup that checks only the quality of the final output will miss most of this, because by the time an answer is final, the evidence of where it went wrong has usually been smoothed over.
Faithfulness evaluation in the RAGAs style works at the level of individual claims: it breaks an answer down into atomic statements and checks each one against the retrieved context separately. That's why it caught the frontier-model failure from CLEF 2026 that sentence-level or document-level grounding scores had missed. A model can produce an answer that reads as grounded in aggregate while three of its ten individual claims have no support in the retrieved material, and claim-by-claim checking is what reveals that.
Chunk attribution solves a narrower but equally necessary problem: it maps each generated claim back to a specific chunk ID, which makes it possible to confirm whether the cited span actually supports the claim attached to it. Without chunk-level attribution, this gap is invisible to every other metric in the pipeline, including faithfulness scoring, because faithfulness checks whether the claim matches the retrieved context in general, not whether the specific cited span backs it up.
The strongest version of this control acts before an evaluation pass: if structured output requires a chunk ID and a quoted span for every claim, and deterministic validation checks that the quoted span exists verbatim in the retrieved context, you catch the problem before the response ever reaches the user. This is a generation-time constraint rather than a post-hoc evaluation: one prevents the failure, the other only reports it after it already happened.
The CRAG and CiteFix combination from DS@GT ARC's CLEF 2026 work applies its two controls at two separate points in the pipeline: CRAG filters chunks before generation, restricting what the model is even allowed to see, while CiteFix prunes citations after generation, removing references that don't hold up under scrutiny. That two-stage design isn't incidental. Grounding failures and citation failures occur at different points in the pipeline, so intervention has to happen at those different points; one late-stage check catching both cannot be assumed.
Retrieval metrics and generation metrics get tracked separately, because conflating them hides which stage produced a given failure. Context Precision, Context Recall, and MRR belong on one side; Groundedness, Faithfulness, and ChunkAttribution belong on the other. Merging them into a single dashboard reproduces the exact averaging problem that a single hallucination score creates, just one layer further into the system.
Pipeline controls that address each mode independently, from retrieval through generation
Reliable grounded output depends on controls placed independently at the retrieval stage, the context-shaping stage, and the generation stage. A constraint bolted onto the generator alone cannot undo a grounding failure that started two stages earlier in retrieval.
At the retrieval stage, hybrid search that combines dense vector matching with sparse keyword search such as BM25 addresses a specific blind spot: purely semantic retrieval often misses exact entity matches, names, case numbers, product codes, because semantic similarity and exact identity aren't the same thing. Fixing this is a precondition for reducing both grounding and citation failures downstream. Summary-Augmented Chunking, validated in the legal RAG research already described, injects document-level context into each chunk specifically to prevent Document-Level Retrieval Mismatch, addressing a retrieval-stage failure with a retrieval-stage fix applied before generation. Index freshness matters: an index that isn't refreshed on a cadence matched to how often the underlying source documents change will ground answers in information that used to be correct and no longer is, a factual failure that originates in the ingestion pipeline, which is what must own fixing it.
At the context-shaping stage, pre-generation chunk filtering, the CRAG approach, strips out distractor documents and low-relevance passages before the generator ever sees them, which narrows the context window down to material that can actually support the answer. One retrieval system pairs this kind of filtering with a recall-biased evidence scoring heuristic that includes dynamic fallback tiers, specifically to avoid over-filtering. Filtering too aggressively starves the generator of supporting evidence, and a generator left without evidence either fabricates an answer or has no basis to answer. Shaping context this way, selecting, filtering, ranking, and structuring what the retriever hands off, is a distinct engineering discipline from retrieval itself. The quality of that context window sets the ceiling on how grounded the final answer can possibly be, regardless of how capable the underlying model is.
At the generation stage, structured output that requires a chunk ID for every claim enforces citation attribution mechanically: a claim without a traceable source span simply can't be produced in valid output. CiteFix-style post-generation pruning, from the CLEF 2026 work, adds a second layer on top of that by checking strict entailment between each generated claim and its cited material after generation, catching cases that the structured-output requirement alone lets through.
No single one of these controls, placed at just one stage, closes the gap between a system that looks grounded and one that actually is. In the CLEF 2026 result, models that scored well on fluency and relevance still failed faithfulness evaluation, because nothing in their generation process enforced a connection to the cited material. Reliable grounded output requires treating retrieval, context shaping, and generation as three separate places where the four failure modes can originate, and building a distinct, independent check for each one.


