Citations that point at the wrong paragraph
In short: a citation can be wrong in at least seven different ways, and only two of them involve the model. Most come from how long documents were cut into chunks. Name each failure, and give each one…
- published
- read time
- 4 min
- words
- 885
- lang
- en
- filed under
- Engineering
In short: a citation can be wrong in at least seven different ways, and only two of them involve the model. Most come from how long documents were cut into chunks. Name each failure, and give each one a check that runs before the answer is shown.
A citation is a promise
In legal research, the citation is the product. A lawyer reads the answer, but what they act on is the pinpoint: this decision, paragraph 14. If they open the decision and paragraph 14 is about costs, they stop trusting every other citation too, including the good ones.
I've built citation engines over a large corpus of long documents, and the hard lesson is that "the citation was wrong" isn't one bug. It's a family of them, with different causes and different fixes. Lumping them together as "hallucination" sends you to fix the prompt when the problem is in ingestion. So I keep a list.
The failure list
Here's the taxonomy I use. It's built from the shape of the problem, not from counting incidents, so the order says nothing about how often each one happens.
| Failure | What the reader sees | Usual cause | Check that catches it |
|---|---|---|---|
| Neighbour paragraph | Cites para 13, the words are in para 14 | A chunk spanned two paragraphs and kept the first number | Quoted text must appear in the cited paragraph |
| Wrong version | Right section, today's text, old event | Index has no in-force dates | Citation carries a date; version window must contain it |
| Summary cited as the court | A headnote quoted as if a judge wrote it | Editorial summary stored in the same text as the reasons | Tag document zones; only reasons are citable |
| Dissent cited as the holding | Real words, from the losing side | The heading saying whose reasons these are was cut off | Carry the section path into every chunk |
| Tidied quotation | Quote marks around a paraphrase | The model smooths the wording | Exact match after whitespace normalising |
| Page read as paragraph | "At 14" means page 14 in one place, para 14 in another | Numbers taken from the PDF layout | Parse paragraph numbers from the text, never the page |
| Never retrieved | A plausible pinpoint to a case the system never saw | Model citing from memory | Every cited ID must be in this turn's retrieved set |
Only two rows involve the model: a tidied quotation and a citation from memory. The other five are faithful answers to text that was cut, labelled or stored badly. That changes where you spend the week.
Where the neighbour-paragraph bug comes from
The first row is the one worth drawing, because it's the most natural thing in the world to build by accident. You chunk a long decision by token count, say a few hundred tokens with some overlap. Paragraphs don't line up with token windows, so a chunk often starts partway through one paragraph and ends partway through the next. You need a paragraph number to cite, so you take the number of the paragraph the chunk starts in.
Now the passage that answers the question sits in the second half of that chunk, in the next paragraph. The model quotes it correctly. The citation says the wrong number. Every individual step was reasonable.
The fix is to chunk on the document's own structure. Decisions are numbered by paragraph, statutes by section and subsection. A chunk is one numbered unit, or a few whole units if they're short, and it carries the exact numbers it contains as metadata. A very long paragraph gets split, and both halves keep the same number. The chunker never decides a paragraph number; it copies one.
Checks that run before the answer ships
The checks in the right-hand column of the table are cheap. None needs a model. They run after generation and before the answer reaches the screen:
- Every citation must name a passage ID from this turn's retrieved set. Anything else is dropped.
- Every quotation must appear, after normalising whitespace and quote marks, inside the passage it cites.
- The paragraph or section number shown to the reader is rendered from that passage's metadata, not copied from the model's text.
- If the question has a date, the cited version's in-force window must contain it.
When a check fails, the honest options are to drop the citation and say so, or to retry with the failure named. What you shouldn't do is show it and hope.
Start with one decision and a highlighter
Pick one long decision from your corpus, the longer the better. Print the chunks your pipeline made from it, next to the paragraph numbers in the original. Mark every chunk that crosses a paragraph boundary, and every chunk that mixes a headnote with the reasons, or a dissent with the majority. If you find even one, you have the neighbour-paragraph bug waiting in production, and the fix is in the chunker, not the prompt.
related