Writing

Citations that point at the wrong paragraph

In short: a citation can be wrong in at least seven different ways, and only two of them involve the model. Most come from how long documents were cut into chunks. Name each failure, and give each one…

published
read time
4 min
words
885
lang
en
filed under
Engineering

In short: a citation can be wrong in at least seven different ways, and only two of them involve the model. Most come from how long documents were cut into chunks. Name each failure, and give each one a check that runs before the answer is shown.

A citation is a promise

In legal research, the citation is the product. A lawyer reads the answer, but what they act on is the pinpoint: this decision, paragraph 14. If they open the decision and paragraph 14 is about costs, they stop trusting every other citation too, including the good ones.

I've built citation engines over a large corpus of long documents, and the hard lesson is that "the citation was wrong" isn't one bug. It's a family of them, with different causes and different fixes. Lumping them together as "hallucination" sends you to fix the prompt when the problem is in ingestion. So I keep a list.

The failure list

Here's the taxonomy I use. It's built from the shape of the problem, not from counting incidents, so the order says nothing about how often each one happens.

FailureWhat the reader seesUsual causeCheck that catches it
Neighbour paragraphCites para 13, the words are in para 14A chunk spanned two paragraphs and kept the first numberQuoted text must appear in the cited paragraph
Wrong versionRight section, today's text, old eventIndex has no in-force datesCitation carries a date; version window must contain it
Summary cited as the courtA headnote quoted as if a judge wrote itEditorial summary stored in the same text as the reasonsTag document zones; only reasons are citable
Dissent cited as the holdingReal words, from the losing sideThe heading saying whose reasons these are was cut offCarry the section path into every chunk
Tidied quotationQuote marks around a paraphraseThe model smooths the wordingExact match after whitespace normalising
Page read as paragraph"At 14" means page 14 in one place, para 14 in anotherNumbers taken from the PDF layoutParse paragraph numbers from the text, never the page
Never retrievedA plausible pinpoint to a case the system never sawModel citing from memoryEvery cited ID must be in this turn's retrieved set

Only two rows involve the model: a tidied quotation and a citation from memory. The other five are faithful answers to text that was cut, labelled or stored badly. That changes where you spend the week.

Where the neighbour-paragraph bug comes from

The first row is the one worth drawing, because it's the most natural thing in the world to build by accident. You chunk a long decision by token count, say a few hundred tokens with some overlap. Paragraphs don't line up with token windows, so a chunk often starts partway through one paragraph and ends partway through the next. You need a paragraph number to cite, so you take the number of the paragraph the chunk starts in.

Now the passage that answers the question sits in the second half of that chunk, in the next paragraph. The model quotes it correctly. The citation says the wrong number. Every individual step was reasonable.

13 14 cited: 13 quote one chunk, two paragraphs
13 14 cited: 14 one chunk, one paragraph
Left, a token window cuts across a paragraph break and inherits the wrong number. Right, chunks follow the paragraphs, and the number comes with the text.

The fix is to chunk on the document's own structure. Decisions are numbered by paragraph, statutes by section and subsection. A chunk is one numbered unit, or a few whole units if they're short, and it carries the exact numbers it contains as metadata. A very long paragraph gets split, and both halves keep the same number. The chunker never decides a paragraph number; it copies one.

Checks that run before the answer ships

The checks in the right-hand column of the table are cheap. None needs a model. They run after generation and before the answer reaches the screen:

  1. Every citation must name a passage ID from this turn's retrieved set. Anything else is dropped.
  2. Every quotation must appear, after normalising whitespace and quote marks, inside the passage it cites.
  3. The paragraph or section number shown to the reader is rendered from that passage's metadata, not copied from the model's text.
  4. If the question has a date, the cited version's in-force window must contain it.

When a check fails, the honest options are to drop the citation and say so, or to retry with the failure named. What you shouldn't do is show it and hope.

CarefulNever let the model type a paragraph number. Give it passage IDs, make it cite those, and turn the IDs into pinpoints yourself. A model asked to write "para 14" will write a number that looks right, and looking right is the problem.

Start with one decision and a highlighter

Pick one long decision from your corpus, the longer the better. Print the chunks your pipeline made from it, next to the paragraph numbers in the original. Mark every chunk that crosses a paragraph boundary, and every chunk that mixes a headnote with the reasons, or a dissent with the majority. If you find even one, you have the neighbour-paragraph bug waiting in production, and the fix is in the chunker, not the prompt.

related

Keep reading