How I evaluate retrieval before anyone asks a question
In short: before a single user types a question, I write the questions myself, mark which passage answers each one, and measure how often that passage comes back in the top k. That one number, recall…
- published
- read time
- 5 min
- words
- 1,039
- lang
- en
- filed under
- Engineering
In short: before a single user types a question, I write the questions myself, mark which passage answers each one, and measure how often that passage comes back in the top k. That one number, recall at k, tells me more about a RAG system than any amount of reading answers.
Why retrieval gets measured first
A RAG answer has two places to go wrong. The retriever can fail to find the passage, or the model can misuse a passage it was given. When you only read final answers, the two look the same: a wrong answer. And the second kind is much easier to talk yourself out of, because the model writes so well.
So I split them. Retrieval gets its own test, with no model in the loop, and it runs before any prompt work starts. In a retrieval project I'm building, this is a small script that runs a fixed question set against the index and prints a handful of numbers. It's the first thing I run after any change to chunking, embeddings or search.
If the right passage isn't in what the model sees, no prompt will save the answer. If it is, the remaining problem is the model's, and that's a different test.
Building the question set
The set is a file of questions, each with the passage or passages that answer it. I call those the gold passages. The work is in picking good questions, and there are a few kinds I always include:
- Plain-language questions. The way a client would ask it, with none of the statute's words. "Can my landlord keep my deposit for normal wear?"
- Lawyer questions. Using the defined terms, the section numbers and the citation formats a practitioner types.
- Questions with more than one gold passage. A rule and its exception, or a section and the definition it relies on. Getting one of the two is not getting the answer.
- Questions with no answer in the corpus. These don't count towards recall, but they show me what the system returns when it should say "nothing here".
I write the questions by starting from the passage and working backwards: open a section, ask what someone would need to know for this to be the answer, and phrase it the way they'd phrase it. That's slower than generating questions with a model, and it matters. A model writing questions from a passage tends to reuse the passage's own words, which makes every retriever look better than it is.
Each row is a line of JSON, so the file diffs cleanly and lives in the repository next to the code it tests.
{"id": "q-017", "question": "Can a landlord keep a deposit for normal wear and tear?", "gold": ["p-0412"], "kind": "plain"}
{"id": "q-018", "question": "s. 38(2) deposit deductions, ordinary wear", "gold": ["p-0412"], "kind": "lawyer"}
{"id": "q-019", "question": "Who counts as a tenant for the notice rules?", "gold": ["p-0007", "p-0530"], "kind": "multi"}
Recall at k, worked through
Recall at k asks one question per row: did a gold passage appear in the top k results? Average that over the set and you have the number. For questions with several gold passages, I count the share of them found, so finding one of two scores a half.
import json
def recall_at_k(results, gold, k):
top = set(results[:k])
return len(top & set(gold)) / len(gold)
def evaluate(search, path, ks=(1, 5, 10, 20)):
rows = [json.loads(line) for line in open(path)]
rows = [r for r in rows if r["gold"]] # unanswerable rows are scored separately
scores = {k: 0.0 for k in ks}
for r in rows:
results = search(r["question"], limit=max(ks))
for k in ks:
scores[k] += recall_at_k(results, r["gold"], k)
return {k: round(s / len(rows), 2) for k, s in scores.items()}
To make the arithmetic concrete, take a toy set of six questions, each with one gold passage. Say the gold passage came back at rank 1, 3, 7, 2 and 12, and for the sixth question it never came back at all. The table below is that illustrative example, not a result from my system.
| k | Questions with gold in top k | Recall at k |
|---|---|---|
| 1 | 1 | 0.17 |
| 5 | 3 | 0.50 |
| 10 | 4 | 0.67 |
| 20 | 5 | 0.83 |
Two things to read from a table like this. First, the k that matters is the one your pipeline actually uses. If you hand the model ten chunks, recall at 10 is your ceiling, and recall at 20 is only telling you what a reranker could rescue. Second, the gap between columns is information. A big jump from 5 to 20 means the passage is being found but ranked badly, which points at ranking. A flat line that never gets high means it isn't being found at all, which points at chunking or at the query not sharing vocabulary with the text.
Reading the misses
The average is the headline. The misses are where the work is. My script writes every row where the gold passage didn't make the top k, with what came back instead. Reading twenty of those is usually enough to see the pattern: a section split in the middle of a sentence, a definition stored apart from the rule that uses it, a citation-style query that the embeddings treat as noise.
I also break the number down by the kind field. One overall recall can hide a system that is great at lawyer questions and poor at plain ones, which is the opposite of what a product for non-lawyers needs.
Start yours with thirty questions
You don't need hundreds to begin. Thirty questions, written by hand from the passages, split across plain, expert and multi-passage kinds, will already tell you whether a change helped. Put the file in the repo, run the script on every change to retrieval, and write the numbers in the pull request. Add a question every time a real user finds a miss. After a few months the set describes your users better than any spec you could have written up front.
related