Reading papers with an LLM without getting lied to
In short: give the model only the part of the paper that holds the study, ask narrow questions that demand quotes and an explicit NA, and check every quote against the text. It still breaks on numbers…
- published
- read time
- 4 min
- words
- 862
- lang
- en
- filed under
- Research
In short: give the model only the part of the paper that holds the study, ask narrow questions that demand quotes and an explicit NA, and check every quote against the text. It still breaks on numbers and on long lists of questions, so keep both small.
A systematic review means reading dozens of papers and filling the same spreadsheet for each one. The one I was working on had a lot of columns. Study type, number of participants, age range, sensors and their exact model names, sampling rates, which features were extracted, how the data was secured, what the authors said their limitations were. It was a review of wearable health systems, so every paper had a different mix of signals and devices.
ChatGPT can read a paper faster than I can. It can also fill a cell with something that sounds right and isn't in the paper at all. So I wrote down a method, put it in a public repo called LitReviewGPT, and adjusted the wording whenever it got something wrong.
Paste less of the paper
The first message gives the model the paper, from the Methods section to the end. Not the abstract, not the introduction.
The reason is simple. An introduction is a summary of other people's work, full of their participant counts and their accuracies. Ask "how many participants?" with the introduction included, and the model may answer with a number from a study the authors cited. Cutting the paper at Methods removes most of that temptation.
The prompts
After the paper goes in, every question follows the same template. Only the list of fields changes.
FIRST = """Read this article completely and answer my questions
in the next message:
{paper_from_methods_to_end}"""
ASK = """Read the article carefully and pay attention to every detail.
Tell me what the authors wrote about:
{fields}
I prefer the exact sentence and quotation from the paper.
I don't want you to write it yourself. Answer in bullet points.
When you mention equipment, devices or techniques, give exact
names and model names. For items the paper does not discuss,
write NA."""
BATCHES = [
"Type of study, Number of participants, Age range, Inclusion criteria",
"Sensors (exact names), Data types, Sampling frequency, Real-time?",
"Data encryption, Authentication, Any other security measure",
"Advantages, Limitations, Future work, Findings with numbers",
]
I pasted these by hand into the chat, one batch per message. The template above is a tidied version of the ones in the repo. The originals also told the model to "set your laziness to zero", which tells you something about how the early answers looked.
Where it broke
Almost every line in that prompt is there because of a specific way the answers went wrong. Here is what each kind of question tended to produce before the fix.
| Question | Failure mode | What helped |
|---|---|---|
| Number of participants | A number from a cited study, or recruited instead of analysed | Methods onward only; ask for the quote |
| Female to male ratio | A ratio the paper never states, worked out by the model | "Don't write it yourself" |
| Devices and sensors | "An EEG headset" instead of the model name | Ask for exact names and model names |
| Fields the paper skips | A plausible guess instead of a blank | An explicit NA rule |
| Long lists of fields | Later fields answered thinly or skipped | Batches of related fields, one per message |
| Findings with numbers | Values rounded or attached to the wrong result | Check every number against the tables |
The last row is the one that no wording fully fixes. Numbers are where a language model is weakest at reading, because a sentence with the wrong number in it reads just as well as the right one. I treated every number as unverified until I'd found it in the paper myself.
Check the quotes, not the answer
Asking for exact quotations has a second benefit. A quote is something you can search for. If the model says the authors wrote a sentence, that sentence is either in the PDF or it isn't.
import re
def norm(s):
return re.sub(r"\s+", " ", s).strip().lower()
def missing_quotes(answer_quotes, paper_text):
"""Return the quotes that do not appear in the paper."""
text = norm(paper_text)
return [q for q in answer_quotes if norm(q) not in text]
A quote that doesn't match is usually a paraphrase dressed up in quotation marks, and those are exactly the cells to read again by hand.
Try it on your next paper
You don't need my repo to use this. Take the next paper on your list and do four things:
- Paste from Methods to the end, nothing earlier.
- Group the fields into a few batches of related questions, one batch per message.
- Demand quotes and an explicit NA for anything not discussed.
- Search the PDF for every quote and every number before it goes in your spreadsheet.
The model saves you the reading. It doesn't save you the checking.
related