Grounded answers or honest refusals: a rule for retrieval over personal files
In short: when a system answers questions over someone's own files, there is one rule I would not trade for anything: every answer quotes the person's own sources verbatim, or it is an honest refusal.…
- published
- read time
- 5 min
- words
- 999
- lang
- en
- filed under
- Engineering
In short: when a system answers questions over someone's own files, there is one rule I would not trade for anything: every answer quotes the person's own sources verbatim, or it is an honest refusal. Here is how I put together a small local system around that rule, and why.
What it is for
Most of what I know about my own life is scattered: PDFs in a downloads folder, years of photos, voice memos I never listened to again, chat exports. I wanted one place that reads all of it, links it into something navigable, and answers questions about it without sending any of it to anyone.
Two requirements shaped everything else. It has to run on a laptop, and it is not allowed to invent things. A second brain that confidently remembers a dinner you never had is worse than no second brain at all.
The architecture
Everything runs locally in containers, except the models, which run natively on the host so they can use the machine's GPU.
- Postgres with pgvector is the source of truth. Your content, its vectors and its metadata live there. If it is not in Postgres, it does not exist.
- The graph is a view of the source store. It is what makes your files navigable as a web of linked things rather than a list. It is also rebuildable from Postgres at any time, which means it can never be the only copy of anything, and a bad graph is a rebuild rather than a data loss.
- Models run through a local model runner on the host. Or, if you prefer, your own key for a hosted model, stored encrypted at rest. Embeddings run locally either way.
- Every model call is traced locally, so I can see exactly what was asked and what came back, and the trace never leaves the machine.
The reason for one source of truth and one projection is backup and sanity. There is one thing to back up and restore. Everything else is derived.
Models by role, not one model
The system does not use one global model. Work is routed to a small number of model roles, and each role can be pointed at a different model.
| Role | Used for | Default | If you have memory to spare |
|---|---|---|---|
| Quick steps | Light extraction and classification during ingest | One small multimodal model | A larger text model, or a hosted key |
| Answering | The main model that answers your questions | The same small model | A larger text model, or a hosted key |
| Vision | Understanding photos during ingest | The same small model | A small vision specialist |
| Embedding | Turning text into vectors for retrieval | A small embedding model | Nothing. It is frozen |
The default covers three roles with one model because a small general model that is also multimodal is good enough for all three. One download covers quick steps, answering and vision on a normal laptop, and the embedder is a small addition on top.
# Models run on the host, not in a container, so they get the GPU.
roles:
quick: small-multimodal # one model covers three roles
answer: small-multimodal
vision: small-multimodal
embed: small-embedder # frozen once anything is ingested
Splitting roles has one trap. In most model families the vision variant is a separate model, not a mode flag, so pointing the vision role at the plain text model of the same family does not work. And the smallest vision models in some families are heavy next to everything else in the stack.
On a weak machine, a small text-only model is a legitimate choice. Photos then degrade to a text-only understanding from the filename, the metadata and any text available, instead of failing outright. Degrading is better than failing.
Refusing to make things up
The contract with the user is short. An answer is grounded only in your own files and comes with verbatim citations, or the system says it does not know.
Verbatim matters. A paraphrased citation can drift from the source and still look cited, and you would have to open the file to notice. A quoted span is something you can check against the file directly. If the system cannot find words in your files that support an answer, it should not have an answer.
The refusal matters just as much. With a small local model, the temptation is to let it fill gaps from what it already knows about the world. For a second brain that is exactly the wrong behaviour. "I don't have anything about that in your files" is a correct answer. A plausible guess about your own life is not.
What you need to run something like this
- A laptop with a capable GPU for the default local setup. On a CPU-only machine a single photo caption takes minutes instead of seconds, which makes local vision too slow to be practical. There, use a hosted key for vision, or accept text-only photos.
- Memory. The containers need a comfortable amount of memory, the tracing stack adds a little more, and the local model needs its own share on top. A mid-range laptop's worth of memory is a realistic floor for running everything at once.
- Patience on first import. Every photo pays a vision call. That is a post of its own.
If you are building anything that answers questions over personal data, write down the refusal sentence before you write the prompt. Then try to make it say something your files do not contain. If you can, fix that before you add a single new feature.
related