RAG with a persona: when the knowledge base and the character disagree
In short: when a character and its knowledge base disagree, let the knowledge base win on content and the character win on voice, and let safety rules beat both. You enforce that in how you build the…
- published
- read time
- 5 min
- words
- 946
- lang
- en
- filed under
- Engineering
In short: when a character and its knowledge base disagree, let the knowledge base win on content and the character win on voice, and let safety rules beat both. You enforce that in how you build the prompt, not by hoping the model sorts it out.
I have been working on the brain of a supervised companion for adults in a supported care programme: a chat character that acts as a calm guide who talks like a person, not a form, grounded in a knowledge base of strategies written by professionals. The two halves pull in different directions all the time. The knowledge base is careful, formal and long. The character is short and warm. Most of the engineering is deciding who gets the last word on what.
The setup
The service is a small graph with three steps behind an HTTP API:
- Retrieve embeds the user's message, searches a vector store, fetches chunk text and metadata from Postgres, keeps only close matches, and returns a few chunks with citations.
- Build the prompt puts the persona first, then the retrieved context, then safety and boundary rules, then the recent history, then the new message.
- Generate calls the model with a low temperature and a short token cap.
Documents go in through an ingest endpoint that splits them into fixed-size chunks with some overlap. Nothing exotic. The interesting part is all in the middle step.
Four ways the two disagree
Voice
The knowledge base is written by professionals for professionals. Give the model a chunk of it and the easiest thing for it to do is paraphrase that chunk closely. The answer comes out correct and sounds like a pamphlet. The character is gone for that turn.
What the character would know
A calm guide who talks like a person does not quote research. If the character suddenly sounds like an expert, it breaks the persona just as badly as getting the facts wrong. The content has to come through as something the character "learned" or "heard", in its own words.
Nothing retrieved
This is the dangerous one. When no chunk clears the similarity floor, the context is empty, but the model still answers. It answers in character, confidently, with advice that came from nowhere. The persona makes made-up advice sound friendlier, which makes it more convincing, not less.
Wanting to please
A warm character wants to agree. Some messages from a participant should not be agreed with. That is not a knowledge question at all. It belongs to the rules.
Who wins what
The split I use:
- Knowledge base wins on content. If the character gives advice, it comes from a retrieved chunk.
- Persona wins on voice. The chunk is raw material. The model must restate it, short and in character.
- Safety rules win over both. They go last in the system prompt, so they are the final instruction the model reads before the conversation.
Here is a generic sketch of the middle step that shows the pattern. The names are made up. The branch for an empty context is the part I would insist on in any system like this.
def build_prompt(turn, persona, rules, floor, top_k, history_len):
chunks = [c for c in turn.chunks if c.score >= floor][:top_k]
if chunks:
notes = "\n\n".join(f"[{i + 1}] {c.text}" for i, c in enumerate(chunks))
grounding = (
"Things you have learned. Use only these for advice. "
"Say them in your own words, short, the way you talk.\n\n" + notes
)
else:
grounding = (
"You have nothing on this topic. Do not give advice. "
"Say you are not sure, ask a question back, or suggest "
"talking to someone on their care team."
)
system = "\n\n".join([persona, grounding, rules])
recent = turn.history[-history_len:]
return [{"role": "system", "content": system},
*recent,
{"role": "user", "content": turn.message}]
The labels matter more than they look. "Things you have learned" lets the character own the content without quoting it. "Use only these for advice" is the grounding rule. The low temperature and the short token cap help too: a short answer has less room to drift from its source and less room to slip out of character.
What to check in your own system
If you are putting a persona on top of RAG, run these before anything else:
- Ask ten questions the knowledge base cannot answer. Read what comes back. Every confident answer is a bug.
- Ask ten it can answer. Check two things separately: is the content from the cited chunk, and does it still sound like the character?
- Look at the similarity floor. A loose threshold on cosine similarity lets weak matches through. Print the scores next to the chunks for a day and move it if the weak ones are noise.
- Put the safety rules last in the system prompt and keep them short. Rules buried in the middle of a long prompt are easier for the model to lose.
- Return the citations with every response, even if the user never sees them. Whoever reviews conversations later needs to know which chunk the character was standing on.
The character is what participants come back for. The knowledge base is why it is safe to let them. Write the prompt so neither one has to pretend to be the other.
related