Writing

Fail closed: designing an AI for people who can't afford a bad answer

In short: when the user can't afford a bad answer, every part of the system that can fail has to fail towards the safe answer, not towards the model. In a supervised care product I'm building, a voice…

published
read time
5 min
words
1,018
lang
en
filed under
Engineering

In short: when the user can't afford a bad answer, every part of the system that can fail has to fail towards the safe answer, not towards the model. In a supervised care product I'm building, a voice and chat companion for adults in a supported care programme, that means a rule table decides the safety tier before the model is called, an error counts as the most serious tier, the model's words are checked before they're sent, and nothing a participant hears in a crisis was written by a model.

Fail open is the default, and it's wrong here

Most software fails open. If the spam filter crashes, the email goes through. If the recommendation service times out, you get a generic list. That's the right call when the cost of a miss is small and the cost of an outage is large.

A companion for people who need extra care, used under supervision, flips that. The cost of a miss is a person who said something frightening and got a cheerful improvised reply. The cost of an outage is a person who can't chat for a while. So the design rule I hold everything to is the opposite of the usual one: when in doubt, take the path a clinician wrote down in advance.

A generic assistant talking to a person in distress is a worse outcome than an outage.

That sentence decides small things too. If the persona file the service reads at startup is missing, malformed, or has lost a required section, the service doesn't start. There's no fallback prompt anywhere in the code, because a fallback prompt is a generic assistant.

Three tiers, decided before the model

Before the model is called, a rule check places every turn a participant takes in one of three tiers. The tiers, their triggers, the crisis scripts and the escalation policy are written by a clinician, not by me. My job is to make the code do exactly what those documents say, and nothing they don't.

user's turn tier decider rules, no model tier 1 approved script only tier 2 reply, the care team is told tier 3 reply, nothing leaves any error
The decider is code and phrase rules over the user's own words. A fault anywhere in it routes the turn to tier one.

Here's what each tier does, as the code does it today.

TierModel calledWhat the user hearsWhat is recordedWho is told
OneNoThe approved crisis script for that trigger, word for wordA safety event, before the first line is sentThe named care contact, before the first line is sent
TwoYesThe model's reply, with one approved line added, once per trigger per conversationA safety event in the day's review listA named member of the care team, the same day
ThreeYesThe model's replyNothingNobody

Three choices in that table are worth explaining.

  • Tier one never reaches the model. Not "the model with a careful prompt". The script is literal text, and the content checks refuse a script that is an instruction to improvise.
  • The care contact is told before the user hears anything. That's the order the approved escalation policy asks for, so the event is written and the notification sent first.
  • Tier two is said once. A participant who mentions feeling unwell on five turns produces one event and hears "I'm going to make sure someone on your care team knows" once. Repeating it every turn would teach them to stop mentioning it.

Stops that only go one way

Sessions follow a set order of steps: a check-in, a conversation, a technique to try. An order creates pressure to continue, which is exactly what you don't want after a person discloses something. So a tier one turn doesn't just change the reply. It ends the session's current step, and that step can't be resumed. The no-resume rule is enforced in more than one layer, from the code down to storage, because one place is a habit and several is a rule.

The stop is not a model decision. The function that decides whether it happens takes a few plain values and returns an answer. It isn't async, so there is nothing in it that could wait on a model.

A guard on the way out

The tier decider reads the user's words. A second, separate check reads the companion's. The persona's constraints file lists phrases the companion never says, marked as hard rules. As the model's reply streams in, an output guard matches it against that list. If one fires, the turn is cut, the generated text reaches neither the user nor the database, and the user hears a short approved line instead.

The guard decides no tier and notifies nobody. It only refuses to say things the persona already says it never says. Turning it off is possible for development, and a production deployment with it off refuses to start.

Content a clinician signs, and a check that notices edits

Safety content is signed off by a clinician, and the sign-off is tied to the exact text they read. If one word of a crisis script changed after sign-off, the content counts as unreviewed, and neither the build nor the live service will accept it.

Before that check existed, an edited line could sit in a signed file with a clinician's name still on it. Nothing would have noticed.

CarefulFail closed has a price, and it's paid in false alarms. A rule table can catch every serious sentence you test it with and still fire on "my sister lent me money for the rent". Test it with ordinary sentences built from the rules' own trigger words, and treat "no misses on my own cases" as exactly that, not as recall.

If you're building for a vulnerable audience

Write down, for each failure in your system, which way it fails. The model times out: what does the user hear? The safety check throws: which tier is that turn? The prompt file is missing: does the service start? If any answer is "the model improvises", change it so the answer is "a script someone qualified approved", and make the code refuse to run any other way.

related

Keep reading