Writing

Why the service speaks English to everyone

In short: the supervised care companion speaks English to every participant because every line it says in a hard moment has to be signed off by a clinician, and a translation is a new text that nobody…

published
read time
5 min
words
996
lang
en
filed under
Product

In short: the supervised care companion speaks English to every participant because every line it says in a hard moment has to be signed off by a clinician, and a translation is a new text that nobody has signed. I wrote that down as an architecture decision, with what it would take to change it, so the question gets answered once instead of in every demo.

The question that comes up in every demo

I live in Montreal. Within five minutes of showing anyone the companion, somebody asks whether it speaks French. It is a fair question, and I have some sympathy for it. I grew up with Azerbaijani and Persian, I work in English, and my French is still at the stage where I apologise before I start.

The short answer is that the model could. A large language model will happily answer in French if a user talks to it in French. That was never the hard part. The hard part is everything around the model that a person in a bad moment depends on, and all of that is written in one language.

The service is the engine behind a conversational companion for people who need extra care, used under supervision in a care programme. Decisions that shape more than one surface go in docs/adr/ as architecture decision records. Localisation was one of the first to get that treatment, because it kept coming back and the answer kept getting improvised.

Where the language actually lives

When people picture localisation they picture the reply. In this service the reply is the least English part of the system. Here is where the language is actually fixed:

  • The persona. The system prompt is a file in the repository, read at startup. There is no fallback prompt. If the file is missing a required section, the service refuses to start.
  • The crisis scripts. When a user says something at the most serious tier, the model is not called at all. The user hears a written script, verbatim, one line at a time.
  • The tier rules. Deciding which tier a message belongs to is a rule table over the user's own words, and every rule quotes the approved sentence it was derived from.
  • The forbidden phrases. An output guard reads the reply as it is produced and cuts it if it matches a list of things the companion never says.
  • The evaluation cases. Sentences a participant might say, including ordinary sentences that happen to contain trigger words, used to measure whether the rules miss or over-fire.
  • The emails to the care team. Templates, with the same approval rules as the safety content.
user tier rules model output guard voice crisis script signed signed signed any language
The model could switch languages tomorrow. The shaded boxes could not.

Every one of those approved files carries four lines of frontmatter: a status, the name of the person who approved it, a date, and a hash of the file as it was when they read it. If anyone changes a single line afterwards, the hash no longer matches and the file counts as an unreviewed draft everywhere. Production refuses to start and names the file.

That machinery is the reason the answer is no. A French crisis script is not a translation of an approved file. It is a new file, and nobody has signed it.

The decision, written like one

An architecture decision record has a boring shape on purpose: the context, the decision, the consequences, and what would change it. Written that way, localisation stops being a feature request and becomes a trade with two honest columns.

One language, signed

  • Every line a user hears in a crisis was read by a clinician
  • The tier rules match the words they were written for
  • The evaluation numbers describe the system that ships
  • A person who speaks another language at home still gets English

Translate at the edges

  • The user hears their own language
  • The crisis script is now text nobody approved
  • The rules see translated words, or miss the untranslated ones
  • The evaluation numbers describe a different system

The left column costs something real. Some participants would be better served in the language they think in. I wrote that consequence into the record rather than leaving it out, because a decision that only lists its benefits is a sales pitch.

What it would take to speak a second language

The useful half of the record is the last section: what would make me reverse it. For one more language, roughly this:

  1. A clinician who reads that language reviews and signs the persona, the crisis scripts, the disclosure flow and the escalation wording in that language. Not a translator. A clinician.
  2. The tier rules are rewritten for that language and traced back, rule by rule, to approved sentences in that language.
  3. The forbidden phrase lists are rewritten, not translated. The phrases a companion must never say do not map one to one across languages.
  4. The evaluation cases are written natively, including the everyday sentences that carry trigger words, so the false alarm rate is measured in the language participants actually use.
  5. Speech recognition and speech synthesis are checked in that language for the voices of the people who will use it, because the voice path is a microphone first.

Seen as a list, a second language is a second launch. It is not a setting.

TipGive every decision record a section called "what would change this". It turns "no" into "not until these five things are true", which is an answer people can plan around.

Count your own strings

If you are building something that talks to people in vulnerable moments, try this before anyone asks you about languages. Search your repository for every piece of text a person in distress might hear or read: scripts, refusals, fallback lines, error messages, emails. Count the files. Then count the rules and tests that depend on the exact words in them.

That count is your localisation estimate. The number of languages your model supports is not.

related

Keep reading