Emotion detection that must never slow the voice
In short: the emotion classifier in the supervised care companion runs beside the reply, never in front of it. It starts when the user's words arrive, nothing waits for it, and when it finishes it…
- published
- read time
- 5 min
- words
- 999
- lang
- en
- filed under
- Engineering
In short: the emotion classifier in the supervised care companion runs beside the reply, never in front of it. It starts when the user's words arrive, nothing waits for it, and when it finishes it shifts a soft tint on the character. If it fails, the user never knows.
The tempting mistake
When you build a companion for adults in a supported care programme, it is natural to want it to read the room. If a participant sounds upset, the reply should be gentler. The obvious way to do that is to classify the person's mood first and then hand the result to the model that writes the reply.
That puts a second model call between a user finishing a sentence and the character starting to speak. In a voice product, that gap is the product. I spent a whole test rig getting the voice to answer quickly and stop when interrupted. I was not going to give that back for a mood score.
So the rule became: the emotion signal is off the hot path. It may never add a millisecond to the reply. It is allowed to be late, and it is allowed to be missing.
Two paths, one transcript
In the voice rig, a bot-side tool classifies each transcript into valence, arousal, intensity and a distress flag, using a small, fast model. The result is streamed to the browser as a server message, and the browser shifts a soft tint on the character, a quiet status indicator rather than a face pulling an expression. The voice never waits for it.
The main service does the same thing for every message a participant sends. The classification starts when the request arrives, runs beside the reply pipeline, and is written by a background task that starts only after the reply has been streamed and stored.
Fire and forget, done properly
"Fire and forget" is easy to get almost right in Python. The shape below is the one I trust:
import asyncio
import logging
log = logging.getLogger(__name__)
TIMEOUT_S = 20 # illustrative
_background: set[asyncio.Task] = set()
def spawn(coro):
# Keep a reference, or the task can be garbage collected mid-flight.
task = asyncio.create_task(coro)
_background.add(task)
task.add_done_callback(_background.discard)
return task
async def handle_turn(message, reply_stream, classify, store):
side = spawn(classify(message)) # starts now, runs beside the reply
async for chunk in reply_stream(message): # the hot path: nothing here awaits side
yield chunk
spawn(record(side, message, store)) # only after the reply is out
async def record(side, message, store):
try:
signal = await asyncio.wait_for(side, TIMEOUT_S)
except Exception:
log.info("no emotion signal for message %s", message.id)
return # no signal is the ordinary case
await store(message.id, signal)
Three details carry the weight. The classifier starts first, so it has the whole reply's duration to finish. The only await on it lives in a task created after the last chunk. And a timeout or an exception produces a message with no signal, which is not an error. Most messages have no signal anyway. A classifier that hangs holds one background task until its timeout and changes nothing else. The timeout in the sketch is illustrative.
A colour, not a label
The classifier produces a word, a valence, an arousal, an intensity and a distress flag. None of that leaves the store as a word. Inside the history store the label is converted into one of a small fixed set of display tints before the record is returned to anything. The only route that serves it is the one the character's tint is drawn from.
That is deliberate. An emotion label shown per message reads as a clinical claim about a vulnerable person, and I do not want any staff screen, or anything the person's family sees, making that claim. Colours are not feelings. If each tint had mapped onto one named feeling, it would just have been a coarser label with a different field name.
| Question | Voice test rig | Main service |
|---|---|---|
| What is classified | Each transcript | Each user message |
| Which model | A small, fast model | A separate emotion model, never the one that answers users |
| When it runs | Off the voice path | Starts on arrival, stored after the reply is out |
| What a screen gets | A colour on the character | One tint from a small fixed set, never a label |
| Distress flag | Tints the character only | Context for a clinician's review, never a trigger |
Two more rules that keep it honest
It never runs on the conversation model. The model that writes words a user reads or hears is for that alone. The classifier reads its own setting, and it refuses to run if that setting would resolve to the conversation model. The check compares against the model the current request asked for, not just the process default, because a request can carry its own model choice and that one wins.
A distress flag is supporting context, not a trigger. In the rig it only tints the character. That seam is where a production build would log, threshold and tell the named care contact. In the main service nothing raises a safety event from it. The clinician review screen shows it beside the turn, with the strength and the model that produced it, labelled as context and never as a trigger. It is an expressiveness signal. It is not an assessment.
The main service also ships with the whole thing switched off. Whether a stored inference about a person's emotional state should exist at all is a question waiting on a clinician, and drafted decisions about the people in the programme fail closed.
Check your own hot path
Open the handler for the moment your user finishes speaking or presses send. List every await between that moment and the first byte of the reply. For each one, ask whether the reply actually needs its result. Anything that does not can move to a side path, start early, and be allowed to fail quietly.
related