Writing

Barge-in is the whole product

In short: a voice companion you cannot interrupt is a recording with extra steps. Barge-in, the user talking over the reply and the reply stopping, is what makes it feel like a conversation, and it…

published
read time
5 min
words
1,051
lang
en
filed under
Engineering

In short: a voice companion you cannot interrupt is a recording with extra steps. Barge-in, the user talking over the reply and the reply stopping, is what makes it feel like a conversation, and it depends on three things: a voice activity detector, a transport that streams both ways at once, and a latency budget you actually keep.

Why I built the boring version first

Before adding memory, retrieval or anything clever to the supervised care companion, a voice character for adults in a supported care programme, I built a small test rig whose only job was to feel right when you talk to it. No retrieval, no memory, no personalisation, on purpose. The safety guardrails still ship in code and fail closed, because a rig that talks to vulnerable people is still a thing that talks to vulnerable people.

The reason is simple. Real people do not wait politely for a turn to end. They start answering halfway through the question. They change their mind mid-sentence. If the character keeps talking over them, the illusion breaks in a way no amount of good writing fixes. You can have the best persona in the world and lose the participant in the first ten seconds because the voice did not stop.

The pipeline

The rig has two halves. The browser side is a web app with the character on screen, talking to the bot over WebRTC. The bot side is Python, running a real-time audio framework over a WebRTC transport, with an open-source voice activity detector and a speech-to-speech model on the other end. The browser opens the call with one HTTP request to an offer endpoint on the server, and from then on audio flows both ways continuously.

mic VAD speech to speech speaker audio out WebRTC barge-in: flush
The interesting arrow is the short one in the middle: hearing the user is what stops the voice.

The detector is the piece that matters most. It decides, many times a second, whether the incoming audio is a person speaking or just a room. When it flips to speaking while the character is talking, three things have to happen, in this order: stop the audio that is playing, throw away the audio that is queued, and tell the model its reply was cut so it does not carry on generating words nobody will hear.

What barge-in looks like in code

Most real-time voice frameworks handle interruptions for you, but it helps to see the shape of it without the framework. This is the logic in plain asyncio:

import asyncio


class Speaker:
    """Plays a reply and can be cut off mid-sentence."""

    def __init__(self, play_chunk):
        self.play_chunk = play_chunk          # async: writes audio to the transport
        self.queue: asyncio.Queue[bytes] = asyncio.Queue()
        self.task = None

    def start(self):
        self.task = asyncio.create_task(self._run())

    async def _run(self):
        while True:
            chunk = await self.queue.get()
            await self.play_chunk(chunk)

    def barge_in(self):
        # The user started talking: drop what is queued, stop what is playing.
        while not self.queue.empty():
            self.queue.get_nowait()
        if self.task:
            self.task.cancel()
        self.start()


async def on_vad(event, speaker, model):
    if event == "speech_started":
        speaker.barge_in()
        await model.cancel_response()         # stop generating what nobody will hear

The order matters. If you cancel the model first and flush the queue second, the user hears the tail of a sentence that has already been abandoned. Flush first. The ear notices.

The latency budget

Barge-in is half of it. The other half is how fast the character answers once the user does stop. In the main service the voice path can run two ways behind one switch: a text path (speech recognition, the same text path that typed chat uses, then sentence-by-sentence speech synthesis) or a speech-to-speech model on one connection. The budget looks different in each.

StageText pathSpeech to speech
Hearing the user stopThe VAD, on the botThe VAD, on the bot
Speech to textA provider round tripInside the model
ThinkingThe text path, same as typingInside the model
Text to speechSentence by sentence, a provider round tripNative audio
Measured so farUnder a second from last word to first audio, without the two provider round tripsRe-taken by a bench script, not quoted here

That figure comes from a real measurement against a real model, but it leaves out speech recognition and speech synthesis, because the machine that took it had no credentials for those providers. I would rather name the gap than guess at it. A bench script runs both engines and records latency percentiles and cost per minute, so the comparison can be re-taken when pricing or latency moves, which it will.

What speech to speech costs you

The speech-to-speech model is faster to feel, and it has a real problem for a product used in care. It speaks as it generates. On the text path the reply exists as moderated text before any audio exists, so a forbidden phrase is never sound. On the speech-to-speech path the audio is already playing while the transcript arrives, so the best a guard can do is stop the sentence part way and say an approved line instead. That is detection, not prevention.

So the default for conversation is the text path, and the faster speech-to-speech path is used only for scripted, pre-approved lines, such as the opening of a check-in, where latency is most of what a participant feels. The test rig is honest about this too: it has no streaming output moderation, and a production build for the care programme would need one.

Where it can and cannot run

  • WebRTC needs a long-lived process. The bot cannot run on a request-scoped serverless runtime. It runs as a long-lived container behind HTTPS, separate from the web front end.
  • The microphone needs HTTPS or localhost. Locally you are fine. In production both halves must be HTTPS.
  • Lock the origins. The offer endpoint should only accept calls from your own web domain before anyone real uses it.
  • Keep the engine swappable. The model is one block in the bot. Swapping it leaves the transport, the detector and the character alone.

If you are building anything voice-first, build the rig before the product. Then test it the way a real user would: start talking while it is mid-sentence, stop, start again, say "wait, no". Time from your last word to its first. If either of those feels wrong, nothing you add later will save it.

related

Keep reading