Streaming speech: where the latency actually goes
In short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages…
- published
- read time
- 5 min
- words
- 900
- lang
- en
- filed under
- Engineering
In short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages overlap, so the first sound waits only for the first sentence.
A few weeks ago I wrote about a voice assistant I built in a weekend: text in, a language model, text-to-speech, audio out over a WebSocket. It works, and for long answers it feels slow. This post is about why, and about how to see where the time goes in your own pipeline instead of guessing.
Two clocks, and only one of them matters
Every voice reply has two times:
- Time to first sound: from the moment the user stops to the moment they hear something.
- Total time: until the last word has played.
People forgive a long answer. They don't forgive silence. If the assistant starts talking quickly, a reply that takes a while to finish feels like conversation. If it is silent for the same total time and then plays everything at once, it feels broken. So the number to chase is the first one, and the first one is all about what each stage waits for.
The waterfall
In the buffered version, text-to-speech can't start until the model has finished the whole answer, the audio can't be sent until the whole file exists, and playback can't start until the whole message has arrived. Every stage adds its full length to the silence.
In the streamed version, the model's tokens are cut into sentences as they arrive. The first sentence goes to text-to-speech while the model is still writing the second. The first chunk of audio starts playing while the rest is still being made. The total time barely changes. The silence shrinks to roughly one sentence of writing plus one sentence of speech.
What each stage waits for
| Stage | Buffered: waits for | Streamed: waits for | Watch out for |
|---|---|---|---|
| Language model | The full answer | The first sentence | A long first sentence |
| Text to speech | The full text | One sentence | Odd pauses between sentences |
| Transport | The full file | One audio chunk | Base64 inside JSON |
| Playback | The full message | The first chunk | Formats that can't play in pieces |
One trap is worth calling out. In one project, the WebSocket handler for speech looked like streaming. It generated the speech, then sent it in small chunks, each with an index and an isLast flag. The client received many messages, which felt like streaming. But the file was finished before the first chunk left the server. Chunking a finished file is not streaming. The first sound still waits for the whole thing.
Measure it, don't guess
The fix starts with seeing it. I stamp each request with a few marks from the moment it arrives, and log them together. The sentence splitter sits between the model's token stream and the speech calls.
import re
import time
SENTENCE_END = re.compile(r"[.!?]\s")
class Marks:
def __init__(self):
self.t0 = time.perf_counter()
self.ms = {}
def mark(self, name):
# keep only the first time each mark is hit
self.ms.setdefault(name, round((time.perf_counter() - self.t0) * 1000))
async def sentences(stream, marks, min_chars=20):
buf = ""
async for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta.content or ""
if delta:
marks.mark("first_token")
buf += delta
while (m := SENTENCE_END.search(buf, min_chars)):
sentence, buf = buf[: m.end()].strip(), buf[m.end():]
marks.mark("first_sentence")
yield sentence
if buf.strip():
yield buf.strip()
async def reply(ws, client, messages, tts):
marks = Marks()
stream = await client.chat.completions.create(
model="gpt-4o-mini", messages=messages, stream=True
)
async for sentence in sentences(stream, marks):
audio = await tts(sentence)
marks.mark("first_audio_ready")
await ws.send_bytes(audio)
marks.mark("first_audio_sent")
marks.mark("done")
return marks.ms
A few notes on this:
min_charsstops the splitter from cutting at "Dr." or "e.g." at the very start of a reply, and from sending a one-word sentence that sounds clipped.- This version still runs speech for sentence two only after sentence one is sent. A small queue, with one task writing sentences and another speaking them, overlaps those too.
- The server can't hear the speaker. Add one more mark on the client, when playback actually starts, and send it back. That is the real first sound.
What streaming costs you
It isn't free. Speaking sentence by sentence can lose the flow between sentences, so the voice sometimes resets its tone at each full stop. Errors get harder: if speech fails on sentence three, the user has already heard one and two. And the client has to handle audio arriving in pieces, in order, and know when the reply is over. All of that is worth it for a voice product. For a tool that just saves an mp3 to disk, it isn't.
Do this next
Add five marks to your pipeline: request received, first token, first sentence, first audio sent, and first sound on the client. Log them for every request for a day. Then find the first stage whose mark grows with the length of the answer. That stage is waiting for "complete", and it's the one to change first.
related