Real-time transcription on a laptop: what Whisper gets wrong
In short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it:…
- published
- read time
- 4 min
- words
- 839
- lang
- en
- filed under
- Engineering
In short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it: gate the silence, overlap the chunks, stop feeding it its own past text, use a headset, and tell it the words to expect.
Earlier this year I put a small tool on GitHub called Transcriber. It turns an audio or video file into a text file with Whisper. On a clean recording it's very good, good enough for notes and first drafts.
Then I tried to make it live. Speak into the laptop, see the text appear a few seconds later, no cloud. That is a different problem. The model is the same. Almost everything around it changes, and the mistakes change with it.
Why live is harder than files
With a file, Whisper sees the whole recording. It works through it in windows of about thirty seconds and has plenty of context on both sides of every word.
Live, you don't have that. To keep the delay down you feed it short chunks of a few seconds as they arrive. Each chunk has little context. Some chunks are pure silence. Some start or end in the middle of a word. On a laptop CPU, every pass also has to finish before the next chunk is ready, or you fall behind.
What it gets wrong
These are the failure patterns I ran into, and the usual ones anyone building live transcription on Whisper will meet. The fixes are the ones that helped.
| Condition | What goes wrong | What helped |
|---|---|---|
| Silence or quiet room | Fluent text that nobody said, often a stock phrase | Drop quiet chunks before the model |
| Steady noise: fan, café | Dropped words and low-confidence guesses | Mic closer to the mouth, cut quiet chunks |
| Echo from laptop speakers | The other side of a call transcribed as you | A headset, or capture the mic only |
| Accents | Names and technical terms replaced by common words | A short glossary in the initial prompt |
| Short chunks | Wrong language picked for a chunk | Set the language, don't detect it |
| Chunk boundaries | Words cut in half or dropped | Overlap chunks by about a second |
| Feeding back past text | The same line repeated again and again | Turn off conditioning on previous text |
The accent row matters to me personally. My first languages are Azerbaijani and Persian, and my English carries that. Whisper handles accented speech far better than older systems I've used. Where it still slips is on words it hasn't seen much: names, project terms, library names. A few of those in the initial prompt goes a long way.
The loop that runs on a laptop
Here is the shape of a live loop with the fixes in place. It uses sounddevice to read the microphone in a callback, so no audio is lost while the model is busy, and the open source whisper package for the model.
import queue
import numpy as np
import sounddevice as sd
import whisper
SR = 16_000
CHUNK = 5 * SR # seconds of new audio per pass
OVERLAP = 1 * SR # carried over so edge words aren't cut
model = whisper.load_model("base.en")
q = queue.Queue()
def on_audio(indata, frames, time, status):
q.put(indata[:, 0].copy())
def loud_enough(x, rms=0.01):
return np.sqrt(np.mean(x ** 2)) > rms
buf = np.zeros(0, dtype=np.float32)
tail = np.zeros(0, dtype=np.float32)
with sd.InputStream(samplerate=SR, channels=1, dtype="float32", callback=on_audio):
while True:
buf = np.concatenate([buf, q.get()])
if len(buf) < CHUNK:
continue
chunk, buf = buf[:CHUNK], buf[CHUNK:]
audio, tail = np.concatenate([tail, chunk]), chunk[-OVERLAP:]
if not loud_enough(chunk):
continue # silence never reaches the model
out = model.transcribe(
audio,
language="en", # don't guess per chunk
fp16=False, # running on CPU
condition_on_previous_text=False, # stops repeat loops
initial_prompt="Glossary: PyTorch, Whisper, Montreal.",
)
print(out["text"].strip())
A few notes on it. The loudness gate is crude. A proper voice activity detector is better at telling speech from a loud fan, and it's the first thing I'd swap in. The overlap means the seam between chunks can produce the same word twice, so compare the start of each new line with the end of the last one and drop the repeat. And the model size is a trade. The small English-only models keep up on a laptop CPU. The larger ones are more accurate and fall behind.
Before you build your own
Record five minutes of your real setup: your room, your microphone, your voice, a stretch of silence and a stretch with someone talking on speaker. Run it through Whisper as a file first. Every mistake you see there will be worse live, and now you know which rows of the table above you need to fix first.
related