July 28, 2026/5 min readEmotion detection that must never slow the voiceIn short: the emotion classifier in the supervised care companion runs beside the reply, never in front of it. It starts when the user's words arrive, nothing waits for it, and when it finishes it…
July 14, 2026/5 min readBarge-in is the whole productIn short: a voice companion you cannot interrupt is a recording with extra steps. Barge-in, the user talking over the reply and the reply stopping, is what makes it feel like a conversation, and it…
November 4, 2025/5 min readStreaming speech: where the latency actually goesIn short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages…
October 21, 2025/4 min readA voice assistant in a weekend: LLM to speech over WebSocketsIn short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's…
September 9, 2025/5 min readSpeech to facial animation: a week with Audio2FaceIn short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a…
April 8, 2025/4 min readMonitoring a model that listens to patientsIn short: for a voice model in healthcare, the true labels arrive late or never, so you can't watch accuracy. You watch the recordings coming in, the scores going out, and the plumbing in between, and…
March 25, 2025/5 min readCloud Run or EC2: how I chose for a model that listensIn short: if the service takes a finished recording, scores it and answers, Cloud Run is usually the simpler and cheaper home. If it holds a live audio stream open, needs a GPU, or has to stay warm…
March 11, 2025/5 min readAudio preprocessing that matters more than the modelIn short: on medical audio, the steps before the model decide most of what the model can learn. Resampling, trimming, filtering, fixed-length windows and a split by patient each remove a shortcut the…
December 17, 2024/4 min readRespiratory sound classification, end to endIn short: I took a public set of lung sound recordings all the way from raw audio to a web demo anyone can upload a file to. The model was the smallest part. Most of the work was deciding what to…
November 5, 2024/4 min readReal-time transcription on a laptop: what Whisper gets wrongIn short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it:…