Writing

A voice assistant in a weekend: LLM to speech over WebSockets

In short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's…

published
read time
4 min
words
848
lang
en
filed under
Engineering

In short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's public, and its latency is honest about every shortcut I took.

I wanted a small, clean version of something I build over and over: text goes in, a character answers, and the answer comes back as speech. No frontend, no database, no accounts. Just the pipe. The code is public as EArts on my GitHub, and this post walks through how it's put together and where the time goes when you use it.

The contract: one message in, one message out

The client opens a WebSocket and sends JSON. Only two fields are required: type, which is always "text", and content, between 1 and 10,000 characters. You can also override the model, the voice (alloy, onyx or nova) and the temperature. The defaults are gpt-4, nova and 0.7.

The server answers with one JSON message: type: "audio", the format (mp3), the audio as base64, and the text the model wrote. If something fails, it answers with type: "error" and one of three codes: INVALID_REQUEST, OPENAI_API_ERROR or INTERNAL_ERROR. Three codes is enough for a client to decide whether to fix its input, retry, or give up.

The shape

client server (FastAPI) validate LLM TTS text mp3 + text, one message
Every step waits for the one before it to finish. That is the whole latency story of this version.

The project has four files that matter: main.py starts the FastAPI app, utils/websocket_handler.py holds the connection handler and the Pydantic models, utils/openai_service.py wraps the two API clients, and utils/config.py loads settings. Secrets live in .env and nothing else does. Models, defaults and the system prompt live in config/config.yaml, so changing the character's personality is an edit to a YAML file, not a code change.

The handler

Stripped down, the loop looks like this. Validation happens in the Pydantic model, so a bad request never reaches the paid APIs.

import base64
from typing import Literal

from fastapi import WebSocket, WebSocketDisconnect
from openai import AsyncOpenAI
from pydantic import BaseModel, Field, ValidationError

client = AsyncOpenAI()
SYSTEM_PROMPT = "You are a friendly character. Keep answers short."

class TextRequest(BaseModel):
    type: Literal["text"]
    content: str = Field(min_length=1, max_length=10_000)
    model: str = "gpt-4"
    voice: Literal["alloy", "onyx", "nova"] = "nova"
    temperature: float = Field(0.7, ge=0.0, le=2.0)

async def handle(ws: WebSocket):
    await ws.accept()
    try:
        while True:
            raw = await ws.receive_json()
            try:
                req = TextRequest(**raw)
            except ValidationError as e:
                await ws.send_json({"type": "error", "error": "INVALID_REQUEST", "message": str(e)})
                continue
            chat = await client.chat.completions.create(
                model=req.model,
                temperature=req.temperature,
                messages=[
                    {"role": "system", "content": SYSTEM_PROMPT},
                    {"role": "user", "content": req.content},
                ],
            )
            text = chat.choices[0].message.content
            speech = await client.audio.speech.create(
                model="tts-1", voice=req.voice, input=text, response_format="mp3"
            )
            await ws.send_json({
                "type": "audio",
                "format": "mp3",
                "data": base64.b64encode(speech.content).decode(),
                "text": text,
            })
    except WebSocketDisconnect:
        pass

The real version also catches API errors and maps them to OPENAI_API_ERROR, and anything else to INTERNAL_ERROR, so the socket stays open after a failure instead of dying with it.

Where the time goes

I didn't put a stopwatch on each stage for this post, so there are no milliseconds here. But you don't need them to see the problem. Look at what each stage waits for.

StageWaits forGrows withDelays first sound?
Receive and validateOne small JSON messageNothing that mattersBarely
LLM replyThe complete answerAnswer lengthYes
Text to speechThe complete answer textAnswer lengthYes
Encode and sendThe complete mp3Audio lengthYes
Decode and playThe complete messageAudio lengthYes

Every row except the first says "complete". The user hears nothing until the model has finished writing, the voice has finished speaking into a file, and that file has crossed the network in one piece. Base64 also makes the audio about a third bigger on the wire. A short answer feels fine. A long one feels like the assistant went to make coffee.

What I'd keep, and what I'd change

Keep:

  • The strict request model. Bad input fails fast, for free, with a clear error.
  • Secrets in .env, everything else in YAML. I can share a config without leaking a key.
  • The test client. test_client.py --save-audio writes the reply to output.mp3, so I can listen to exactly what a user would hear.

Change:

  • Stream the model's tokens and cut them at sentence boundaries.
  • Send each sentence to TTS as soon as it's complete, and play the first one while the rest are still being written.
  • Send audio as binary frames instead of base64 inside JSON.
  • Add conversation history. The request schema has no field for it, so every message starts from zero.

Run it yourself

Clone EArts, copy .env.example to .env and add your OpenAI key, then start the server with python main.py. In a second terminal run python test_client.py --text "Tell me a joke" --save-audio. Then ask for a long story and compare how long you wait for each. That difference is the next thing to fix, and it's what my next post is about.

related

Keep reading