Writing

Speech to facial animation: a week with Audio2Face

In short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a…

published
read time
5 min
words
1,019
lang
en
filed under
Engineering

In short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a pipeline, and the latency you feel comes mostly from the stages around it.

I've built voice companions before. In 2023 I worked on one for patients in a hospital, with a chat model behind it and a text-to-speech voice in front. The voice worked. The face never did. A static avatar that talks feels wrong in a way people notice in the first few seconds. So I spent a week of evenings with Audio2Face, as a user, to see what it would take to put a moving face on a voice.

What it actually gives you

Audio2Face-3D runs as a microservice. You send it audio, and it sends back facial animation as ARKit blendshapes: named weights like jawOpen, mouthSmileLeft or eyeBlinkRight, each between zero and one, one set per frame. It is not a renderer. Your engine takes those weights and moves the mesh.

Two details matter more than they look at first:

  • Emotion is part of the output. Where it can detect emotion in the audio, the animation carries it. You can also pass the emotion in yourself, which is useful when the language model upstream already knows the tone of what it is saying.
  • It talks gRPC. The repository ships the proto definitions, a docker compose quick-start and sample scripts. You will be writing a small client, not calling a REST endpoint.
neutral
jawOpen
mouthSmile
Three blendshapes pushed to the top of their range. Real speech is a blend of dozens of them, moving every frame.

Getting it running

The order I'd follow again:

  1. Try the hosted demo on NVIDIA's build site first. It costs nothing and tells you in five minutes whether the output style suits your character.
  2. Read the licensing before you plan anything. The GitHub resources are Apache 2. The microservice itself comes through an evaluation license of NVIDIA AI Enterprise on NGC, under NVIDIA's own product license.
  3. Start the quick-start docker compose. It brings up the service and a telemetry collector together.
  4. Run the sample scripts against the example audio in the repo, before your own audio. If something looks off, you want to know whether it is the service or your recording.

You need a GPU for the service. That shaped every later decision for me: it is a separate box to run, with its own cost, and it is not something you drop onto a small cloud container next to your API.

Where the latency goes

I didn't run a careful benchmark, so I won't give you milliseconds. What I can give you is the shape. The face is the last thing in a long chain, and most of the waiting happens before Audio2Face sees a single sample.

StageRuns onWaits forCan stream?
Reply textLanguage model APIThe first tokens of the answerYes
SpeechText-to-speech serviceEnough text to say somethingYes, by sentence
Audio to blendshapesGPU serviceA short window of audioYes
Frames back to clientNetworkEach batch of framesYes
Apply and renderYour engineThe next display framePer frame

The one rule that took me a day to accept: hold the audio back until the face is ready. If you play the sound the moment it arrives, the lips trail behind it and the whole thing looks dubbed. It is better to start talking a little later and stay in sync than to start early and drift.

Cleaning up what comes back

The raw weights are good, but my rig did not use ARKit names, and a frame or two of jitter on the mouth is very visible on a close-up face. A small pass between the service and the engine fixed both.

# Map ARKit blendshape names onto the rig, clamp to [0, 1],
# and smooth the mouth so single-frame spikes don't show.
RIG_NAMES = {
    "jawOpen": "Jaw_Open",
    "mouthSmileLeft": "Smile_L",
    "mouthSmileRight": "Smile_R",
    "eyeBlinkLeft": "Blink_L",
    "eyeBlinkRight": "Blink_R",
}
SMOOTHED = {"Jaw_Open", "Smile_L", "Smile_R"}

def to_rig(frames, alpha=0.6):
    """frames: list of (time_s, {arkit_name: weight})."""
    out, prev = [], {}
    for t, weights in frames:
        cur = {}
        for arkit, rig in RIG_NAMES.items():
            w = min(max(weights.get(arkit, 0.0), 0.0), 1.0)
            if rig in SMOOTHED and rig in prev:
                w = alpha * w + (1 - alpha) * prev[rig]
            cur[rig] = w
        out.append((t, cur))
        prev = cur
    return out

Notice the blinks are left alone. Smoothing a blink turns it into a slow, sleepy droop. Smoothing also adds a little lag of its own, so keep alpha high and only smooth the shapes that need it.

What it is good for, and what it isn't

It is good for:

  • Talking characters in games, training scenes and VR, where nobody has the budget to hand-animate every line.
  • Voice assistants with a face. If the speech is generated, the animation has to be generated too.
  • Prototypes. You can find out whether a face helps your product before you hire anyone to build one.

It is less of a fit when the character is a flat 2D cartoon, where a handful of mouth shapes picked from the phonemes looks fine and costs nothing. And it doesn't remove the hard part of a live avatar, which is the time between the user finishing a sentence and the character starting to answer. That time belongs to the language model and the voice.

TipPass the emotion in when you already know it. A model that just wrote a sad sentence is a better judge of its tone than a detector listening to the synthetic voice afterwards.

Try this first

Take one of the example audio files, run it through the sample script, and plot jawOpen over time on top of the waveform. If the peaks line up with the loud syllables, the service is doing its job, and every problem you see after that is in your pipeline. That one plot will tell you more than a day of watching the avatar and guessing.

related

Keep reading