Speech to facial animation: a week with Audio2Face
In short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a…
- published
- read time
- 5 min
- words
- 1,019
- lang
- en
- filed under
- Engineering
In short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a pipeline, and the latency you feel comes mostly from the stages around it.
I've built voice companions before. In 2023 I worked on one for patients in a hospital, with a chat model behind it and a text-to-speech voice in front. The voice worked. The face never did. A static avatar that talks feels wrong in a way people notice in the first few seconds. So I spent a week of evenings with Audio2Face, as a user, to see what it would take to put a moving face on a voice.
What it actually gives you
Audio2Face-3D runs as a microservice. You send it audio, and it sends back facial animation as ARKit blendshapes: named weights like jawOpen, mouthSmileLeft or eyeBlinkRight, each between zero and one, one set per frame. It is not a renderer. Your engine takes those weights and moves the mesh.
Two details matter more than they look at first:
- Emotion is part of the output. Where it can detect emotion in the audio, the animation carries it. You can also pass the emotion in yourself, which is useful when the language model upstream already knows the tone of what it is saying.
- It talks gRPC. The repository ships the proto definitions, a docker compose quick-start and sample scripts. You will be writing a small client, not calling a REST endpoint.
Getting it running
The order I'd follow again:
- Try the hosted demo on NVIDIA's build site first. It costs nothing and tells you in five minutes whether the output style suits your character.
- Read the licensing before you plan anything. The GitHub resources are Apache 2. The microservice itself comes through an evaluation license of NVIDIA AI Enterprise on NGC, under NVIDIA's own product license.
- Start the quick-start docker compose. It brings up the service and a telemetry collector together.
- Run the sample scripts against the example audio in the repo, before your own audio. If something looks off, you want to know whether it is the service or your recording.
You need a GPU for the service. That shaped every later decision for me: it is a separate box to run, with its own cost, and it is not something you drop onto a small cloud container next to your API.
Where the latency goes
I didn't run a careful benchmark, so I won't give you milliseconds. What I can give you is the shape. The face is the last thing in a long chain, and most of the waiting happens before Audio2Face sees a single sample.
| Stage | Runs on | Waits for | Can stream? |
|---|---|---|---|
| Reply text | Language model API | The first tokens of the answer | Yes |
| Speech | Text-to-speech service | Enough text to say something | Yes, by sentence |
| Audio to blendshapes | GPU service | A short window of audio | Yes |
| Frames back to client | Network | Each batch of frames | Yes |
| Apply and render | Your engine | The next display frame | Per frame |
The one rule that took me a day to accept: hold the audio back until the face is ready. If you play the sound the moment it arrives, the lips trail behind it and the whole thing looks dubbed. It is better to start talking a little later and stay in sync than to start early and drift.
Cleaning up what comes back
The raw weights are good, but my rig did not use ARKit names, and a frame or two of jitter on the mouth is very visible on a close-up face. A small pass between the service and the engine fixed both.
# Map ARKit blendshape names onto the rig, clamp to [0, 1],
# and smooth the mouth so single-frame spikes don't show.
RIG_NAMES = {
"jawOpen": "Jaw_Open",
"mouthSmileLeft": "Smile_L",
"mouthSmileRight": "Smile_R",
"eyeBlinkLeft": "Blink_L",
"eyeBlinkRight": "Blink_R",
}
SMOOTHED = {"Jaw_Open", "Smile_L", "Smile_R"}
def to_rig(frames, alpha=0.6):
"""frames: list of (time_s, {arkit_name: weight})."""
out, prev = [], {}
for t, weights in frames:
cur = {}
for arkit, rig in RIG_NAMES.items():
w = min(max(weights.get(arkit, 0.0), 0.0), 1.0)
if rig in SMOOTHED and rig in prev:
w = alpha * w + (1 - alpha) * prev[rig]
cur[rig] = w
out.append((t, cur))
prev = cur
return out
Notice the blinks are left alone. Smoothing a blink turns it into a slow, sleepy droop. Smoothing also adds a little lag of its own, so keep alpha high and only smooth the shapes that need it.
What it is good for, and what it isn't
It is good for:
- Talking characters in games, training scenes and VR, where nobody has the budget to hand-animate every line.
- Voice assistants with a face. If the speech is generated, the animation has to be generated too.
- Prototypes. You can find out whether a face helps your product before you hire anyone to build one.
It is less of a fit when the character is a flat 2D cartoon, where a handful of mouth shapes picked from the phonemes looks fine and costs nothing. And it doesn't remove the hard part of a live avatar, which is the time between the user finishing a sentence and the character starting to answer. That time belongs to the language model and the voice.
Try this first
Take one of the example audio files, run it through the sample script, and plot jawOpen over time on top of the waveform. If the peaks line up with the loud syllables, the service is doing its job, and every problem you see after that is in your pipeline. That one plot will tell you more than a day of watching the avatar and guessing.
related