Writing

Flat image in, living character out

In short: give the pipeline a JPG or PNG of a character and it produces a layered rig plus a small web player that breathes, blinks, shows emotion and lip-syncs to a real voice. The steps are…

published
read time
5 min
words
901
lang
en
filed under
Engineering

In short: give the pipeline a JPG or PNG of a character and it produces a layered rig plus a small web player that breathes, blinks, shows emotion and lip-syncs to a real voice. The steps are background removal, rigging, generating the layers the drawing never had, and an audit that refuses to ship a rig that is wrong.

Why a pipeline and not a one-off

A voice companion for adults in a supported care programme needed an on-screen character that feels alive while it talks: a mouth that moves with its voice, eyes that blink, a body that shifts a little while it listens. The first character's rig came out of tooling written for that one character, and it worked. Then more characters arrived and needed the same treatment, and I had two choices: do it all again for each one, or generalise what I had into a pipeline.

I generalised it. The reference rig is the first character, with a full face, mouth shapes and a wave, and two more were built straight from new drawings. One of those came from a flat JPG whose face colour matches its own background, which is the hardest common case for background removal.

The repository has one rule I keep coming back to: every rule in the code was paid for with a shipped bug, and the comments say which one. Keep the comments.

Four stages

cutout + alpha
pivot layers + pivots
film strip
Extract, rig and verify. Generating the missing layers happens between the second and third.

Extract

Background removal sounds solved until the character's colours match the background. The extractor builds an alpha ramp from each pixel's distance to the background colour, keeps the largest connected component so stray specks do not become part of the character, fills in the interior so a pale face does not turn into a hole, and cleans up colour fringes only at the true outline, not across the whole character.

Rig

The rig is a manifest, and the manifest is the product. It lists the layers, the pivots they rotate around, the zones, the draw order, and the invariants the player relies on. The player is dumb on purpose. If the manifest is right, the character is right.

Generate

A flat drawing has no closed eyelids, no mouth shapes for speech, no eyes for different emotions and no keyframes for an arm. Those layers are completed with AI image generation. They are never stubbed with placeholders, because a placeholder that ships is a bug a real person sees.

Verify

A rig that fails the audit does not ship. The audit looks for leftover box edges from a bad cutout, transparency holes inside the body, and draw orders that do not make sense. It also captures a film strip, several frames in a row, because one screenshot cannot show a flicker.

Lip sync from the real voice

The mouth does not follow the text. It follows the actual amplitude of the voice coming out of the speaker. When the voice is loud, the mouth opens wider. When it is quiet, the mouth overlay fades and the painted smile from the original drawing shows through.

voice amplitude mouth quiet: painted smile
Loud opens the mouth, quiet lets the original drawing's smile show through.

In the browser this is the Web Audio API and a CSS variable. The important decision is that it is imperative. Writing amplitude into React state many times a second would re-render the character constantly, so it never touches React at all:

const ctx = new AudioContext();
const analyser = ctx.createAnalyser();
analyser.fftSize = 1024;
ctx.createMediaStreamSource(remoteVoice).connect(analyser);

const samples = new Uint8Array(analyser.fftSize);
const mouth = document.querySelector<HTMLElement>(".mouth")!;

function sample() {
  analyser.getByteTimeDomainData(samples);
  let sum = 0;
  for (const s of samples) {
    const x = (s - 128) / 128;
    sum += x * x;
  }
  const rms = Math.sqrt(sum / samples.length);
  // Straight to a CSS variable: no state, no re-render.
  mouth.style.setProperty("--open", Math.min(1, rms * 5).toFixed(3));
}

// An interval, not requestAnimationFrame: a throttled tab slows it down but never freezes it.
setInterval(sample, 33);

The comment on the last line is one of those rules paid for with a bug. A character that freezes mid-word when the tab loses focus looks broken in a way that is hard to explain to someone who just wanted to talk. The player's own clock runs on an interval for the same reason.

Two more rules from the player

  • Swap sprites, never crossfade. Changing from one mouth shape to another is a pure replacement. A crossfade between two layers shows both of them half transparent for a few frames, and on a face that reads as a ghost.
  • Colour washes follow the body. The player can lay a soft, neutral tint over the character. The wash is masked to the character's exact silhouette and tracks the base layer's transform, so when the body leans, the colour leans with it instead of staying behind as a rectangle.

The player is one JavaScript file with no dependencies. Running the pipeline takes one command-line call on an image file, and the output folder includes a preview page you can serve with python3 -m http.server.

If you are animating a character for anything interactive, try one thing before anything else: put a film strip capture in your test loop. Record eight or ten frames in a row while the character talks, and look at them side by side. A flicker that one screenshot hides is obvious there.

related

Keep reading