Skip to content

Talking Head

Takes a script plus a presenter image and returns a clip of that presenter delivering it. The audio drives the mouth, so audio quality caps video quality — a muddy voice track produces mushy mouth shapes no parameter fixes.

Presenter image. Front-facing, mouth closed or barely open, even lighting, no hand near the jaw. Three-quarter profiles work but drift more across long takes. A mouth-open source is the single most common cause of a “chewing” look.

Script or audio. Supply audio directly when you have it. If you supply text, the app runs Text to Speech first — same result as doing it yourself, but you cannot audition the voice before the video renders.

  1. Audition the voice in Text to Speech first. Cheaper than discovering the delivery is wrong after a video render.
  2. Upload the presenter image. Check the crop — output framing follows the source, so a shot cropped at the chin stays cropped at the chin.
  3. Generate ~10 seconds before committing the full script. Identity and mouth behaviour are visible immediately.
  4. Run the full script once the short take looks right.

Mouth moves but does not match words. Audio has music or noise under the voice. Isolate the voice track first.

Face changes partway through. Clip too long — split it. See the caution above.

Presenter looks stiff. Expected. The model animates speech, not gesture. Cut between angles rather than asking for more motion.

Teeth or lips smear on fast speech. Slow the delivery in the audio; the video follows what it is given.