Workspaces are coming: shared projects, assets and credits for your whole team

Follow along
Video

Make the mouth match the audio.

Fourteen lipsync models across four modes — video-to-video dubbing, avatars, image-to-video and reference-to-video.

Real output

A face, and something for it to say

The face comes from a picture or a clip. The words arrive as a recording you upload or as text the model speaks itself — which is what changes how much you have to supply.

What you give it
  1. A face A photograph on most models, a clip on the two that take one. This is who is speaking.
  2. The words Either an audio file you upload, or text the model speaks in a voice of its own. The model you pick decides which.
  3. Direction How it is delivered rather than what is said — held look, small nods, hands just inside the frame. On a model given audio, this is all the prompt carries.
What comes back
One take A single clip of that face saying those words, the mouth matched to the audio and the delivery following the direction.
How it works

A face, a track, and the direction to play it

01

Upload the source

A video for the video-to-video models, a portrait for the avatar and image-to-video models, or a reference clip. The source decides which of the four modes you are in, and the mode decides which models are available. 200 MB per slot.

02

Supply audio, or don't

Most of the fourteen take an audio track and drive the mouth from it. Six do not: Veo 3.1, Seedance v1.5, Kling v2.6 and all three Wan models generate what is spoken rather than following a track you give them.

03

Write the prompt

The prompt is the app's central control, required on seven of the fourteen and capped at 2000 characters. It is direction rather than dialogue — pace, where they look, what their hands do. On the models that speak it, it is both. A negative prompt of up to 600 characters names what to keep out.

04

Check the face, then render

Front-on, mouth visible, nothing crossing it. A hand near the jaw or hair over the mouth is the most common reason a take comes back wrong. Clip length is validated before the job is submitted, so a paid run does not fail at the provider, and failed runs are refunded.

Inputs and outputs

What Lipsync takes, and what it returns

What Lipsync needs from you

Source
A video, a portrait or a reference clip, depending on the mode. 200 MB per slot
The face
Front-on, mouth visible, nothing crossing it — a hand near the jaw or hair over the mouth is the top cause of a bad take
Audio
Required by most models; six generate their own — Veo 3.1, Seedance v1.5, Kling v2.6 and the three Wan models
Prompt
Required on seven of the fourteen, up to 2000 characters. Direction, not dialogue
Rights
Permission to use the likeness and the voice in the clip

Export specifications

Output
A video clip with the mouth driven by your audio, or by the model's own speech
Modes
Image-to-video, video-to-video, ai-avatar and reference-to-video
Longest take
Up to ten minutes on Infinitalk
Price
Per second, by model and variant. Veo 3.1 is 980 at 720p and at 1080p, 1470 at 4K, 368 on the fast variant
Failed runs
Refunded
Delivery
Rendered clip in your workspace library
What you can control

Fourteen models, each declaring what it takes

Nothing is common to all fourteen — each model declares its own inputs and the app renders exactly those. The controls below are the ones you will meet, with the models that offer them named.

  • Mode

    Image-to-video, video-to-video, ai-avatar or reference-to-video

    Set by what you upload, and it narrows the catalogue to the models that can do that job.

  • Prompt

    Up to 2000 characters, required on seven models

    The app's central control, and it is direction rather than dialogue — pace, where they look, what their hands do. On the models that generate their own speech, it is both.

  • Negative prompt

    Up to 600 characters

    What to keep out of the take when the prompt alone keeps pulling it in.

  • Audio track

    Taken by eight of the fourteen

    The track the mouth follows. Veo 3.1, Seedance v1.5, Kling v2.6 and the three Wan models take none and produce the speech themselves.

  • Resolution

    Per model

    The output size the clip renders at. Each model publishes its own set, and it is usually the price control — though not on Veo 3.1, where 720p and 1080p cost the same.

  • Duration

    Per model

    How long the synced clip runs, and it is validated against your source before the job is submitted so a paid run does not fail at the provider. Infinitalk is the long-form option, holding up to ten minutes.

  • Aspect ratio

    Per model

    The frame the render lands in, so a vertical cut does not need reframing afterwards.

  • Mode (HeyGen)

    Speed or precision

    On HeyGen Lipsync v3. Speed for drafts, precision when the mouth has to survive a close-up.

Engine choice

Pick the one that matches your source.

  • HeyGen Lipsync v3

    Video-to-video

    Video-to-video, with a speed mode and a precision mode.

    Dubbing existing footage

  • Sync Lipsync v2

    Dubbing

    Video-to-video dubbing across a whole clip.

    Replacing a track on a finished cut

  • OmniHuman

    Faces

    ByteDance's model, built around human subjects and faces.

    Close-ups of real people

  • Infinitalk

    Long-form

    Long-form, up to ten minutes, and it supports singing.

    Talks, lessons and vocal takes

  • Kling AI Avatar

    Avatar

    Multilingual avatar delivery from a portrait.

    A speaker you do not have on film

  • Kling Avatar v2

    Avatar

    The newer Kling avatar model, also multilingual.

    Avatar work in more than one language

  • Kling v2.6 Lipsync

    Own audio

    Takes no audio track — it generates what is spoken from the prompt rather than following a supplied one.

    When you have no audio to sync to

  • Veo 3.1 Lipsync

    Own audio

    Google's model, generating its own audio. 980 credits per second at both 720p and 1080p — the two cost the same — 1470 at 4K, and 368 on the fast variant.

    A finished shot with speech and no source track

  • Wan 2.5

    Own audio

    Talking avatar generation, with no audio track taken.

    Avatar clips

  • Wan 2.5 Fast

    Own audio

    The quicker Wan 2.5 variant, also without an audio input.

    Drafts and iteration

  • Wan 2.7

    Own audio

    The newer Wan talking-avatar model, also without an audio input.

    Avatar clips

  • Seedance v1.5

    Own audio

    Seedance's talking output, taking no audio track.

    Stylised delivery

  • Seedance 2.0 Lipsync

    Seedance

    The newer Seedance lipsync model, driven by a supplied track.

    Stylised delivery

  • P-Video Avatar

    Affordable

    Fast and affordable avatar delivery.

    Volume work

Who it's for

Built for footage whose audio changed after the shoot

  1. Localisation teams

    Put a translated track on the original performance rather than subtitling over it.

  2. Course and training teams

    Update a line in a lesson without bringing the presenter back in front of a camera.

  3. Marketing teams

    Recut a spot for a new claim or a new market from the footage you already have.

One product. Pick your volume.

Built for Growth at Every Stage

Every plan includes every model and every feature. Plans only change how many credits you get and how many generations run at once.

Pro

Occasional projects

2,900/month

Get Started

Credits per month

60,000

≈1,000 images or ≈12 videos


At once

6 parallel generations

Max

Daily production

5,900/month

Get Started

Credits per month

150,000

≈2,700 images or ≈30 videos


At once

8 parallel generations

Ultimate

High-volume

9,900/month

Get Started

Credits per month

260,000

≈4,700 images or ≈55 videos


At once

10 parallel generations

Enterprise

Teams & agencies

1,00,000+/month

Contact Us

Credits per month

Custom volume

≈1,000 images or ≈12 videos


At once

Custom concurrency

  • Credits refresh monthly
  • Top-up additional credits anytime
  • Unused credits don't roll over

Not ready for the commitment?

Pay as you go

Free to sign up. Buy credits when you need them, same models, same features as every plan above.

Buy Credits

Credit packages

15 credits per 1

You pay You get
500 7,500 credits
1,000 15,000 credits
2,000 30,000 credits

FAQ

Frequently asked questions

Everything you need to know about Lipsync on BeHooked.

Do I need rights to the face and the voice?

Yes. You need permission to use the likeness in the source and the voice in the audio. Do not sync a person who has not agreed to it.

Is there a prompt, and what goes in it?

Yes, and it is the app's central control — required on seven of the fourteen models and capped at 2000 characters. It is direction rather than dialogue: pace, where they look, what their hands do. On the models that generate their own speech it is both, because it also decides what is said. A separate negative prompt of up to 600 characters names what to keep out.

Do I have to supply an audio track?

It depends on the model. Eight of the fourteen take a track and drive the mouth from it. Six take none and produce the speech themselves: Veo 3.1, Seedance v1.5, Kling v2.6 and all three Wan models. The app renders the audio slot only on the models that accept one.

Why does my take come back wrong?

Almost always the framing of the face. Front-on, mouth visible, nothing crossing it — a hand near the jaw or hair over the mouth is the most common reason a take comes back wrong. Clip length is the other one, and the app validates that before submitting rather than after, so a run that would fail at the provider does not get paid for.

How many models are there, and how do I choose?

Fourteen, across four modes — image-to-video, video-to-video, ai-avatar and reference-to-video — and what you upload decides which of them you are in. From there: HeyGen Lipsync v3 or Sync Lipsync v2 for video-to-video dubbing, OmniHuman for close-ups of human faces, the Kling, Wan and P-Video models for avatar work, Infinitalk when the take is long, and Veo 3.1 when you have no audio at all. Nothing is common to all fourteen — each model declares what it takes and the app renders exactly that.

What does it cost?

Per second, and the rate depends on the model and the variant you pick. Veo 3.1 Lipsync is 980 credits per second at 720p and the same 980 at 1080p — the two resolutions cost the same, so there is no reason to render the smaller one — rising to 1470 at 4K, while its fast variant is 368. Failed runs are refunded.

What is the longest clip it can hold?

Infinitalk is the long-form model, running up to ten minutes. It also supports singing, which most of the catalogue does not. Uploads are capped at 200 MB per slot.

What is the difference between speed and precision mode?

They are the two modes on HeyGen Lipsync v3, not a setting the whole catalogue shares. Speed is for drafts and review passes; precision is what you want when the mouth is on screen in a close-up.

Put the right words in the right mouth

Experience Lipsync and every specialized AI app in your workspace.

Try Lipsync