AI Lip Sync for Photos and Videos

Make a portrait photo speak with MiniMax H3 Max, or re-sync the mouth in a video you already have with Video Lip Sync. Add an audio track and the mouth follows your recording.

Cargando componente del modelo...

Vista previa

Clips made with this exact tool

Both clips below started as a single still photo and one audio track — nothing else. Turn the sound on; the point is whether the mouth matches the voice.

Talking-head ad read

12s

A 768x1024 portrait plus a 12.05-second voice track. The output came back at 768x1024 and 12.05 seconds — the canvas follows the photo, the length follows the audio.

Spokesperson intro

11.5s

Same setup, a different face and an 11.50-second read. Head movement and blinking are generated too, so it does not read as a static photo with a moving mouth.

Generated on MiniMax H3 Max Lip Sync at 768P. The portraits are AI-generated and the voices are synthetic — no real person is depicted. Both clips are compressed for the web; the originals are about 10 MB each.

What these two lip sync models actually do

Make a photo speak, or put a new voice in footage you already shot.

A still photo starts speaking

Upload a single front-facing portrait and MiniMax H3 Max animates the mouth to your audio. No avatar to build, no rig to set up, no video footage needed.

Re-dub a video you already have

Video Lip Sync re-syncs the mouth in real footage to a new voice track. The result keeps your video's resolution up to 1080p and comes back as MP4 at 25fps, as long as your audio: a longer video is trimmed, a shorter one loops.

Photo clips follow your audio

There is no duration to guess at. MiniMax H3 Max returns a clip exactly as long as the audio you gave it, up to 14.8 seconds, so the sync never drifts out at the end. Video Lip Sync follows your audio too, up to 60 seconds.

Two scenes, one switch

On the Video tab, Front-facing, single speaker renders fastest on talking-head footage and unlocks the looping and start-time options. Complex scenes, multiple shots detects the cuts and identifies the speaker in each shot, so wide shots and edits are handled instead of smeared - and you can switch that detection off in Advanced settings when the footage is one continuous take.

Transcript-guided mouth shapes

The model can transcribe your audio and use the words to drive the mouth, which sharpens consonants in spoken lines. Switch it off and it syncs to the waveform instead.

480P through 2K on photos

On the Photo tab, draft at 480P, work at 768P, and render at 1080P or 2K when the clip is going somewhere public. 2K is available on lip sync specifically. Video Lip Sync has no resolution picker: it keeps whatever your footage is, up to 1080p.

How to lip sync a photo or a video

Three steps in either tab, and you bring your own audio.

1

Pick a tab and upload

Photo takes one clear, front-facing portrait with an aspect ratio between 0.4 and 2.5. Video takes MP4 or MOV footage at 24 to 60fps, 360p and up; anything above 1080p is scaled down to 1080p.

2

Add your audio track

Upload a recording or capture one from your mic. Photo needs at least 5 seconds and clips anything past 14.8. Video accepts up to 60 seconds and keeps all of it.

3

Set the options and generate

Photo picks a resolution from 480P to 2K. Video picks the scene mode, and Front-facing, single speaker adds looping and a start time under Advanced settings. Both show the exact credit cost before you spend anything.

AI Lip Sync FAQ

What it does, and what it does not.

Give a photo a voice, or a video a new one

Upload a face photo or the footage you already shot, add your audio, and get a lip-synced clip back.

See how it works