AI Lip Sync from a Photo
Upload a face photo and an audio track. MiniMax H3 Max animates the mouth to match your recording, up to 14.8 seconds per clip.
Preview
Clips made with this exact tool
Both clips below started as a single still photo and one audio track — nothing else. Turn the sound on; the point is whether the mouth matches the voice.
Talking-head ad read
12sA 768x1024 portrait plus a 12.05-second voice track. The output came back at 768x1024 and 12.05 seconds — the canvas follows the photo, the length follows the audio.
Spokesperson intro
11.5sSame setup, a different face and an 11.50-second read. Head movement and blinking are generated too, so it does not read as a static photo with a moving mouth.
Generated on MiniMax H3 Max Lip Sync at 768P. The portraits are AI-generated and the voices are synthetic — no real person is depicted. Both clips are compressed for the web; the originals are about 10 MB each.
What this lip sync model actually does
One photo, one audio track, and a clip whose mouth matches the recording.
A still photo starts speaking
Upload a single front-facing portrait and MiniMax H3 Max animates the mouth to your audio. No avatar to build, no rig to set up, no video footage needed.
Length follows your audio
There is no duration to guess at. The clip comes back exactly as long as the audio you gave it, up to 14.8 seconds, so the sync never drifts out at the end.
Transcript-guided mouth shapes
The model can transcribe your audio and use the words to drive the mouth, which sharpens consonants in spoken lines. Switch it off and it syncs to the waveform instead.
480P through 2K
Draft at 480P, work at 768P, and render at 1080P or 2K when the clip is going somewhere public. 2K is available on lip sync specifically.
How to make a photo talk
Three steps, and you bring your own audio.
Upload a face photo
One clear, front-facing portrait works best. The aspect ratio has to sit between 0.4 and 2.5, and the output keeps whatever shape you upload.
Add your audio track
Upload a recording or capture one from your mic. It needs to be at least 5 seconds; anything past 14.8 seconds is clipped to the first 14.8.
Pick a resolution and generate
Choose 480P, 768P, 1080P or 2K. The panel shows the exact clip length and credit cost before you spend anything, because both are set by your audio.
AI Lip Sync FAQ
What it does, and what it does not.