AI Talking Avatar
Upload a face and a voice clip. The audio drives the mouth, the timing and the delivery.
Upload a face photo and an audio clip
Face photo
Voice clip
Tap to choose a file
BeforeThe audio does the acting
This is the one video tool here where you do not describe what should happen. The recording decides it. Pace, pauses, emphasis and the shape of each sound all come out of the audio and into the face.
That has a practical consequence worth knowing before you start: improving the recording improves the video far more than changing the photo does. A clean take of ordinary speech beats a dramatic take recorded in a noisy room every time.
How to make a talking avatar
Upload a face photo
Front-facing, whole face visible, sharp and evenly lit.
Upload a voice clip
Clean speech. This is what drives the mouth and the timing.
Generate
The face is animated to the audio. No prompt needed.
Download
An mp4 you can post, send or drop into an edit.
Recording audio that works
Record in a small room with soft furnishings. A kitchen or a bathroom adds echo, and echo smears the boundaries between sounds, which is exactly the information the model uses to shape the mouth.
Keep the phone about a hand\u2019s width from your mouth and speak at a normal pace. Too close gives you plosive thumps on every P and B; too far brings the room back in.
Leave a half second of silence at the start and end. Audio that begins mid-word gives the face nothing to settle from, and the first moment of the video is the one people judge.
What to use it for
Narrated explainers, product walkthroughs, course intros and messages where you want a face on screen without filming one. It is also the quickest way to put a presenter in front of a script you have already written.
Use your own face and your own voice, or someone who has agreed to both. Our AI restricted policy sets out what is not allowed. If you want movement rather than speech, the photo animate tool is the other half of this.
Frequently asked questions
A front-facing photo where the whole face is visible, sharp and evenly lit, with the mouth closed or only slightly open. Sunglasses, a hand near the face, a strong side angle or a hat over the brow all make the result worse.
Clean speech with little background noise. Music behind the voice, echo from a large room and two people talking at once all confuse the timing, because the model is reading the speech to decide how the mouth should move.
The mouth follows the sound rather than the words, so it is not tied to one language. Very fast speech is harder than measured speech in any language.
Only with their permission. Putting words in the mouth of a real, identifiable person who did not agree is exactly what this must not be used for, whatever the tool makes technically possible.
Video has its own daily free allowance, separate from image editing and generating. An Unlimited plan removes the daily limit.