Translation that matches the mouth
Upload a video of someone speaking. HeyGen transcribes it, translates it, synthesises the speech in the same person’s voice in the target language, and reanimates their mouth so the lip movement fits.
The result is the same person, recognisably themselves, apparently fluent. For a founder’s message, a course, or a product explanation, that is materially different from a subtitle or a dubbed voice that plainly is not them.
Three things decide whether it works, and all three are about the source rather than the tool. A clear frontal face — profile angles and heads that turn produce visible artefacts. Clean single-speaker audio — overlapping speech confuses both transcription and voice synthesis. Stable framing — heavy camera movement or a subject moving in and out of frame degrades the mouth reanimation.
A talking head recorded on a decent webcam works well. Conference footage from the back of a room does not.
Consent for a face is not the same as consent for a voice
This tool clones a likeness and a voice together, and puts words in someone’s mouth that they never said. That is precisely the mechanism behind synthetic media abuse, and the fact that your use is legitimate does not change what the output is.
Get explicit written permission covering both the likeness and the voice, the specific languages, and the content. Someone consenting to an English training video has not consented to being made fluent in eight languages saying things they cannot check.
Three practical points that catch organisations out.
Verification of a translation. The person whose face and voice are used cannot check what they appear to be saying in a language they do not speak. Someone who does speak it should review before publication — that is a duty to the person as much as a quality step.
What happens when they leave. A custom avatar of a departed employee or a former spokesperson is an asset your organisation holds and they cannot retract. Settle that in the agreement rather than discovering it during an exit negotiation.
Client work. Put it in the contract, and be specific about who holds the avatar after the engagement ends.
HeyGen operates verification and content controls. Those protect the platform. The obligation to the person on screen is yours.
The avatar workflow
Stock presenters, or a custom avatar recorded from a real person, driven by a typed script — the same economics as Synthesia, where a script change is a re-render rather than a re-shoot.
An interactive avatar mode responds in real time, aimed at kiosks, support and demos.
What you have to supply
- A browser and an account. A free tier exists for trying it; paid plans for anything real.
- Video minutes are metered — the resource that decides plan size.
- Clean source video for translation, as described above.
- For a custom avatar, a recording session and written consent.
- An API for programmatic generation, which is how personalised-video-at-scale is built.
Translation quality is the model’s, not a translator’s
The translation is machine translation, with the fluency and the failure modes that implies.
Idiom, humour, technical terminology, product names and anything culturally specific are where it goes wrong — and it goes wrong fluently, producing confident, natural-sounding text that means something slightly different.
For internal communication that is usually acceptable. For marketing, legal or safety-relevant content it is not, and a native speaker should read the transcript before it is rendered. Reviewing the text costs minutes; reviewing after rendering means paying for the render twice.
Where it earns its cost
Companies taking existing video to international markets — one recording becomes a dozen localised versions without re-filming or hiring voice talent per language.
Course creators and educators with a library worth translating, where the presenter’s presence is part of the product.
Sales and marketing teams personalising outreach video at volume, and support teams building interactive avatar experiences.
It is a poor fit for cinematic work, for multi-speaker footage, and for organisations that cannot obtain unambiguous consent from everyone on screen.
Its real advantages
- Lip-synced translation keeping the original speaker’s voice — the strongest single feature in this batch.
- Fast custom avatars, with a lower setup burden than competitors.
- A usable free tier, unusual in presenter software.
- Interactive avatars for live use rather than rendered video only.
- An API, for personalised video at scale.
The limitations to plan around
- Source quality decides everything — poor lighting, side angles or overlapping speech produce visibly wrong lip-sync.
- Machine translation errors are fluent and need a native speaker to catch.
- Avatars still read as synthetic over longer durations.
- Metered minutes that a translation library consumes quickly.
- The consent burden is real and sits with you, not the platform.
What happens to an avatar you built
Rendered videos download and are yours. Custom avatars and cloned voices stay in HeyGen and cannot be exported.
That has a consequence people do not anticipate: if you stop paying, you cannot produce new video of your own spokesperson, and the recording session would have to be repeated elsewhere.
Keep the original source recordings used to build any avatar. With them, the asset can be rebuilt; without them it cannot. And keep the signed consents alongside them — the recordings and the permissions belong together, and organisations routinely retain one and lose the other.
Related tools
- Synthesia — the direct competitor, stronger on enterprise controls and templates.
- ElevenLabs — dubbing without the face, and better raw voice quality.
- Captions — mobile-first, aimed at social rather than corporate video.
- Descript — when the job is editing footage rather than translating it.
Compiled from HeyGen’s documentation and public sources. We have not hands-on tested this tool. Last reviewed 16 August 2026.