- Four constraints decide whether a dub works: isochrony, isometry, lip sync and the preserved performance. Vendor demos blur them together on purpose.
- A study of 319.57 hours of professionally dubbed video across 54 titles concluded that vocal naturalness and translation quality matter more than the character length and lip sync constraints the field optimises hardest.
- Engineering effort has gone the other way. Recent systems paraphrase for duration and warp vowels for sync, and report beating voice actors on those objective measures.
- That gap explains the queasy feeling: a dub can win every metric a machine can compute and still sound wrong to a native ear.
- Your footage decides the difficulty. A visible mouth is the hard case, a voiceover with the speaker off camera is materially easier, and a screen recording with no face removes the hardest constraint entirely.
- If an identifiable performer's voice is involved, consent is a contracted matter under union agreements now, and it has to be settled before the work, not after.
The demo is always the same demo. Somebody plays thirty seconds of a talking head in English, clicks once, and the same person is speaking Spanish in what sounds like their own voice, mouth roughly in agreement with the sound. It is genuinely impressive. Then you send it to a friend in Madrid and they watch for eleven seconds and say, without being able to say why, that something is off.
That gap between the demo and the native ear is not mysterious. It has been measured, and the measurement points somewhere counterintuitive. The parts of dubbing that are easiest to score are the parts that matter least, and the parts that decide whether a viewer forgets they are watching a dub are the parts nobody has a good number for.
Four words that vendor pages blur together
Before evaluating anything, you need the vocabulary, because dubbing marketing uses "sync" to mean at least three different things.
Isochrony is whether the dub occupies the same time as the original, including its pauses and its speech rate. Virkar, Federico, Enyedi and Barra-Chicote define it in their work on prosodic alignment for automatic dubbing as translating the original speech while also matching its prosodic structure of phrases and pauses. A dub can be the right length overall and still break isochrony if it puts its breath in the wrong place.
Isometry is narrower: the translated text has roughly the same character or syllable count as the source. It is a text property, measurable without ever rendering audio, which is precisely why so much research optimises it.
Lip sync is whether the visible mouth shapes agree with the sound coming out. It only exists as a problem when there is a mouth on screen.
Performance preservation is whether emphasis, emotion and voice character survive the crossing. It is the hardest to measure and, as the evidence below shows, the one that decides how the result feels.
What did 319 hours of professional dubbing actually show?
It showed that the field has its priorities inverted. William Brannon, Yogesh Virkar and Brian Thompson assembled a corpus of 319.57 hours of video from 54 professionally produced titles and studied how human localisers actually make their choices, publishing the result in the Transactions of the Association for Computational Linguistics.
Their conclusion is stated flatly and it contradicts a decade of engineering emphasis. They argue for the importance of vocal naturalness and translation quality over the commonly emphasised isometric and lip sync constraints, and for a more qualified view of how much isochronic constraints matter. They also found substantial influence of the source audio on human dubs through channels other than the words of the translation, which is a careful way of saying that professionals listen to the original performance and carry something across that is not in the script.
Now put that next to what the systems are being built to do. PS-TTS, a recent phonetic synchronisation approach, paraphrases the translated text with a language model until the target duration matches the source, then applies dynamic time warping with local costs based on vowel distances so the target vowels resemble the source vowels visually. Tested across Korean, English and French, the authors report that their systems outperform voice actors on objective metrics in the Korean to English and English to Korean directions.
Read those two findings together and the contradiction resolves itself cleanly. The machines are winning, convincingly, on exactly the measures that professional practice treats as secondary. Nothing about that is fraud, and the engineering is real. It is simply optimisation against the available scoreboard, and the scoreboard is missing the column that matters.
One more thing about the corpus is worth stating, because it changes how you read every vendor claim built on a research paper. The 54 titles studied were professionally produced dubs, meaning the reference point is not an amateur baseline but the output of people who do this for a living in an industry with decades of craft behind it. When a system reports beating voice actors, check which voice actors and on which measure. Beating a professional on vowel alignment is a real result. It is not the same claim as beating one on whether the scene works.
The practical reading for a seller is that the research is telling you where to spend a limited budget. If the money can only cover one human intervention, the evidence says buy the translation review and the voice direction, not the sync pass. The sync is the part the machine already does better than you could specify.
How do you check a sample clip in ten seconds?
Play it with your eyes closed first. Almost every failure that a native speaker reacts to is audible without the picture, and the picture is what makes a bad dub look acceptable in a demo. Once you have listened blind, run the table below against the same clip.
| Constraint | What it is | How the field measures it | Machines today | Your ten second check |
|---|---|---|---|---|
| Isochrony | Same duration, same pauses, same rate | Speech overlap and pause alignment against the source | Strong, and improving fastest | Does the speaker still breathe where they breathed? |
| Isometry | Translated text matches source length | Character or syllable count ratio | Solved, it is a text operation | Does the sentence sound compressed or padded? |
| Lip sync | Visible mouth agrees with the sound | Lip sync error confidence and distance scores | Good on frontal faces, weak in profile and motion | Watch the closing consonants, b, p and m |
| Translation quality | Meaning, register and idiom survive | Learned metrics such as COMET, plus human review | Fluent, and confidently wrong on register | Would a customer say it that way to a friend? |
| Vocal naturalness | Voice sounds like a person, not a reader | Human rating, no reliable automatic proxy | The weak point, and the one that decides | Close your eyes. Is anyone in the room? |
| Performance transfer | Emphasis and emotion cross over | Barely measured at all | Largely absent | Does the joke land, or just get said? |
The last column is the one to actually use. Every vendor will quote you a sync score. None of them will quote you a number for the bottom two rows, because a trustworthy one does not exist.
Why does a technically excellent dub still feel wrong?
Because the errors that survive automated scoring are the ones a listener is most sensitive to. A machine can place every syllable inside the source duration and still put the stress on the wrong word in the sentence, and stress is how a listener knows what the speaker means. Nothing in an isometry score notices that.
The same failure appears in text translation, where fluency has outrun accuracy and produced errors that read beautifully. We wrote about that specific trap in the piece on why fluent machine translation is harder to catch than clumsy machine translation, and dubbing inherits the whole problem then adds a clock and a face to it.
There is a second reason, and it is structural. The corpus study found that human dubbers take cues from the source audio that never appear in the words. A professional hears hesitation, sarcasm or restraint and reproduces it. A pipeline that translates a transcript and then speaks the translation has thrown that signal away before the voice is ever generated, and it inherits every error in that transcript too, which is why a vendor accuracy figure needs its test set and normalisation rules before it means anything, no matter how good the voice model is.
Which of the three kinds of footage do you have?
This is the most useful question in the article, and it has direct research backing rather than being a matter of taste. The prosodic alignment work draws an explicit line between on screen dubbing, where a visible mouth demands strict synchronisation, and off screen dubbing, which needs materially less stringent synchronisation constraints. The authors extended one model to handle both inside a single video, and tested it on English into French, Italian, German and Spanish over TED Talks and public video, reporting a better subjective viewing experience on mixed content. The full paper sits in the ISCA archive for Interspeech 2022.
Turn that straight into a decision about your own library.
A talking head with the face on camera is the hard case, and the only one where lip sync is a real constraint rather than a marketing line. Budget for a human pass, or accept that a portion of your audience will notice.
A voiceover with the speaker off camera is a different and much easier problem, and the research says so in as many words. Product walkthroughs, tutorials with the narrator absent, most course modules: the sync requirement drops to keeping the words over the right part of the picture.
A screen recording with no face at all removes lip sync entirely. This is where automated dubbing is genuinely close to solved, and where a small seller with a course catalogue can localise at a cost that makes a second market viable.
Saying which of the three you have is more useful than any blanket verdict about the technology. Most catalogues are a mix, and the honest plan splits them rather than treating the library as one job. If your product pages travel alongside the video, the same split logic applies to the text, which we covered in the piece on what actually blocks a sale in a multilingual online store.
Who has to consent before a voice is cloned?
If an identifiable performer is involved, consent is now a contracted matter rather than a courtesy, and it has to be settled before the work rather than discovered afterwards. Union agreements in the United States have moved on this specifically, and the standard is stricter than most buyers assume.
Under the SAG-AFTRA Interactive Media Agreement, a vocal digital replica is defined as an AI generated voice created primarily from a performer's union covered work in order to generate new dialogue. The consent standard, as summarised by the media law practice at Frankfurt Kurnit, requires clear and conspicuous writing in a separate document or a specifically signed rider, a reasonably specific description of how the replica will be used, and disclosure of whether it will drive real time generation. Blanket consent obtained at the point of hiring does not carry forward to future productions. Compensation is calculated per line generated, with a line treated as roughly ten words, and usage reports are owed after release.
Two practical consequences for anyone with a video library. If you hired a narrator, your original contract almost certainly does not grant you the right to build a synthetic model of their voice, because that right did not exist as a category when most freelance narration contracts were written. And if you are the person whose face and voice are in the videos, you are the performer here, which makes this an asset question rather than a compliance one.
The adjacent risk is worth naming because it lands on ordinary businesses rather than studios. The same synthesis quality that makes dubbing viable, built on the pipeline we describe in how an AI voice is assembled stage by stage, makes impersonation cheap, which is the subject of our piece on the phone checks that survive a cloned voice. A seller publishing a library of their own narrated video is publishing training data for their own voice. That is not a reason to stop. It is a reason to have the callback procedure in place before the library is public.
Is AI dubbing good enough to use commercially?
For screen recordings and off camera narration, yes, and the ceiling is higher than most sellers assume. For anything where a face is talking to the camera in close up, treat the machine output as a draft that a native speaker reviews, because the failures cluster exactly where your credibility lives.
Should I use subtitles instead?
Subtitles cost a fraction as much, carry none of the consent complexity, and are the correct answer whenever the video is short or the budget is tight. Dubbing earns its cost when the viewer needs their eyes on what is happening on screen, which is the ordinary case for a demonstration or a hands on tutorial.
Can I trust a vendor that publishes benchmark numbers?
Publishing numbers is better than not publishing them, and it still tells you little about what you will hear. The measures that get published are the ones that can be computed, which are the sync and length constraints the corpus study puts second. Ask instead for three raw clips in your target language, chosen by you, and hand them to a native speaker with no context.
The test nobody has built
No benchmark in this field measures whether a native speaker forgets they are watching a dub. Every published metric answers a proxy question: how close is the timing, how similar are the vowels, how well does a learned model score the translation. The commercial question is a different one, and it is binary. Did the viewer stay inside the video, or did they notice the machine and drop out?
Until somebody builds that measurement, the practical method is the unglamorous one. Split your library by footage type. Run the easy tier through automation and spend the saving on a human pass over the hard tier. Test with three real clips and one real native speaker, not with the vendor's showreel. And do the consent paperwork before you upload a single file, because that is the only part of this that cannot be fixed afterwards.
The upside is worth the discipline. A course catalogue or a product video library that opens into a second language is the cheapest new market a small seller can buy, far cheaper than the equivalent in advertising. If the storefront around that catalogue needs rebuilding to carry a second language properly, our AI store builder generates the multilingual structure on code you own outright, which keeps the localisation decisions yours rather than a platform's.
The technology is further along than the sceptics say and less finished than the demos suggest. Both things are true, and which one governs your project depends almost entirely on whether there is a mouth on screen.