BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/AI Dubbing: Why It Sounds Wrong to a Native Ear
ToolsAugust 26, 2026
Read · 5 min
ai dubbing · video dubbing

AI Dubbing: Why It Sounds Wrong to a Native Ear

AI dubbing optimises lip sync and timing, yet 319 hours of professional dubs say those matter least. What actually decides whether a dub is usable.

Key takeaways
  • Four constraints decide whether a dub works: isochrony, isometry, lip sync and the preserved performance. Vendor demos blur them together on purpose.
  • A study of 319.57 hours of professionally dubbed video across 54 titles concluded that vocal naturalness and translation quality matter more than the character length and lip sync constraints the field optimises hardest.
  • Engineering effort has gone the other way. Recent systems paraphrase for duration and warp vowels for sync, and report beating voice actors on those objective measures.
  • That gap explains the queasy feeling: a dub can win every metric a machine can compute and still sound wrong to a native ear.
  • Your footage decides the difficulty. A visible mouth is the hard case, a voiceover with the speaker off camera is materially easier, and a screen recording with no face removes the hardest constraint entirely.
  • If an identifiable performer's voice is involved, consent is a contracted matter under union agreements now, and it has to be settled before the work, not after.

The demo is always the same demo. Somebody plays thirty seconds of a talking head in English, clicks once, and the same person is speaking Spanish in what sounds like their own voice, mouth roughly in agreement with the sound. It is genuinely impressive. Then you send it to a friend in Madrid and they watch for eleven seconds and say, without being able to say why, that something is off.

That gap between the demo and the native ear is not mysterious. It has been measured, and the measurement points somewhere counterintuitive. The parts of dubbing that are easiest to score are the parts that matter least, and the parts that decide whether a viewer forgets they are watching a dub are the parts nobody has a good number for.

Four words that vendor pages blur together

Before evaluating anything, you need the vocabulary, because dubbing marketing uses "sync" to mean at least three different things.

Isochrony is whether the dub occupies the same time as the original, including its pauses and its speech rate. Virkar, Federico, Enyedi and Barra-Chicote define it in their work on prosodic alignment for automatic dubbing as translating the original speech while also matching its prosodic structure of phrases and pauses. A dub can be the right length overall and still break isochrony if it puts its breath in the wrong place.

Isometry is narrower: the translated text has roughly the same character or syllable count as the source. It is a text property, measurable without ever rendering audio, which is precisely why so much research optimises it.

Lip sync is whether the visible mouth shapes agree with the sound coming out. It only exists as a problem when there is a mouth on screen.

Performance preservation is whether emphasis, emotion and voice character survive the crossing. It is the hardest to measure and, as the evidence below shows, the one that decides how the result feels.

Breakdown diagram showing the four constraints that together make a usable dub, isochrony, isometry, lip sync and preserved performance

What did 319 hours of professional dubbing actually show?

It showed that the field has its priorities inverted. William Brannon, Yogesh Virkar and Brian Thompson assembled a corpus of 319.57 hours of video from 54 professionally produced titles and studied how human localisers actually make their choices, publishing the result in the Transactions of the Association for Computational Linguistics.

Their conclusion is stated flatly and it contradicts a decade of engineering emphasis. They argue for the importance of vocal naturalness and translation quality over the commonly emphasised isometric and lip sync constraints, and for a more qualified view of how much isochronic constraints matter. They also found substantial influence of the source audio on human dubs through channels other than the words of the translation, which is a careful way of saying that professionals listen to the original performance and carry something across that is not in the script.

"...the importance of vocal naturalness and translation quality over commonly emphasized isometric (character length) and lip-sync constraints."Brannon, Virkar and Thompson, TACL vol. 11, 2023

Now put that next to what the systems are being built to do. PS-TTS, a recent phonetic synchronisation approach, paraphrases the translated text with a language model until the target duration matches the source, then applies dynamic time warping with local costs based on vowel distances so the target vowels resemble the source vowels visually. Tested across Korean, English and French, the authors report that their systems outperform voice actors on objective metrics in the Korean to English and English to Korean directions.

Read those two findings together and the contradiction resolves itself cleanly. The machines are winning, convincingly, on exactly the measures that professional practice treats as secondary. Nothing about that is fraud, and the engineering is real. It is simply optimisation against the available scoreboard, and the scoreboard is missing the column that matters.

One more thing about the corpus is worth stating, because it changes how you read every vendor claim built on a research paper. The 54 titles studied were professionally produced dubs, meaning the reference point is not an amateur baseline but the output of people who do this for a living in an industry with decades of craft behind it. When a system reports beating voice actors, check which voice actors and on which measure. Beating a professional on vowel alignment is a real result. It is not the same claim as beating one on whether the scene works.

The practical reading for a seller is that the research is telling you where to spend a limited budget. If the money can only cover one human intervention, the evidence says buy the translation review and the voice direction, not the sync pass. The sync is the part the machine already does better than you could specify.

How do you check a sample clip in ten seconds?

Play it with your eyes closed first. Almost every failure that a native speaker reacts to is audible without the picture, and the picture is what makes a bad dub look acceptable in a demo. Once you have listened blind, run the table below against the same clip.

ConstraintWhat it isHow the field measures itMachines todayYour ten second check
IsochronySame duration, same pauses, same rateSpeech overlap and pause alignment against the sourceStrong, and improving fastestDoes the speaker still breathe where they breathed?
IsometryTranslated text matches source lengthCharacter or syllable count ratioSolved, it is a text operationDoes the sentence sound compressed or padded?
Lip syncVisible mouth agrees with the soundLip sync error confidence and distance scoresGood on frontal faces, weak in profile and motionWatch the closing consonants, b, p and m
Translation qualityMeaning, register and idiom surviveLearned metrics such as COMET, plus human reviewFluent, and confidently wrong on registerWould a customer say it that way to a friend?
Vocal naturalnessVoice sounds like a person, not a readerHuman rating, no reliable automatic proxyThe weak point, and the one that decidesClose your eyes. Is anyone in the room?
Performance transferEmphasis and emotion cross overBarely measured at allLargely absentDoes the joke land, or just get said?

The last column is the one to actually use. Every vendor will quote you a sync score. None of them will quote you a number for the bottom two rows, because a trustworthy one does not exist.

Why does a technically excellent dub still feel wrong?

Because the errors that survive automated scoring are the ones a listener is most sensitive to. A machine can place every syllable inside the source duration and still put the stress on the wrong word in the sentence, and stress is how a listener knows what the speaker means. Nothing in an isometry score notices that.

The same failure appears in text translation, where fluency has outrun accuracy and produced errors that read beautifully. We wrote about that specific trap in the piece on why fluent machine translation is harder to catch than clumsy machine translation, and dubbing inherits the whole problem then adds a clock and a face to it.

There is a second reason, and it is structural. The corpus study found that human dubbers take cues from the source audio that never appear in the words. A professional hears hesitation, sarcasm or restraint and reproduces it. A pipeline that translates a transcript and then speaks the translation has thrown that signal away before the voice is ever generated, and it inherits every error in that transcript too, which is why a vendor accuracy figure needs its test set and normalisation rules before it means anything, no matter how good the voice model is.

Which of the three kinds of footage do you have?

This is the most useful question in the article, and it has direct research backing rather than being a matter of taste. The prosodic alignment work draws an explicit line between on screen dubbing, where a visible mouth demands strict synchronisation, and off screen dubbing, which needs materially less stringent synchronisation constraints. The authors extended one model to handle both inside a single video, and tested it on English into French, Italian, German and Spanish over TED Talks and public video, reporting a better subjective viewing experience on mixed content. The full paper sits in the ISCA archive for Interspeech 2022.

Turn that straight into a decision about your own library.

A talking head with the face on camera is the hard case, and the only one where lip sync is a real constraint rather than a marketing line. Budget for a human pass, or accept that a portion of your audience will notice.

A voiceover with the speaker off camera is a different and much easier problem, and the research says so in as many words. Product walkthroughs, tutorials with the narrator absent, most course modules: the sync requirement drops to keeping the words over the right part of the picture.

A screen recording with no face at all removes lip sync entirely. This is where automated dubbing is genuinely close to solved, and where a small seller with a course catalogue can localise at a cost that makes a second market viable.

Saying which of the three you have is more useful than any blanket verdict about the technology. Most catalogues are a mix, and the honest plan splits them rather than treating the library as one job. If your product pages travel alongside the video, the same split logic applies to the text, which we covered in the piece on what actually blocks a sale in a multilingual online store.

Card showing the three footage types that decide dubbing difficulty, a face on camera, a voiceover off camera, and a screen recording with no face

Who has to consent before a voice is cloned?

If an identifiable performer is involved, consent is now a contracted matter rather than a courtesy, and it has to be settled before the work rather than discovered afterwards. Union agreements in the United States have moved on this specifically, and the standard is stricter than most buyers assume.

Under the SAG-AFTRA Interactive Media Agreement, a vocal digital replica is defined as an AI generated voice created primarily from a performer's union covered work in order to generate new dialogue. The consent standard, as summarised by the media law practice at Frankfurt Kurnit, requires clear and conspicuous writing in a separate document or a specifically signed rider, a reasonably specific description of how the replica will be used, and disclosure of whether it will drive real time generation. Blanket consent obtained at the point of hiring does not carry forward to future productions. Compensation is calculated per line generated, with a line treated as roughly ten words, and usage reports are owed after release.

Two practical consequences for anyone with a video library. If you hired a narrator, your original contract almost certainly does not grant you the right to build a synthetic model of their voice, because that right did not exist as a category when most freelance narration contracts were written. And if you are the person whose face and voice are in the videos, you are the performer here, which makes this an asset question rather than a compliance one.

The adjacent risk is worth naming because it lands on ordinary businesses rather than studios. The same synthesis quality that makes dubbing viable, built on the pipeline we describe in how an AI voice is assembled stage by stage, makes impersonation cheap, which is the subject of our piece on the phone checks that survive a cloned voice. A seller publishing a library of their own narrated video is publishing training data for their own voice. That is not a reason to stop. It is a reason to have the callback procedure in place before the library is public.

Is AI dubbing good enough to use commercially?

For screen recordings and off camera narration, yes, and the ceiling is higher than most sellers assume. For anything where a face is talking to the camera in close up, treat the machine output as a draft that a native speaker reviews, because the failures cluster exactly where your credibility lives.

Should I use subtitles instead?

Subtitles cost a fraction as much, carry none of the consent complexity, and are the correct answer whenever the video is short or the budget is tight. Dubbing earns its cost when the viewer needs their eyes on what is happening on screen, which is the ordinary case for a demonstration or a hands on tutorial.

Can I trust a vendor that publishes benchmark numbers?

Publishing numbers is better than not publishing them, and it still tells you little about what you will hear. The measures that get published are the ones that can be computed, which are the sync and length constraints the corpus study puts second. Ask instead for three raw clips in your target language, chosen by you, and hand them to a native speaker with no context.

The test nobody has built

No benchmark in this field measures whether a native speaker forgets they are watching a dub. Every published metric answers a proxy question: how close is the timing, how similar are the vowels, how well does a learned model score the translation. The commercial question is a different one, and it is binary. Did the viewer stay inside the video, or did they notice the machine and drop out?

Until somebody builds that measurement, the practical method is the unglamorous one. Split your library by footage type. Run the easy tier through automation and spend the saving on a human pass over the hard tier. Test with three real clips and one real native speaker, not with the vendor's showreel. And do the consent paperwork before you upload a single file, because that is the only part of this that cannot be fixed afterwards.

The upside is worth the discipline. A course catalogue or a product video library that opens into a second language is the cheapest new market a small seller can buy, far cheaper than the equivalent in advertising. If the storefront around that catalogue needs rebuilding to carry a second language properly, our AI store builder generates the multilingual structure on code you own outright, which keeps the localisation decisions yours rather than a platform's.

The technology is further along than the sceptics say and less finished than the demos suggest. Both things are true, and which one governs your project depends almost entirely on whether there is a mouth on screen.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building