Auto-dubbing in 27 languages: what it fixes and what still breaks

YouTube opened auto-dubbing to all creators across 27 languages, and added a layer called Expressive Speech that tries to preserve tone and emotion in eight of them. I spent a few days testing what that actually changes in a 45-second clip, and the answer is less magical and more interesting than the announcement suggests.

Auto-dubbing in 27 languages: what it fixes and what still breaks

Auto-dubbing in 27 languages: what it fixes and what still breaks

The fact: YouTube auto-dubbing stopped being a pilot and became a general feature, with 27 languages supported. Alongside it came Expressive Speech, a layer that tries to preserve the original speaker's tone, emotion and delivery, announced for eight languages.

I spent a few days testing this from the engineering side, with our own material and with clients'. Here is what I found, including the parts that do not work, because those are the ones that decide whether it is worth it in your case.

What the pipeline actually does

Auto-dubbing is not one model. It is four chained stages, and each has its own way of failing.

1. Transcription. Audio becomes timestamped text. The most mature of the four. It stumbles on proper nouns, regional slang and overlapping speech.

2. Translation. The text moves to the target language. Here comes video's first structural problem: a good translation changes sentence length. "I don't know" is three syllables; "no tengo la menor idea" is not. A sentence longer than the available window forces the system either to speed up the delivery or to drop words.

3. Speech synthesis. The translated text becomes audio. This is where Expressive Speech acts: instead of generating a neutral newsreader delivery, it tries to carry the original's prosody, meaning where the voice rises, where it pauses, where it laughs mid-sentence.

4. Alignment. The new audio is fitted into the video timeline. This is the stage nobody talks about and the one that ruins more output than the other three combined.

What Expressive Speech really delivers

Better than I expected, and less than the name promises.

What it catches well: question intonation, emphasis on a specific word, the general rhythm of a fast or slow talker, and the volume lift when someone gets excited. A streamer reacting to something unexpected comes out recognizably excited, not robotic.

What it still misses: genuine laughter mid-sentence, a cracking voice, irony delivered deadpan, and the kind of drawl that is a signature for certain creators. In those cases the result is not bad, it is neutral, and neutral is the worst place for a video that runs on personality.

One technical note that matters: prosody preservation works best when the source audio is clean. On a stream with background music, chat read out loud and a clipping microphone, prosody estimation degrades and Expressive Speech collapses into an ordinary read. If you want to dub, the decision starts upstream, at capture.

Where dubbing breaks in short-form

This is the part that matters to clippers, and the short answer is: short-form is the worst case for dubbing.

Three reasons, in order of severity.

There is no adaptation time. In a twenty-minute video, a viewer finds the voice odd for the first thirty seconds and then forgets about it. In a 45-second clip, the first thirty seconds are the whole video. The oddness never wears off.

The face is in close-up. A well-framed vertical clip puts the speaker's mouth large on screen. Any drift between lips and sound is obvious. Long-form usually sits on a wider shot and forgives more.

Retention is decided in the first three seconds. If the hook depends on the way a person said a sentence, and the dub delivers that sentence half a second late with approximate intonation, you lost exactly what made the clip work.

The test I would run before deciding: take your best clip, dub it, and watch ten seconds without sound, then with sound. If the dubbed version makes you watch the mouth instead of the sentence, the format is not for you.

Where it works very well

Now the other side, because there are cases where dubbing clearly beats captions.

Narration over footage. Tutorials, screen-share analysis, commentated gameplay, faceless channels. There are no visible lips to contradict the voice, and sync stops being a problem. Here dubbing is nearly invisible, in the good sense.

Information-dense content. When the viewer needs to look at the screen to follow (a chart, a map, a snippet of code), captions compete with the image for attention. Dubbing frees the eyes.

Long-form you want to expand without re-recording. A one-hour podcast in three more languages is a month of work. Dubbed, it is an afternoon. Even with lost nuance, the math favors it.

We covered the captioned version of this same decision in translated captions in three languages, and the comparison still holds: on clips, captions win most of the time.

The mistake I saw most in testing

Dubbing and not translating everything else.

A video with Spanish audio, burned-in English captions, an English title and English cover text. The Spanish-speaking viewer clicks, hears their own language and sees three visual elements they cannot read. Drop-off in that case is worse than the original with no dubbing at all, because the video looks broken.

If you are going into another language, go all the way: audio, captions, title, description and on-screen text. Half a translation is worse than none. That connects to what we wrote about titles and descriptions as search signals.

The authenticity question, which is coming

Worth flagging a discomfort that has not become a big argument yet, but will.

Voice is identity. When a creator's voice is synthesized in another language, the viewer is hearing something that person never said, not in that way. That is acceptable when it is declared and unsettling when it is not.

YouTube marks the track as dubbed, which handles the formal part. The cultural part is slower: in some markets, AI-dubbed video already carries a stigma similar to the one we discussed in generic AI-made content. That is not a reason not to use it. It is a reason not to hide it.

How I would decide, in practice

A three-line rule that has held up in our testing:

  • Large face on screen and continuous speech: captions. Always.
  • Narration over footage, no face: dubbing. Wins easily.
  • Mixed format: dub the long video, caption the clips pulled from it.

That last line is the most useful one for clippers. The dubbed original reaches a new audience, and the clips from that same video stay captioned, because it is in the clip that the original voice matters most.

In our own pipeline, that is the path: the transcript is produced once, and from it come both the translated captions for the clip and the text that feeds a dub of the long video. Paste the link into Cut.Pro and you get proposed clips already transcribed, with vertical reframing and captions, and the translation reuses that same text instead of starting over.

The short version

  • YouTube auto-dubbing is now general, with 27 languages; Expressive Speech covers eight.
  • The pipeline has four stages, and alignment is the one that ruins the most output.
  • Expressive Speech catches intonation and rhythm. It misses laughter, deadpan irony and a cracking voice.
  • Short-form is the worst case: no adaptation time, large face, retention decided in three seconds.
  • Narration over footage is the best case: no lips to contradict the voice.
  • Translating halfway is worse than not translating.

The technology genuinely matured this cycle, and that changes the math of international expansion for anyone with a long-form library. It just does not change the clip, which remains the place where a person's own voice is half the value.

Sources: Times of AI, YouTube auto-dubbing across 27 languages · GSMArena, YouTube's new AI tools and monetization options

Share

Keep reading

More insights and tutorials to help you grow as a content creator.