YouTube is testing an AI tool that could make it easier for creators to publish videos in more languages. Developed with Aloud, a service from Google’s in-house incubator Area 120, the tool handles several steps of dubbing that creators previously had to arrange with outside providers.
How Aloud handles dubbing
Aloud starts by transcribing a video. The creator can review and edit that transcription before the tool translates it and produces a dubbed audio version. That gives creators a chance to correct the text before it becomes another language’s spoken track.
The process connects transcription, translation and audio production in one tool. Previously, creators had to work directly with third-party dubbing providers to create audio tracks, a process that could be time-consuming and expensive. YouTube says Aloud lets creators dub videos at no additional cost.
The tool is being tested with hundreds of creators. YouTube said it plans to make it available to all creators soon. The current language options are English, Spanish and Portuguese, with Hindi and Bahasa Indonesian among the languages expected in the future.
More languages, wider audiences
YouTube introduced support for multi-language audio tracks earlier this year, enabling creators to add dubbing to both new and existing videos. The feature is intended to help videos reach audiences beyond the language in which they were originally made.
As of June 2023, creators had dubbed more than 10,000 videos in over 70 languages, the company told TechCrunch. Those figures show that creators were already using multi-language audio, while YouTube’s new test aims to make producing those tracks simpler.
For creators, the practical appeal is straightforward: a translated audio track can offer viewers another way to follow a video without requiring the creator to produce every version manually. Lowering the effort and cost of dubbing could make that option more accessible, though the tool is still being tested and is not yet open to everyone.
The challenge of making translated audio feel natural
A translated track has to do more than convey the words. YouTube VP of Creator Products Amjad Hanif said the company is working to make translated audio sound like the creator’s voice, with more expression and lip sync.
YouTube also confirmed that it expects generative AI to support future Aloud features, including voice preservation, better emotion transfer and lip reanimation. These are planned capabilities, not features the source says are already available. The goal is to make dubbed speech feel closer to the original performance as well as understandable in another language.
That distinction matters because a creator’s delivery is part of a video, not just its transcript. Preserving voice and emotion could help translated audio retain more of that delivery, while improved lip sync and lip reanimation could make the visuals align more closely with dubbed speech. YouTube’s comments describe a direction for future development rather than a finished result.
A test with broader ambitions
Aloud builds on YouTube’s multi-language audio feature by addressing the work involved in creating a dub. Its transcription and editing step gives creators a review point, while translation and audio generation handle the next parts of the process. The company’s stated plan to expand availability and add languages would extend that workflow beyond the current test.
For now, the tool remains in testing with hundreds of creators and supports three languages. YouTube says a wider rollout will come soon and has named additional languages it may offer later. Whether the tool makes dubbing routine for more creators will depend on how useful the generated audio is and how well future improvements preserve the original voice and performance.