Innovation

Why Amazon Pairs AI Lip-Sync With Human Dubbing on Maxton Hall

By automating post-production synchronization of dubbed audio, Amazon aims to reduce localization costs and make smaller titles economically viable for translation.

By Hollywood Feature · September 30, 2026 · 8 min read
Why Amazon Pairs AI Lip-Sync With Human Dubbing on Maxton Hall

Amazon Prime Video began testing a different approach to dubbing in March 2025, then deployed it commercially on Maxton Hall, its most-watched international series to date, starting September 9, 2026. Rather than relying entirely on new voice recordings, the company uses artificial intelligence to adjust actors’ on-screen mouth movements to match professionally produced dubbed dialogue. This hybrid approach—preserving human voice performance while automating visual adjustments—is designed to cut the post-production labor involved in dubbing and lower the cost of localizing content for global audiences.

The technique addresses a persistent problem in dubbed entertainment: when translators render dialogue into another language, the words no longer match the actors’ lip movements. Studios traditionally hire dialogue adapters to rewrite lines to fit the original mouth positions, then record new voices, then synchronize everything by hand—a process that can span weeks and cost thousands to tens of thousands of dollars per language. Amazon’s system performs the visual adjustment computationally, analyzing and regenerating the actors’ mouth shapes for each dubbed version after humans have already recorded the dialogue.

How the dubbing industry works

Traditional dubbing follows a standardized workflow across studios worldwide. Pre-production begins with script translation and adaptation—a specialized role where translators render dialogue into the target language while a dialogue adapter rewrites lines to match the original actors’ mouth movements as closely as possible, a constraint that often requires sacrificing literal accuracy for lip-sync fidelity.

Production involves casting voice actors selected to match the original cast’s vocal qualities. Actors perform in soundproof studios, watching the original footage and timing their dialogue to match actors on screen. The goal is to find voices that sound convincingly like they’re coming from the original actors’ mouths—a significant constraint on casting.

Post-production requires mixing and editing to integrate the new dialogue with the original video. If dialogue falls out of sync, actors may need to re-record lines, extending the schedule. Studio time alone runs $100 to $400 per hour.

The dubbing cost baseline
Dubbing costs $20 to $60 per minute of finished audio, with broadcast-quality work starting at $50 to $60 per minute. Studio time alone runs $100 to $400 per hour. Full-length films or television series may instead be billed a flat project fee ranging from a few thousand to tens of thousands of dollars. These costs cover translation, dialogue adaptation, voice talent, direction, recording, mixing, and quality control.

The cost structure that limits dubbing to major titles

Traditional dubbing costs $20 to $60 per minute of finished audio, with broadcast-quality work for television or film starting at $50 to $60 per minute, and studio time running $100 to $400 per hour. Full-length films or television series may instead be billed a flat project fee ranging from a few thousand to tens of thousands of dollars, bundling translation, dialogue adaptation, voice talent (often charged per session or per character), a dubbing director, recording engineers, mixing and mastering, and quality control. For a global rollout across even a dozen languages, the economics favor only the largest productions and most commercially promising titles.

This cost structure explains why international streaming platforms strategically choose which content to dub and which to subtitle-only. Netflix, with more than 70% of its subscribers outside the United States, has invested $1 billion in content localization including both dubbing and subtitling in a single year. The company offers dubbing in more than 34 languages, but only for titles expected to justify the expense. “Money Heist,” dubbed into multiple languages, generated over 180 million viewing hours in the month after its release, validating that investment.

Smaller titles, independent productions, and lower-profile licensed content rarely get dubbed at all. Studios calculate that the cost to dub into, say, five languages would exceed projected revenue from those markets. This creates a silent divide in streaming: viewers of major-studio productions access dubbed versions; viewers of everything else rely on subtitles or foreign-language audio.

Visual dubbing vs. voice synthesis

Prime Video’s technology, known as visual dubbing or lip-sync adjustment, operates on dubbed audio that humans have already recorded. The system scans the original video frame-by-frame, identifying the actors’ mouth movements, then generates new mouth shapes to match the dubbed language’s phonemes and rhythm, addressing the disconnect that occurs when an actor’s mouth keeps moving after the translated dialogue has ended.

This differs fundamentally from voice synthesis dubbing, where artificial intelligence generates the voices themselves. Amazon’s approach keeps voice talent in the process—localization professionals record dialogue with human actors, then use AI to solve the synchronization problem. Amazon has described the hybrid approach to dubbing as one in which “localization professionals collaborate with AI to ensure quality control.” The technology addresses what professional translators have traditionally solved manually: matching dialogue duration and mouth position without losing meaning or performance.

From pilot to commercial deployment

Amazon launched its AI-aided dubbing pilot in March 2025 across 12 licensed titles in English and Latin American Spanish, including El Cid: La Leyenda, Mi Mamá Lora, and Long Lost. These titles had no existing dubbed versions, making them candidates for automation that studios would not have translated through traditional methods. The pilot represented a strategic choice: apply AI not to replace dubbing on major titles, but to enable dubbing on content that would otherwise never receive it.

The September 2026 deployment marked a different milestone. Visual dubbing appeared on Maxton Hall Seasons 1 and 2, available globally in English. Maxton Hall is Prime Video’s most-watched international series to date. This was not a pilot for unserved content but a commercial application to content that would have received traditional dubbing anyway—a test of whether Amazon could maintain quality while reducing post-production labor on its flagship international series.

When Season 3 releases December 9, 2026, it will be the first new content produced with this visual dubbing workflow, available across more than 240 countries and regions. This timing matters: Maxton Hall’s performance with AI lip-sync will determine whether competitors follow or dismiss the technology as insufficient for major productions.

Economics: automation of revision, not elimination of talent

Amazon’s approach does not eliminate the voice actors, directors, and engineers that drive traditional dubbing costs. Human voices still require professional actors and studio time, the largest line items in dubbing budgets. Instead, the technology reduces post-production revision cycles. Traditionally, after dialogue is recorded, engineers synchronize it to picture and flag timing mismatches. Dialogue adapters then re-record lines that fall out of sync, extending the schedule and requiring actors to return to studios for reshoots.

Visual dubbing automation bypasses that revision step for most content, compressing the post-production pipeline. A title that might have required additional rounds of recording and re-recording may now require fewer. This reduces studio time, engineer hours, and re-cast sessions—labor costs that don’t affect quality but do affect schedule and economics.

By targeting content “that would not have been dubbed otherwise,” the March 2025 pilot demonstrated Amazon’s intent to enable dubbing for titles where traditional economics didn’t justify translation. A 90-minute film that would cost a few thousand to tens of thousands of dollars to dub traditionally—money many independent productions cannot justify—becomes economically viable when post-production labor shrinks to hours rather than days. The economics shift from “only major productions warrant dubbing” to “dubbing becomes viable for mid-tier and independent content.” Amazon has not disclosed specific cost savings or timeline improvements from either the pilot or the Maxton Hall deployment.

Reducing the cost of each dubbed version might increase demand for dubbing across more titles, particularly lower-budget content where dubbing was previously unaffordable.

Implications for dubbing departments and voice careers

The technology primarily affects post-production roles rather than on-set talent. Dialogue adapters, dubbing directors, and audio engineers in post-production face potential workflow changes. If visual dubbing becomes standard, the labor-intensive revision step compresses, reducing billable hours for specialized roles like dialogue adaptation and post-production supervision.

Voice actors themselves remain central to the process—the system requires professionally recorded performances, and no automation changes the need to cast actors who sound convincing as the original characters. However, the technology could shift the total volume of voice acting work. Reducing the cost of each dubbed version might increase demand for dubbing across more titles, particularly lower-budget content where dubbing was previously unaffordable. Conversely, automated post-production could reduce the total number of personnel needed per project, offsetting volume gains.

Localization professionals also maintain oversight roles. Amazon’s pilot model retained human review of AI output before release, suggesting the company views quality control as irreplaceable. This suggests any scaling of the technology will require human review infrastructure, preserving some of the traditional quality-control roles in localization departments.

Strategic implications for global streaming competition

The deployment signals a shift in streaming economics across the industry. Netflix built its international dominance partly on dubbing—the platform now offers dubbed content in over 34 languages. Nearly half of non-English viewers prefer dubbed versions over subtitles, particularly for drama and action content. But Netflix’s high dubbing costs limit which titles receive translation.

Visual dubbing expands the economically viable dubbing catalog without proportional cost increases. If Amazon can scale the technology to more languages and titles, it could offer dubbed versions of lower-profile content that Netflix subtitles-only, strengthening Amazon’s position in regions where viewers strongly prefer dubbing. Netflix, Apple TV+, and other streamers must now determine whether to follow with similar technologies or maintain traditional dubbing budgets for major titles only.

For Amazon specifically, expanding dubbed content into more than 240 countries and regions without proportional cost increases improves margins on international content and potentially accelerates releases. If visual dubbing reduces post-production timelines, Amazon could release dubbed versions of international content simultaneously across multiple regions rather than staggering them over weeks—a competitive advantage for momentum and word-of-mouth.

What comes next

Amazon has committed to expanding beyond Maxton Hall but provided no timeline or scope. Unanswered questions shape industry strategy: Will the technology scale reliably to dozens of languages with vastly different phonetic systems? Can quality hold for genres beyond drama? Will voice acting unions and localization professionals accept the workflow change?

The rollout will likely depend on technical performance with Seasons 1, 2, and 3 of Maxton Hall. If viewers accept the visual adjustments without noticing discontinuity, Amazon will have validated a model that competitors can follow. Quality issues—visible artifacts, mismatched mouth movements, or inconsistent compositing—would limit the approach to specific content or languages where visual synchronization is less perceptually critical. Industry adoption will follow viewer reception and technical reliability, not just cost savings alone.

Photo: Jessicaperkins99 · CC BY-SA 4.0 · via Wikimedia Commons

Read next