Microsoft’s new tools can turn a written script into a video of a person-like avatar speaking, or use a short speech sample to create a replica of a user’s voice. The features, announced at Microsoft Ignite 2023, offer uses such as training videos, virtual assistants and translated narration. They also put consent, disclosure and control of AI-generated media in focus.
From script to speaking avatar
Azure AI Speech text-to-speech avatar lets users upload images of someone they want an avatar to resemble and provide a script. Microsoft trains one model to animate the avatar, while a separate text-to-speech model reads the words aloud. That speech model can be prebuilt or trained on the person’s voice.
The tool was available in public preview at launch. Microsoft described possible applications including training videos, product introductions and customer testimonials. Avatars can speak multiple languages, and chatbot versions can use AI models such as OpenAI’s GPT-3.5 to answer questions beyond a prepared script.
Those capabilities could make some video and customer-service content easier to produce. They also make it possible to show a realistic likeness saying words that the person did not necessarily say. That gap between appearance and authorship is central to the ethical questions surrounding synthetic media.
Access depends on the type of avatar
Most Azure subscribers could access prebuilt avatars at launch. Custom avatars, based on a specific person, were a limited-access capability available through registration and restricted to certain use cases.
Microsoft later clarified that customers creating custom avatars must obtain explicit written permission and consent statements from the person whose likeness is used. The customer’s agreement with that person must cover how long the avatar can be used, the purposes for which it can be used and any content limits. Microsoft also requires customers to disclose that the avatars were created with AI and are AI-generated.
The safeguards address permission and labeling, but the article raised questions about compensation and notification. Those concerns echoed a sticking point in the recent SAG-AFTRA strike: the use of AI-generated digital likenesses. Studios ultimately agreed to pay actors for their AI-generated likenesses. Microsoft did not initially respond to questions about how companies using actors’ likenesses should compensate them.
A separate tool can recreate a voice
Personal voice, another feature introduced at Ignite, belongs to Microsoft’s custom neural voice service. It can replicate a user’s voice from a one-minute speech sample. Microsoft presented it for personalized voice assistants, dubbing content into different languages and producing narration for stories, audio books and podcasts.
Microsoft set conditions for access. Users cannot provide prerecorded speech; they must give explicit consent through a recorded statement. The company checks whether that statement matches separate, one-time-use training data before the customer can synthesize new speech. Access was gated behind registration.
Customers also must agree to use personal voice only in applications where the voice does not read user-generated or open-ended content. Microsoft says the voice model must stay within an application, and the output must not be publishable or shareable from it. Customers who meet limited-access eligibility criteria retain sole control over creating, accessing and using voice models and their output for dubbing in entertainment scenarios involving films, TV, video and audio.
Watermarks offer a clue, with a dependency
Microsoft later said it would automatically add watermarks to personal voices. The watermarks are intended to help identify synthesized speech and determine which voice it came from. That adds a way to trace generated audio, but using Microsoft’s watermark detection service in an app or platform requires the company’s approval.
The features therefore come with several layers of rules: permission for custom likenesses, disclosure that avatars are AI-generated, consent and sample checks for personal voice, and limits on where voice output can be used. The source also leaves open practical questions about compensation and how broadly watermark detection can be used. For organizations considering synthetic avatars or voices, the terms of access and the conditions attached to each use are part of the product’s implications.