Microsoft’s AI Avatars Raise Questions About Consent and Control

Microsoft introduced Azure AI Speech text-to-speech avatar and a voice replication feature at Ignite 2023. The tools can generate speech and video from scripts or voice samples, with access limits and consent requirements that leave questions about likenesses, disclosure and watermark detection.

WTF Index TERMINATOR
◄ Terminator 3 Idiocracy 1 ►

The tools raise meaningful concerns about consent, likeness control and deceptive synthetic media, though access limits and consent rules temper the risk.

Microsoft’s AI Avatars Raise Questions About Consent and Control

Microsoft’s new tools can turn a written script into a video of a person-like avatar speaking, or use a short speech sample to create a replica of a user’s voice. The features, announced at Microsoft Ignite 2023, offer uses such as training videos, virtual assistants and translated narration. They also put consent, disclosure and control of AI-generated media in focus.

From script to speaking avatar

Azure AI Speech text-to-speech avatar lets users upload images of someone they want an avatar to resemble and provide a script. Microsoft trains one model to animate the avatar, while a separate text-to-speech model reads the words aloud. That speech model can be prebuilt or trained on the person’s voice.

The tool was available in public preview at launch. Microsoft described possible applications including training videos, product introductions and customer testimonials. Avatars can speak multiple languages, and chatbot versions can use AI models such as OpenAI’s GPT-3.5 to answer questions beyond a prepared script.

Those capabilities could make some video and customer-service content easier to produce. They also make it possible to show a realistic likeness saying words that the person did not necessarily say. That gap between appearance and authorship is central to the ethical questions surrounding synthetic media.

Access depends on the type of avatar

Most Azure subscribers could access prebuilt avatars at launch. Custom avatars, based on a specific person, were a limited-access capability available through registration and restricted to certain use cases.

Microsoft later clarified that customers creating custom avatars must obtain explicit written permission and consent statements from the person whose likeness is used. The customer’s agreement with that person must cover how long the avatar can be used, the purposes for which it can be used and any content limits. Microsoft also requires customers to disclose that the avatars were created with AI and are AI-generated.

The safeguards address permission and labeling, but the article raised questions about compensation and notification. Those concerns echoed a sticking point in the recent SAG-AFTRA strike: the use of AI-generated digital likenesses. Studios ultimately agreed to pay actors for their AI-generated likenesses. Microsoft did not initially respond to questions about how companies using actors’ likenesses should compensate them.

A separate tool can recreate a voice

Personal voice, another feature introduced at Ignite, belongs to Microsoft’s custom neural voice service. It can replicate a user’s voice from a one-minute speech sample. Microsoft presented it for personalized voice assistants, dubbing content into different languages and producing narration for stories, audio books and podcasts.

Microsoft set conditions for access. Users cannot provide prerecorded speech; they must give explicit consent through a recorded statement. The company checks whether that statement matches separate, one-time-use training data before the customer can synthesize new speech. Access was gated behind registration.

Customers also must agree to use personal voice only in applications where the voice does not read user-generated or open-ended content. Microsoft says the voice model must stay within an application, and the output must not be publishable or shareable from it. Customers who meet limited-access eligibility criteria retain sole control over creating, accessing and using voice models and their output for dubbing in entertainment scenarios involving films, TV, video and audio.

Watermarks offer a clue, with a dependency

Microsoft later said it would automatically add watermarks to personal voices. The watermarks are intended to help identify synthesized speech and determine which voice it came from. That adds a way to trace generated audio, but using Microsoft’s watermark detection service in an app or platform requires the company’s approval.

The features therefore come with several layers of rules: permission for custom likenesses, disclosure that avatars are AI-generated, consent and sample checks for personal voice, and limits on where voice output can be used. The source also leaves open practical questions about compensation and how broadly watermark detection can be used. For organizations considering synthetic avatars or voices, the terms of access and the conditions attached to each use are part of the product’s implications.