How an inaudible watermark could help identify AI voices

Resemble AI’s PerTh system is designed to embed identifying data in AI-generated speech, where listeners are unlikely to hear it. The proposal aims to make that data survive common audio changes, though it currently identifies only speech generated by Resemble.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 0 ►

The watermark is a mildly Terminator-leaning tool for tracing AI-generated voices and addressing potential misuse.

How an inaudible watermark could help identify AI voices

AI-generated speech can serve useful purposes, including screen readers and replacing voice actors with their permission. The same ability to imitate a voice can also be misused to create fake quotes. Resemble AI has proposed an inaudible watermark that could help identify speech generated by its system, even after common audio changes.

Why mark generated speech?

A watermark is a pattern embedded in media to indicate where it came from. Some watermarks are visible, such as a logo on an image. Others are hidden: a computer can detect the pattern even when a person sees or hears no obvious difference.

That distinction matters for generated voices. A verification method based only on close listening or a statement from a publicist would be difficult to apply consistently. An embedded signal could provide another way to check whether a clip came from a particular speech-generation service.

For companies that build voice models, verification may also help address the risks of misuse. Resemble AI and other startups use speech models to create things such as dubs and audiobooks. If realistic generated voices are used deceptively, the companies behind them could face reputational damage and potential liability.

How PerTh is designed to work

Resemble calls its proposed method PerTh, a name combining “perceptual” and “threshold.” The company says machine learning models embed packets of data into generated speech and can later recover them. That data is intended to show whether a clip was generated by Resemble.

The approach uses a feature of human hearing: a louder tone can mask a nearby, quieter one. The article describes an example in which laughter produces peaks at 5,000 Hz, 8,000 Hz, and 9,200 Hz. Structured tones placed at nearby frequencies could be hard for a listener to notice while still carrying identifying information.

Keeping those tones close to important parts of the audio is also intended to make them harder to remove. The system would need to find suitable sections of a waveform, add the signal, and then detect it later. The article says Resemble’s method is designed to tolerate changes such as speeding up or slowing down audio and converting it to compressed formats like MP3.

The challenge of surviving audio changes

Hidden watermarks can be fragile. In an image, resizing may disrupt a pattern encoded at the pixel level. In audio, compression for streaming can remove quiet tones. A watermark that disappears during routine processing would be of limited use for checking clips that have been shared or edited.

PerTh’s proposed advantage is that it couples the identifying data closely to the speech information. Resemble says this makes the watermark difficult to remove while allowing it to persist through several common manipulations. The goal is a mark that is unobtrusive to people but recoverable by the company’s detection system.

In examples provided by Resemble, the article’s author could not distinguish the watermarked clip by listening or by inspecting its waveform for obvious anomalies. That observation suggests the signal was not readily apparent in those examples, but it does not establish how the method performs across every recording or edit.

What the proposal can—and cannot—identify

At the time described in the article, Resemble planned to roll PerTh out to its customers. Its detection was limited to speech generated by Resemble itself. A mark from one provider therefore could not, by itself, identify AI-generated speech from every other system.

If other speech-generation companies develop similar methods, watermarks could become more closely tied to the models that produce voices. That would give people and services another clue about a clip’s origin. It would not guarantee that deceptive uses disappear: the article acknowledges that malicious actors may find ways around safeguards.

The proposal also applies specifically to audio. The article says similar techniques will not work for text or images, so identifying generated content in those formats remains a separate challenge. For voice clips, PerTh offers a possible verification layer, with its usefulness depending on whether the signal can be recovered after real-world handling and whether other providers adopt comparable systems.