For people who may lose the ability to speak, a synthetic voice that sounds like their own can help preserve something personal. Acapela Group’s “my own voice” service makes recording a voice profile a short process, with no charge for creating and storing it.
From hours of training to a short recording session
Acapela has worked in text-to-speech for around 25 years. The company was recently acquired by Tobii Dynavox, though it continues to operate independently.
Acapela co-founder Remy Cadic described how much the process has changed. Seven or 8 years ago, customizing a synthetic voice could require 8 hours of patient training, and the results were not especially good. With neural text-to-speech techniques, the service can now create a voice profile from 50 sentences recorded in about 10 minutes. The voice is ready the next day.
The sentences come from a selection of novels, recipe books, and articles. In a demonstration, the recording interface was straightforward, and the resulting voice could read different sentences while remaining recognizable as the speaker.
Why keeping a personal voice matters
People who expect to lose speech because of degenerative conditions, cancer, or certain procedures may want a synthetic voice based on their own. A voice selected from a standard list may not feel like a suitable replacement, especially when someone has a voice they would prefer to keep using.
The service is designed around that need. It is not presented as a way to clone a celebrity or as a showcase for voice technology. It gives people a way to prepare a voice profile while they can still record themselves, making the process less demanding than earlier approaches.
Acapela also says it can train a voice using recordings from someone who has already lost the ability to read and speak. That route is possible, though the article describes it as less simple than recording the 50 sentences directly.
Access, devices, and customization
Creating and banking a voice is free. A user pays only if they want to download and install the voice for use on a device. The profile can be used with compatible speech-generation systems, including Tobii Dynavox’s TD Talk and devices.
Acapela says compatibility with offline devices is an important feature. Some online services make voice creation easy but provide access only through the cloud, which Cadic said is not practical for every use. Being able to use a voice on compatible devices can matter when the speech system needs to work without the latest neural processing chip.
The company also works on synthetic voices for children. Cadic said Acapela has made the recording script easier for children to read and adjusted the system to improve the quality of children’s synthetic voices. The article also notes that recording and re-recording a voice, or artificially aging a banked voice, is a newer and challenging capability.
Training quality and representation matter
Making a voice profile quickly is only useful if the result resembles the speaker. Cadic warned that some fast training methods select a speaker in the training material who seems closest to the user. If the material lacks a sufficiently similar voice, the synthetic result may not sound like the original.
Acapela product manager Nicolas Mazars said this issue can affect people unevenly. He described approaches that work well for an “average 50-year-old white guy” but may work less well for an African-American man or someone who does not speak English well.
Acapela says it works in 23 languages and uses feedback from people with disabilities to guide development. That focus matters because voice banking is intended for people with varied voices and needs. A quick recording session is a starting point; the training material and the people it represents also shape whether the resulting voice feels personal.