How Voicemod is turning AI voice tools into a creator platform

Voicemod raised $14.5 million to develop real-time AI voice tools and expand its creator platform. Its plans include singing voice conversion and audio watermarking, while questions about misuse, identity and voice licensing remain.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story mildly leans toward Terminator because real-time voice generation raises identity and misuse concerns, though it also describes creator tools and audio watermarking.

How Voicemod is turning AI voice tools into a creator platform

Voicemod is using a new funding round to expand from gaming voice effects into AI tools for speech and music. The company says it wants to help people create and transform voices in real time, while also developing ways for platforms to identify synthetic audio.

From voice effects to generated audio

Voicemod began with digital signal processing tools that let gamers add effects to their voices during chats. Its main audience remains gamers, but the company is pursuing wider uses for audio tools as generative AI advances.

Traditional effects modify a person’s voice. AI can also generate a voice or change how someone sounds while they speak. That opens possibilities for creators, from playful voice changes in social content to tools for games, apps and other digital spaces.

Music is another focus. Voicemod launched a text-to-song feature in December, allowing users to turn their lyrics into a vocal composition with generative AI. The company says it is also developing sing-to-sing conversion, which could let a person perform a song in another voice.

Combining tools and reaching more creators

The company’s plans draw on its acquisition of audio startup Voctro Labs last year. Voicemod CEO and co-founder Jamie Bosch says the teams are combining Voctro’s singing technology with Voicemod’s speech tools and digital signal processing. The goal is a hybrid platform that could support real-time speech and singing effects.

That combination could also bring together singing conversion and effects such as autotune. Bosch described potential uses for artists and for people creating voices for films or game characters. Those are possible directions for the technology; the company’s current work still involves sound engineers and designers when it needs to synthesize a whole voice.

Voicemod sells tools directly to consumers and creators, and also offers software development kits and APIs. Those options let third parties add its audio technology to products such as games, apps and hardware. Its business therefore reaches beyond a standalone desktop voice changer.

Funding growth and new platforms

Voicemod raised $14.5 million in expansion funding after an $8M Series A in summer 2020. The 2014-founded company says its desktop product has had more than 40 million downloads to-date, and Bosch says it has 3.3 million monthly active users. The company has generated revenue for years through paid versions of its tools.

Madrid-based Kfund’s growth fund Leadwind led the round, with participation from Minifund and Bitkraft Ventures. Voicemod says it will use the funds to develop real-time AI voice identity capabilities and creation tools for gamers, Gen Z, content creators and professionals.

The next planned expansion is platform access. Voicemod said it expected to bring its desktop product to macOS next month, extending availability beyond PC. Bosch also described a mobile creation app targeted for the beginning of next quarter, as the company works toward a product that spans devices.

Identity, misuse and the role of people

AI voice tools raise questions alongside their creative potential. A synthesized voice could be used to impersonate someone, and voice-based systems used for identity checks could face new challenges if voices are easy to change. Voicemod says it is working on audio watermarking that could help platforms tell whether audio was created with a synthetic voice.

Bosch said the company expects platforms to handle moderation in their own spaces, while Voicemod provides a way to identify synthetic audio. Watermarking is intended as one tool for managing misuse, including scams, fraud, manipulation, abuse, bullying and trolling.

Questions about training data and intellectual property also remain unsettled in the source article’s account. Bosch says Voicemod uses paid voice actors to build datasets for its models, and that when it wants to use original content, the team approaches the rights holder to discuss licensing.

The company also expects human performance to remain important. Bosch argued that a convincing voice depends on cadence, rhythm, tone and expression, not only on the identity of the voice being imitated. For now, Voicemod presents AI as a tool for people to create with, while acknowledging that the technology’s direction is still unfolding.