AI Turns Project Gutenberg Books Into 5,000 Free Audiobooks

Microsoft and Project Gutenberg used AI to create more than 5,000 free audiobooks from digital books. The process combines book-structure parsing, speech synthesis and dialogue-aware narration, and the team says it plans to release the audio data as open source without restrictions.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The project uses AI to make books more accessible through free audiobooks, with no clear lean toward control or human skill erosion.

AI Turns Project Gutenberg Books Into 5,000 Free Audiobooks

More than 5,000 free audiobooks have been created through a project by Microsoft and Project Gutenberg. The team used AI to turn digital books into spoken audio, with systems designed to identify the text worth reading and produce natural-sounding narration.

From ebook structure to spoken words

Before a book can be read aloud, the system needs to separate its main text from the other material in an ebook. The researchers developed an algorithm to interpret the structure of HTML-based books and filter out elements such as footnotes, page numbers and tables.

This parsing step prepares the text for text-to-speech, or TTS. The project used WaveNet, Tacotron and FastSpeech to generate speech described as natural and human-like. Together, these stages turn a structured digital text into audio that listeners can follow.

That preparation matters because an ebook contains more than the sentences of its story. Page numbers and notes may be useful on screen, but they can interrupt a spoken version. By distinguishing them from the main text, the system can focus its narration on the material intended to be read aloud.

Narration that follows dialogue

The team also developed a system to tell narration apart from dialogue. It can distinguish between individual characters and their emotions, then adapt the generated voice accordingly.

That gives the audiobook process a way to handle changes in who is speaking, rather than treating every line as the same kind of text. The system combines this interpretation with speech synthesis, so the generated audio can reflect whether a passage is narration or dialogue.

The complete process runs on SynapseML, a machine learning framework designed to divide tasks and process them in parallel. In this project, that framework connects the stages involved in preparing and generating audiobooks.

Personal voices and a large audio collection

For a conference presentation, the team demonstrated a zero-shot text-to-speech approach. It can capture the character of a user's voice from a few recorded sentences and apply it to audiobook narration.

Users could choose a book from the digital library and hear it in their own voice, or use a voice of their choice if they had audio files. The article says it was unclear whether this service would continue beyond the conference, and suggests that potential costs could make wider availability unlikely.

Across the project, the team collected more than 35,000 hours of audio data covering classical literature, plays, biographies and more. The recordings were read in a clear and consistent voice. The team intends to make all of this audio data available as open source without restrictions, which could also make it useful for other AI projects.

Where to listen and what the project offers

The audiobooks are available on Spotify, Apple Podcasts and Google Podcasts. The project pairs a free digital library with a way to listen to its books, extending access beyond reading the ebooks on a screen.

Project Gutenberg is a volunteer-created digital library accessible through the internet. Its website offers more than 70,000 ebooks to read and download for free. The audiobook project adds spoken versions to that collection, while its planned open release of audio data could support further work with speech and literature.

The team says the effort has the potential to improve the accessibility and availability of audiobooks. Its approach addresses several parts of the task: finding the main text, generating speech, and making dialogue and character changes intelligible in audio.