Written descriptions can now serve as the starting point for custom music and sound effects in Stability AI’s Stable Audio. The system turns prompts about genre, instruments and other characteristics into audio, giving creative professionals a way to produce background material through a web interface.
From descriptions to audio
Stable Audio uses a diffusion-based model to generate songs, sound effects and instrument stems. A user describes what they want, including details such as genre, instruments and tuning, and the system composes an audio file to match.
Stability AI’s examples show how detailed a prompt can be. One test requested “Post-Rock, Guitars, Drum Kit, Bass, Strings, Euphoric, Up-Lifting, Moody, Flowing, Raw, Epic, Sentimental, 125 BPM”. The resulting piece was described as a fast, atmospheric rock song. The company says the system can also create music in genres such as ambient, techno and trance.
Other demonstrations used simpler scenes or musical directions, including “People talking in a crowded restaurant” and “Piano solo chord progression in major, upbeat 90 BPM”. These examples show the range of material the tool is intended to produce, from environmental sound to a defined musical arrangement.
Longer compositions and the model behind them
Stable Audio can generate tracks of up to 90 seconds at 44.1 kHz. The article describes its samples as musically coherent over longer stretches, and says they sound authentic. These are notable qualities for a text-to-audio tool, especially when the intended use is a piece of background music rather than a brief sound.
Stability AI says generation can be fast: on an Nvidia A100 GPU, 95 seconds of audio should take less than a second to generate. That is the company’s stated performance, and the article notes that people trying the service may encounter heavy server demand.
The system combines several components. A Variational Autoencoder compresses stereo audio into a lossy, noise-resistant and invertible representation. A text encoder processes prompts, while a U-net-based diffusion model generates the audio. Timing embeddings, calculated during training, help control the output length.
The diffusion model is a 907 million parameter U-net based on the Moûsai model. The broader idea of using diffusion for music is not new, but Stable Audio’s ability to produce pieces of varying lengths was a focus of its training. Its basis includes Dance Diffusion, a text-to-music model released by Harmonai in 2022 with support from Stability; Stability says Stable Audio itself was developed from scratch by its audio division.
Training data and creator payments
Stable Audio was trained on a music library provided by AudioSparx. The partnership covers approximately 800,000 songs, audio effects and instrument snippets. AudioSparx receives a cut of Stable Audio’s revenue, and creators whose work was included can share in profits through the company.
The article says creators were allegedly asked whether they wanted to make their songs available before training. That arrangement speaks to one of the central questions around generative audio: how the works used to train a system relate to the revenue it earns. Stability AI has faced opposition over training material in the copyright debate around Stable Diffusion, its image model.
Access, licensing and open questions
Stable Audio is available through a web interface and is not open source. The article reports that personal use is free, with 20 songs per month of up to 45 seconds. A subscription costing $11.99 per month includes 500 songs of up to 90 seconds and a commercial license. Stability AI says it is aiming at creative professionals, including filmmakers and game developers who need background music quickly.
The company also says it plans to release an open-source music model trained on different datasets. That would give the project a separate route for people interested in an open model, though the article does not give a release date.
Questions remain about misuse. The article notes that Stable Audio could be used to imitate popular artists, and that the legal situation for AI-generated songs is unclear. Stability AI told TechCrunch it wants to use the technology responsibly. The AudioSparx library does not contain pop songs, according to the article, but includes tracks labeled as being in the style of well-known artists. Unlike Google’s MusicLM, Stable Audio did not block famous artists’ names at the time described.
For now, Stable Audio combines prompt-based creation, longer audio outputs and a revenue arrangement involving contributors to its training library. Its appeal to working creators will depend on how well it fits their needs, while questions about artist imitation and rights remain unresolved.