Stable Diffusion XL, or SDXL, is Stability AI’s latest text-to-image system, and early results suggest a noticeable step forward from Stable Diffusion 2.1. The beta model produces more convincing hands and richer images, though its ability to generate text still needs work.
More detailed images from shorter prompts
Stability AI says SDXL brings improvements in image composition and face generation. The company also says users can create descriptive imagery without the long, detailed prompts its predecessor often needed. Tom Mason, Stability AI’s CTO, described the results as adding a “richness” that was missing from the earlier model, particularly for graphic design and architecture.
In a brief test, the new model appeared comparable to or better than the latest Midjourney model. A prompt as short as “Balenciaga Pope” produced runway models in designer-like clothing. The older Stable Diffusion instead generated more plainly religious-looking apparel.
Hands have long been a conspicuous weakness in text-to-image tools. SDXL’s results are not always realistic, but they are a marked improvement over the malformed hands that its predecessor often generated. Text inside images remains less reliable: brief testing found that SDXL still had some way to go on that task.
More ways to work with images
SDXL is designed to do more than turn text prompts into images. It also supports image-to-image prompting, where an input image guides variations; inpainting, which reconstructs missing parts; and outpainting, which extends an existing image.
These features let users start with an image and modify or expand it, rather than relying only on a written description. Together with the reported improvements to composition and faces, they give the model several routes to produce or revise visual material.
The system is available in beta through DreamStudio, Stability AI’s generative art tool, and through the company’s API in early access. Stability AI says SDXL will be open sourced after it exits beta, following earlier Stable Diffusion releases.
Progress comes with unresolved concerns
SDXL arrives as generative image tools face scrutiny over how their training data was collected and how the models are used. A legal case alleges that Stability AI infringed artists’ rights by building its tools with web-scraped copyrighted images. Getty Images has also taken the company to court over reportedly using images from its site without permission to create the original Stable Diffusion.
The open-source release of Stable Diffusion has drawn criticism over its relatively light usage restrictions. Some online communities have used it to create pornographic celebrity deepfakes and graphic depictions of violence. At least one U.S. lawmaker has called for regulation to address models that “don’t sufficiently moderate content.”
Stability AI has pledged to respect artists’ requests to remove their work from the Stable Diffusion training dataset. That pledge did not cover SDXL; it applied to the next-generation models code-named “Stable Diffusion 3.0.” Spawning, the organization leading the opt-out effort, says artists have removed more than 78 million works of art from the training dataset to date.
Business pressure remains part of the picture
Stability AI is also under pressure to make money from a broad range of AI projects spanning art, animation, biomed and generative audio. CEO Emad Mostaque has hinted at plans to IPO. Semafor recently reported that the company, which raised over $100 million in venture capital last October at a reported valuation of more than $1 billion, “is burning through cash and has been slow to generate revenue.”
Those pressures form part of the setting for SDXL’s release. The model offers visible improvements in image generation and adds ways to edit images, while its weak text rendering and the disputes around training data, moderation and commercialization remain unresolved.