Stability AI’s latest image generator is designed to make text-to-image creation more capable and easier to steer. Stable Diffusion XL 1.0 adds improved image quality, more flexible editing and a beta fine-tuning feature, while leaving open difficult questions about how such tools are used and what data they learn from.
A faster route to detailed images
The company describes Stable Diffusion XL 1.0 as its most advanced release to date. It says the model produces more vibrant and accurate colors, with improvements to contrast, shadows and lighting compared with its predecessor, Stable Diffusion XL 0.9.
Joe Penna, Stability AI’s head of applied machine learning, said the model has 3.5 billion parameters and can produce full 1-megapixel images in seconds, across multiple aspect ratios. Parameters are parts of a model learned from training data that help define what it can do. The previous version could also make higher-resolution images, but needed more computational power.
The release is available as open source on GitHub and through Stability AI’s API and consumer apps, ClipDrop and DreamStudio. That gives developers and users several ways to access the same model, including opportunities to customize it and fine-tune it for particular concepts or styles.
More ways to guide and edit results
Stable Diffusion XL 1.0 is intended to respond to shorter, more natural prompts, including instructions with several parts. Earlier Stable Diffusion models often needed longer prompts to convey the same kind of request. The change could make image generation easier to direct for people who do not want to build elaborate descriptions.
The model also adds capabilities for working with existing images. Inpainting can reconstruct missing sections, while outpainting can extend an image beyond its current edges. With image-to-image prompts, a user can provide a picture and text to generate a more detailed variation.
Text within generated images is another area Stability AI says has improved. Many text-to-image systems struggle to produce legible logos, lettering and calligraphy. Penna said this release supports advanced text generation and improved legibility, though the article does not describe how consistently it performs across different designs.
Access comes with safety and consent questions
Making the model open source also means it could be used to create harmful material, including nonconsensual deepfakes. The source article connects this risk partly to the training set, which contains millions of images from around the web. It also notes that online tutorials show how Stability AI tools can be used to make deepfakes and how base models can be fine-tuned to generate pornography.
Penna acknowledged the possibility of abuse and the presence of biases. Stability AI says it has filtered training data for unsafe imagery, added warnings about problematic prompts and blocked as many problematic terms as possible. Penna said the company plans to keep improving those safety measures.
There is also an unresolved dispute over artists’ work in training data. The article says the dataset includes artwork by artists who have protested its use by AI companies. Stability AI argues that fair use doctrine shields it from legal liability in the U.S., while artists and Getty Images have filed lawsuits to stop the practice.
Stability AI has partnered with Spawning to respect artists’ opt-out requests. The company says it has not removed all flagged artwork from its datasets, but continues to incorporate artists’ requests. That leaves a practical tension: the company describes a process for honoring removal requests while acknowledging that the process is not complete.
New product plans and a competitive market
Alongside the release, Stability AI is introducing a beta API fine-tuning feature. It is designed to let users specialize image generation for people, products and other subjects using as few as five images. The model is also coming to Amazon Bedrock, expanding the company’s collaboration with AWS.
These moves arrive as Stability AI faces competition from OpenAI, Midjourney and others. The article reports that the company had raised over $100 million in venture capital, was burning through cash, closed a $25 million convertible note in June and was seeking an executive to help ramp up sales. Wider availability, new features and commercial partnerships may support its position, but the release also makes its approach to safety and artists’ requests part of the product’s public test.