DeepFloyd IF aims to make open text-to-image generation better at understanding prompts and producing readable text within images. The model follows an approach similar to Google's Imagen and combines a text encoder with diffusion models that generate and then enlarge an image.
Language understanding is central to the design
Google demonstrated Imagen in May 2022. The source article says the model outperformed OpenAI's DALL-E 2 in accuracy and image quality in the team's examples, and could generate text inside images—a capability open-source models had not handled reliably.
Imagen uses the T5-XXL language model to encode a prompt before a diffusion model turns that representation into an image. This differs from approaches that use the multimodally trained CLIP encoder. The Google team also reported that increasing the language model's size improved image quality more than additional training of the diffusion model.
DeepFloyd IF replicates this general architecture as an open-source model. According to its team, IF brings strong image quality and language understanding, with text handling as a notable capability.
How IF generates and enlarges images
The model was trained on around 1.2 Billion images from the LAION-5B dataset. It uses a staged process: after image generation, two super-resolution models increase the output resolution to 1,024 x 1,024 pixels.
Different model sizes are available, up to 4.3 billion parameters. The largest model paired with the 1,024-pixel upscaler calls for 24 gigabytes of VRAM, according to the team's recommendation. The largest option with a 256-pixel upscaler still requires 16 gigabytes of VRAM.
Beyond creating an image from a text prompt, the team says IF supports Image-to-Image-Translation and Impainting. These capabilities let users work from an existing image or modify part of one, extending the model's use beyond prompt-only generation.
Reported results and what they suggest
In tests on the COCO dataset, IF reached a Zero-Shot FID score of 6.66. The source reports that this result was ahead of Google Imagen and available models such as Stable Diffusion.
DeepFloyd describes the work as evidence for the potential of larger UNet architectures in the first stage of cascaded diffusion models. In plain terms, the design puts substantial capacity into an early image-generation stage, followed by models that increase resolution.
The reported benchmark and capabilities point to an open-source effort aiming to approach leading text-to-image systems. They do not, by themselves, establish how the model will perform across every prompt or use case.
Access comes with an initial license limit
The first IF release has a restricted license intended for research and other non-commercial purposes while the team gathers feedback. DeepFloyd and StabilityAI said they would release a completely free commercially compatible version after that feedback period.
The model has a Github page, a demo on HuggingFace, and further information on the DeepFloyd website. For anyone assessing the model, the release terms matter alongside its prompt understanding, image quality, hardware requirements, and reported benchmark results.