Diffusion models now lead image generation, but GigaGAN suggests that generative adversarial networks still have room to develop. The system can create images from text prompts and, in the reported comparison, does so substantially faster than Stable Diffusion. Its current image quality still trails high-quality diffusion models, leaving speed and editing control as its clearest strengths.
A GAN scaled for text-to-image generation
Generative adversarial networks, or GANs, were widely used for image generation before diffusion models surpassed them in image quality in 2021. GANs have continued to offer two practical advantages: fast image synthesis and a structure that can make outputs easier to control.
Researchers from POSTECH, Carnegie Mellon University, and Adobe Research developed GigaGAN as a large-scale text-to-image GAN. The model has a billion parameters, making it six times larger than the largest GAN to date. The team trained it on LAION-2B, a dataset of over 2 billion image-text pairs, and COYO-700M. A separate GigaGAN-based upscaler was trained using Adobe Stock photos.
According to the paper, reaching this scale required architectural changes, including some ideas inspired by diffusion models. The result can generate images at 512 x 512 pixels from text descriptions. In the examples provided, the subjects are recognizable, although the results do not yet reach the quality of high-quality diffusion models.
Generation speed is GigaGAN’s clearest advantage
The reported timing comparison shows a large speed gap. On an Nvidia A100, GigaGAN generates an image in 0.13 seconds. Muse-3B takes 1.3 seconds, while Stable Diffusion (v.1.5) takes 2.9 seconds.
That difference could matter for tasks where people need to generate or revise many images quickly. The comparison positions GigaGAN as between 10 and 20 times faster than comparable diffusion models. It does not, by itself, establish that GigaGAN matches their image quality: the article describes its current output as clear but behind the strongest diffusion examples.
The researchers say further scaling could improve quality, and expect performance to improve with larger models. This is a prospect rather than a demonstrated result; the examples described show where the model stands now, not what a future version will achieve.
GAN structure can make edits more direct
GigaGAN also brings back editing capabilities that became challenging with the move to autoregressive and diffusion models, according to the researchers. Its GAN architecture allows changes such as altering an object’s material or shifting the time of day in an image.
Diffusion models can support similar edits, but the source says they often depend on external methods, tricks, or manual work to do so. A system with more direct controls could make those adjustments easier to handle. That advantage is about the editing process; it does not remove the current gap in image quality.
An upscaler points to another use
The GigaGAN upscaling variant takes a 128-pixel image and produces a high-resolution 4K image in 3.66 seconds. In the examples shown, the added details appear photo-realistic. This gives the work a second application beyond creating images from text: increasing the resolution and detail of an existing image.
There are no plans so far to release the models. The article notes that a variant of the upscaler could potentially be integrated into Adobe Firefly or Photoshop, but presents this only as a possibility. For now, GigaGAN is evidence that GANs may still offer a useful combination of speed and editing control, while its image quality and availability remain open questions.