How Stability AI’s Model Turns Still Images Into Video

Stable Video Diffusion can animate a still image into a short clip, but results are limited and generation can take time on a local GPU. Stability AI describes it as a research preview, not a tool for commercial or real-world use.

WTF Index NEUTRAL
◄ Terminator 0 Idiocracy 1 ►

This limited research preview offers a modest creative capability without a clear lean toward harm or human decline.

How Stability AI’s Model Turns Still Images Into Video

Stable Video Diffusion gives researchers and hobbyists a way to turn a still image into a short video clip. The free, open-weights preview comes in two versions, but its output is brief, its motion can be modest, and Stability AI says it is intended for research rather than real-world or commercial applications.

Two models, short clips

The release includes SVD, which generates 14 frames, and SVD-XT, which generates 25. Both work from an image and can generate at different speeds, from 3 to 30 frames per second. The resulting MP4 clips are typically 2-4 seconds long and have a resolution of 576×1024.

The model can run locally on a machine with an Nvidia GPU. That makes it possible to experiment on personal hardware, though the time required can be a practical constraint. In Ars Technica’s local test, generating 14 frames took about 30 minutes on an Nvidia RTX 3060 graphics card.

Cloud services such as Hugging Face and Replicate also offer ways to try the models, though some may charge for access. For people without suitable local hardware, those services provide another route to experimentation.

Motion can be subtle or uneven

Image-to-video generation does not necessarily make every part of a picture move. In Ars Technica’s experiments, much of a scene often stayed still while the model introduced panning or zooming, or animated elements such as smoke and fire. People in photos often remained motionless, though one image of Steve Wozniak showed slight movement.

That makes the model better understood as an early image animation tool than a way to produce a fully moving scene on demand. A clip may add movement that changes how a still image feels, but the results can vary, and the source article describes them as mixed.

The limits matter for anyone judging what the system can do. A short clip with a camera-like shift or movement in one part of the frame is a different result from a scene where people and objects behave naturally throughout.

A research preview with open weights

Stability AI presents Stable Video Diffusion as an early research release. Its website says the model is not intended for real-world or commercial applications at this stage, and asks users for feedback on safety and quality as the company works toward a later release.

The model weights and source are available on GitHub. The Pinokio platform offers another way to run it locally, handling installation dependencies and keeping the model in its own environment. Together, these options make the preview accessible for hands-on exploration, while leaving its intended use and limitations clear.

Training data and the next step

The research paper describes a large video dataset comprising roughly 600 million samples. The team curated this into the Large Video Dataset, or LVD, which consists of 580 million annotated video clips spanning 212 years of content in duration. The paper does not identify the sources of the training data.

Stable Video Diffusion is one of several efforts to generate video with AI. Other systems have come from Meta, Google, Adobe, ModelScope, Runway’s Gen-2 model, and Pika Labs. Stability AI says it is also working on a text-to-video model, which would create short clips from written prompts instead of starting with an image.

For now, the release offers a way to explore image-to-video generation, with a clear tradeoff: it can add movement to a still picture, but clips are short, motion may be limited, and the company positions the technology as research-only.