Stability AI has introduced Stable Video Diffusion, a model family that turns existing images into short videos. The release gives researchers access to one of the few video-generation models offered in open source, but it arrives as a research preview with usage conditions and technical limits.
Two models, with different clip lengths
Stable Video Diffusion builds on Stability AI’s Stable Diffusion text-to-image model. It comes in two versions: SVD and SVD-XT. Both create videos at 576×1024 resolution, with SVD generating 14 frames and SVD-XT generating 24.
The models can produce video at between three and 30 frames per second. The source describes their output as roughly four-second clips. SVD-XT uses the same architecture as SVD, but generates more frames, offering a longer sequence of images to animate.
Access is not open without conditions. Stability AI calls the release a “research preview,” and people who want to run the models must agree to terms describing intended uses such as educational or creative tools and artistic processes. The terms also identify factual or true representations of people or events as a non-intended use.
Useful capabilities, visible limitations
The model’s central task is image-to-video: users supply a still image, and the system creates motion around it. That differs from text-to-video generation, which Stable Video Diffusion does not yet support. Stability AI says it plans a web tool that will add text prompting later.
The company also lists several limitations. The models cannot generate video without motion or slow camera pans, cannot be controlled with text, and do not render text legibly. They also do not consistently generate faces and people “properly.” These constraints matter for anyone considering the system for scenes that need precise direction or reliable human likenesses.
Stability AI says the models can be extended for other uses, including creating 360-degree views of objects. That suggests the initial release may serve as a base for further experiments, though the article does not describe those applications as ready-made features.
Training data and safeguards raise questions
A whitepaper released with the models says they were first trained on a dataset of millions of videos, then fine-tuned on a smaller collection ranging from hundreds of thousands to around a million clips. The precise origins of those videos are unclear. The paper suggests many came from public research datasets, but does not settle whether any were copyrighted.
That uncertainty leaves open legal and ethical questions about usage rights for the training material and for people who use the models. The preview’s restrictions set expectations for use, but the source says the model does not appear to have a built-in content filter. That absence raises concerns about how generated video could be misused if the model circulates beyond its intended audience.
The concern is informed by earlier image-generation tools: after Stable Diffusion was released, people used it to create nonconsensual deepfake pornography and other harmful material. The source raises this as a risk to consider for Stable Video Diffusion, rather than reporting that the new model has already been used in the same way.
A product launch amid business pressure
Stability AI says it intends to develop more models that build on SVD and SVD-XT. It has also pointed to possible applications in advertising, education and entertainment, indicating a commercial direction for the technology.
The company’s financial position forms part of that context. Stability AI recently raised $25 million through a convertible note, bringing its total raised to over $125 million. It was last valued at $1 billion, while the source reports that it had not closed new funding at a higher valuation and was said to be seeking quadruple that amid low revenues and a high burn rate.
The article also reports concerns about cash use, delayed or unpaid wages and payroll taxes, and AWS threatening to revoke access to GPU instances used to train models. Ed Newton-Rex, Stability AI’s former VP of audio, left after a disagreement about copyright and the use of copyrighted data in AI training. Those pressures make Stable Video Diffusion both a technical launch and part of a broader effort to build products that can support the company.