A Free Video Model Brings Text-to-Video to More GPUs

Zeroscope offers a free text-to-video model with a smaller version for exploring ideas and a larger one for upscaling clips. Its hardware requirements make the smaller model usable on many standard graphics cards, while video generation still faces limits in quality, length, and computing cost.

WTF Index IDIOCRACY
◄ Terminator 0 Idiocracy 1 ►

Zeroscope makes AI video generation easier to try, with no clear sign of harm or eroded human skills.

A Free Video Model Brings Text-to-Video to More GPUs

Zeroscope makes text-to-video generation available as free software, with two model versions for different stages of creating a clip. The smaller version is meant for quick experiments, while a larger model can upscale selected results. That setup lowers one barrier to trying video generation, though the technology still demands substantial graphics memory and produces clips with visible limitations.

Two models serve different stages

The project builds on Modelscope, a text-to-video diffusion model with 1.7 billion parameters. Zeroscope refines that approach with higher resolution, a frame closer to a 16:9 aspect ratio, and videos without the Shutterstock watermark.

Zeroscope_v2 567w is designed for rapid content creation at 576x320 pixels. It gives users a lower-resolution way to explore video concepts before choosing results to improve. Those selected clips can then be upscaled with zeroscope_v2 XL to 1024x576 pixels, described as a high-definition resolution.

The two-step workflow separates experimentation from improving the output. Users can explore ideas with the smaller model and reserve the larger version for clips they want to upscale. The demo video’s music was added in post-production, so the model itself is being shown as a visual generator.

Graphics memory sets the practical limit

At 30 frames per second, generation at 576x320 pixels requires 7.9 GB of VRam. The 1024x576 version requires 15.3 GB of VRam at the same frame rate. The article says the smaller model should run on many standard graphics cards, giving more people a chance to try text-to-video locally.

Those requirements also show why resolution is a meaningful choice. A user can begin with the smaller output to test a prompt, then use the larger model to upscale a promising result. The higher-resolution option needs more than twice the graphics memory, so access to the XL model depends on having enough VRam available.

Zeroscope developer Cerspense, who has experience with Modelscope, said fine-tuning a model with 24 GB of VRam is “not super hard.” He removed Modelscope watermarks during the fine-tuning process. The developer described Zeroscope as “designed to take on Gen-2,” Runway ML’s commercial text-to-video model, and said Zeroscope is completely free for public use.

Training aimed to broaden the results

Zeroscope’s training involved offset noise applied to 9,923 clips and 29,769 tagged frames, each comprising 24 frames. The source explains that introducing noise during training can help the model understand the data distribution. It says this can support a more diverse range of realistic videos and help the model interpret variations in text descriptions.

In practical terms, the training approach is intended to help the model respond across different prompts rather than reproduce one narrow visual pattern. That does not remove the broader challenges of generated video. The article notes that text-to-video is still in its infancy: clips are typically only a few seconds long and often contain visual flaws.

Open access arrives amid a resource-heavy field

The article compares the trajectory of video generation with text-to-image systems, which also faced problems before achieving photorealism within months. But video requires far more resources for both training and generation. Progress in one area therefore does not guarantee that video models will advance at the same speed.

Several other systems were presented but had not been released: Google’s Phenaki and Imagen Video, and Meta’s Make-a-Video. At the time described in the source, Runway’s Gen-2 was commercially available, while Zeroscope was presented as the first high-quality open-source model in the field.

Both 567w and Zeroscope v2 XL can be downloaded for free from Hugging Face, which provides usage instructions. A Colab version with a tutorial is also available. Together, those options make the model accessible to people who want to experiment, while the memory requirements and short, imperfect clips show how much work remains for text-to-video tools.