Google VideoPoet Brings Video, Audio and Image AI Into One Model

Google has introduced VideoPoet, a generative AI system built as a large language model for video creation and editing. It can handle text-to-video, image-to-video, stylization, inpainting, outpainting and video-to-audio tasks within one model.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

This is mostly a routine AI media-generation launch with only mild implications for synthetic content and creative dependence.

Google VideoPoet Brings Video, Audio and Image AI Into One Model

Google has unveiled VideoPoet, a generative AI system that uses a large language model approach to create and edit video from text and other inputs. The system is designed to handle several media-generation tasks in one framework, including video, image, audio and text.

What VideoPoet Is Built To Do

VideoPoet is described by Google as a large language model for a wide range of video generation tasks. Its capabilities include text-to-video, image-to-video, video stylization, video inpainting and outpainting, and video-to-audio.

The key difference highlighted by Google is that VideoPoet brings these functions into a single model. Instead of depending on separately trained components for each task, the system is built to work across multiple kinds of media inside one language-model framework.

That matters because video generation is not just about making frames appear from a prompt. A useful system may also need to animate an image, extend or repair a scene, change the look of a video, or produce sound that fits what is happening on screen. VideoPoet is presented as a step toward that broader kind of media generation.

How The Model Handles Media

VideoPoet uses multiple tokenizers to represent different media types. Google uses MAGVIT V2 for video and image, and SoundStream for audio. These tokenizers allow an autoregressive language model to train across video, image, audio and text modalities.

In plain terms, the model works with tokens rather than raw media alone. After VideoPoet generates tokens based on a given context, tokenizer decoders can convert those tokens back into a viewable or listenable output.

This approach lets the same overall model operate across different inputs and outputs. A text prompt can guide a video. An image can become the basis for animation. Video-related information can also be used for stylization or audio generation.

What It Can Generate

VideoPoet can create videos of variable length, with different motions and styles depending on the text content. Google also says the model can animate an input image when paired with a prompt.

The system can predict optical flow and depth information for video stylization. It can also generate audio, including examples where a video is paired with sound.

By default, VideoPoet creates videos in portrait orientation. Google frames that choice around short-form content, where vertical video is often the expected format.

Text prompts can also be used to guide camera movement. That means the prompt can describe not only the subject of a video, but also how the camera should move through or around the scene.

How Google Says It Compared

Google evaluated VideoPoet against several benchmarks and compared its generated videos with outputs from other models. The models named in the source include Phenaki, VideoCrafter and Show-1.

On average, participants preferred between 24 and 35% of VideoPoet's examples because they matched the prompt better than competing models. That evaluation focused on how closely the generated results followed the prompt.

Those results do not mean every VideoPoet output is preferred in every case. They do show that Google is positioning prompt alignment as one of the system's strengths.

A Step Toward Broader Generation

Google says the framework could support "any-to-any" generation in the future. The company also says it could be extended to text-to-audio, audio-to-video and video captioning, "among many others".

Google has also produced a short film using VideoPoet, with Bard used as the scriptwriter. The source does not say whether VideoPoet will be released publicly.

The company has not revealed whether it has plans to make the model available. For now, VideoPoet is best understood as a research system showing how a single large language model can be applied across video, image, audio and text generation tasks.