Meta’s Emu Models Turn Text Prompts Into Videos and Edits

Meta introduced Emu Video, which creates short videos from text or image prompts, and Emu Edit, which changes images using written instructions. Both remain research projects, while the company reports strong results in human evaluations and image-editing tests.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The research tools make visual creation easier, but the article describes a routine technology update without a clear societal harm or dependency trend.

Meta’s Emu Models Turn Text Prompts Into Videos and Edits

Meta’s Emu research introduces two ways to work with visual media through text: Emu Video generates short clips, while Emu Edit applies image changes based on written instructions. Together, they point to a workflow where people can describe a visual result instead of assembling every change by hand.

Emu Video builds a clip in two stages

Emu Video builds on Meta’s Emu image model, introduced at Connect 2023. It accepts text prompts, image prompts, or a combination of the two. For text-led generation, it first creates an image from the description, then uses the text and that image to generate a video.

Meta says this staged process carries over the visual variety and style of its image model. The researchers describe it as a factorized approach, designed to make training more efficient while supporting direct generation of higher-resolution video.

The system uses two diffusion models to create 512x512 videos that run for four seconds at 16 frames per second. The team also experimented with clips up to eight seconds long and reported good results. Examples described in the research include an American flag waving during the moon landing as the camera pans, and a ship leaving a harbor.

Human ratings compare quality and prompt accuracy

To assess generated clips, the researchers developed JUICE, short for JUstify their choICE. People comparing videos are asked to explain their choices against set criteria, an approach intended to make the ratings more reliable.

The quality criteria include pixel sharpness, smooth motion, recognizable objects and scenes, image consistency, and range of motion. Prompt accuracy is evaluated through spatial and temporal text alignment: whether the clip reflects what the prompt describes in its visual content and over time.

In those human evaluations, the researchers say Emu Video was preferred to Pika Labs videos in quality and quantity in more than 95 percent of cases. Google’s Imagen Video came closer on prompt accuracy, but the reported figures were 56.4% for that measure and 81.8% for quality. These are results from the team’s evaluation; they describe comparisons under its chosen criteria.

Emu Edit targets changes described in plain language

Emu Edit applies text instructions to images across tasks such as local or global edits, adding or removing backgrounds, changing color or geometry, and detection and segmentation. Meta says the system aims to alter only the pixels relevant to a request, leaving other parts of the image unaffected.

The model was trained using a dataset of tens of millions of synthesized examples covering 16 image-processing tasks. Each example pairs an input image with a description of the task and the desired output. Learned task embeddings help steer the generation toward the requested type of processing.

The researchers also report that Emu Edit can generalize to tasks including image inpainting and super-resolution, as well as combinations of processing tasks, from only a few labeled examples. That may help in settings where high-quality examples are scarce. They found that computer vision tasks improved editing performance, and that performance also rose as the number of training tasks increased.

Research results, with product use still ahead

Meta reports that Emu Edit outperformed existing methods in qualitative and quantitative evaluations across several image-processing tasks. The researchers also say it followed editing instructions more effectively while preserving the original image’s visual quality. Those findings describe research evaluations, not a guarantee that every requested edit will work as intended.

The potential uses range from making animated stickers or GIFs to editing photos and other images. But Emu Video and Emu Edit are still research projects. The source article suggests Meta could eventually bring such capabilities into communication products such as Instagram and WhatsApp, though it does not say that this has happened.

For now, the two models show how text could serve as a shared control for creating and changing visual content: one system turns prompts into motion, while the other directs edits to an image. Their practical reach will depend on what happens after the research stage.