Meta’s Make-A-Video3D, or MAV3D, is designed to turn a text description into an animated 3D scene. The approach combines a model that represents dynamic scenes with Meta’s Make-A-Video system, using text to guide what the scene should show.
The work points toward a way to create content that can move beyond a generated video: the learned scene representation can be converted into animated meshes and rendered in real time in a 3D engine. The process is still inefficient, and Meta says it wants to improve both efficiency and resolution.
How MAV3D connects text, video and 3D
MAV3D builds on Neural Radiance Fields, or NeRFs, which represent a scene in a form that can be used to generate images from different camera positions. Meta uses a variant called HexPlane, suited to dynamic scenes, to produce a sequence of images from a sequence of camera positions.
That sequence is treated as video and supplied, along with the text prompt, to Meta’s Make-A-Video model. Make-A-Video evaluates how well the content from HexPlane matches the prompt and other parameters. Its score then acts as a learning signal: HexPlane adjusts its parameters over repeated passes until its representation better corresponds to the description.
The key idea is the feedback loop between the scene representation and the video model. The video model assesses whether the generated content fits the words, while HexPlane uses that assessment to refine the 3D scene.
From a written prompt to an animated scene
Meta’s examples include a singing cat, a baby panda eating ice cream, and a squirrel playing the saxophone. In each case, MAV3D produces a three-dimensional representation of the described subject and action. The source says the results match their text prompts, while also noting that there is not yet a qualitatively comparable model.
These examples illustrate the system’s goal: a prompt can describe both what is present and what is happening, and the model generates a dynamic scene rather than only a still image. The output is represented in a way that can be viewed from camera positions, linking the text-to-video guidance with a 3D scene model.
Why real-time rendering matters
Meta says the learned HexPlane representation can be converted into animated meshes. Those meshes could then be rendered in real time in a standard 3D engine. That capability could make generated scenes useful in virtual reality or in traditional video games, where content needs to function inside an interactive environment.
Real-time rendering changes the potential role of generated content. A scene that can be placed into a 3D engine may be used as an asset in an application, rather than remaining only a video to watch. The source describes this as a possible fit for VR and games; it does not say that MAV3D is already a production-ready tool for either.
What remains unfinished
The current process is inefficient, and Meta’s team is working to make it more efficient and improve scene resolution. Those limits matter because an interactive 3D scene must be generated and rendered in a form that is practical for its intended use.
The model and code are not available. Meta’s project page includes more video examples and renderings, but the described research remains a demonstration of a method rather than a publicly available system readers can run.
MAV3D brings together text-guided video evaluation and a dynamic 3D representation. Its promise lies in producing scenes that can be converted for real-time use, while its efficiency and resolution remain open challenges.