A robot may need to follow a spoken-style instruction, copy a demonstrated action, or arrange objects to match a picture. VIMA is designed to bring those different jobs together in one model, using prompts that can mix words and images to describe what the robot should do.
One prompt format for different robot tasks
Many robotics tasks have traditionally been handled by specialized models. One system might imitate an action shown once, while another follows language instructions or works toward a visual goal. VIMA instead treats these tasks as variations on a shared problem: interpreting a prompt and producing a sequence of actions.
The prompt can contain both text and visual information. For instance, a user can ask the system to rearrange objects to match a pictured scene. VIMA processes the instruction and image together, then controls a robotic arm in a simulation to attempt the arrangement.
This approach gives users a way to convey intent through more than words alone. An image can show the desired outcome, while the accompanying text explains the task. The researchers describe this as multimodal prompting, intended to provide an intuitive interface for a general-purpose robot agent.
Training in a simulated tabletop world
The team created VIMA-Bench, a simulation benchmark made up of thousands of procedurally generated tabletop tasks. The tasks span 17 categories and have corresponding prompts that combine different kinds of input.
For imitation learning, the benchmark includes more than 600,000 expert trajectories. These examples provide demonstrations of how tasks can be carried out, giving the model experience with a broad range of simulated situations.
VIMA is based on a Transformer architecture. The researchers frame manipulation tasks as a uniform sequence modeling problem, allowing one model to learn from prompts and action sequences across tasks rather than relying on a separate specialized system for each capability.
Reported gains in performance and data efficiency
According to the researchers, VIMA outperforms Gato, Flamingo and Decision Transformer by up to 2.9 times across model sizes and levels of generalization. The largest VIMA model has 200 million parameters.
The team also reports a data-efficiency advantage in imitation training. VIMA achieves performance comparable to other methods while using 10 times less data. That result matters because collecting expert demonstrations is part of how the model learns to carry out tasks; using fewer examples for similar performance could make training more efficient.
The reported results come from the researchers' benchmark and simulation setting. They show how the model handles the tasks represented in that environment, including visual goals, one-shot video imitation and grounding novel concepts. The article does not establish how the same results would transfer to every real-world robot or setting.
A direction for multimodal robotics
VIMA follows a broader effort to apply language and multimodal models to robotics. Other projects described in the source include Google's PaLM-SayCan, real-time robot control using a language model, and Robotics Transformer 1, or RT-1, a multimodally trained robotics model. VIMA's distinction is that its prompts themselves can combine text and images.
The researchers present the work as a foundation for further development. Google described multimodal prompting as a promising future direction for RT-1, suggesting that combining different forms of input is of interest beyond VIMA itself.
For people directing robots, the central idea is straightforward: express a task in a format that can show both what to do and what the intended result looks like. VIMA explores how one model can interpret that mixed instruction and translate it into actions in a simulated setting.