Microsoft researchers have introduced Composable Diffusion, or CoDi, a model that can process and generate text, images, video and audio. The system is designed to combine these forms of content in a single process, including producing video and sound that are synchronized.
Combining different kinds of input and output
CoDi is part of Microsoft's i-Code project, which aims to develop integrative and composable multimodal AI. Unlike systems built around a specific input type, CoDi is designed to work across several modalities at once.
That flexibility matters because training data for many combinations of modalities is scarce. The researchers describe an alignment strategy that connects modalities in both the input and output spaces. This allows the model to condition on different combinations of inputs and generate different combinations of outputs, including combinations not represented in the training data.
In practical terms, the design is intended to let the system use a mix of information rather than require a separate generative model for each format. The approach addresses a limitation of single-modality models and the cumbersome, slow process of assembling modality-specific systems.
How CoDi builds a shared representation
During training, CoDi projects inputs such as images, video, audio and language into a common semantic space. This gives the model a shared basis for working with information that arrives in different forms.
A cross-attention module and an environment encoder then support the generation of multiple output modalities at the same time. The researchers also describe a composable generation strategy that links alignment with the diffusion process. Its purpose is to support intertwined outputs, such as video paired with audio that matches it in time.
Synchronization is a central part of the model's design. Generating a clip and a separate soundtrack would not be enough for the example described by the researchers: the sound needs to correspond with what is happening in the video.
A demonstration with a rainy Times Square scene
One demonstration combined a text prompt, an image and a sound. The text prompt was "teddy bear on skateboard, 4k, high resolution"; the other inputs were an image of Times Square and the sound of rain.
CoDi produced a short video of a teddy bear skateboarding in the rain at Times Square, with synchronized rain and street noise. The source describes the video as low quality, a reminder that the demonstration illustrates the model's ability to combine inputs and outputs rather than polished production quality.
The example brings together three distinct cues: a written description, a visual setting and environmental sound. CoDi's generated scene reflects all three, while coordinating its video and audio outputs. It offers a concrete illustration of what a model with cross-modal generation can do.
Possible uses and open-ended work
The researchers point to education and accessibility for people with disabilities as possible areas of application. The article does not detail specific products or use cases, but the model's ability to handle multiple forms of content suggests why these fields are relevant to the work.
For education, the broader idea is that information could be represented through more than one modality. For accessibility, flexible processing and generation across text, image, video and audio could matter when interactions need to accommodate different ways of receiving or creating information. These are potential directions identified by the researchers, not evidence of deployed applications.
The team frames CoDi as a step toward more engaging and holistic human-computer interaction, and as a foundation for further investigation in generative artificial intelligence. Its key contribution, as presented, is a composable approach to multimodal AI: one designed to combine varied inputs and produce coordinated outputs even when training examples for every possible pairing are unavailable.