Meta AI has developed a system that turns magnetoencephalography (MEG) recordings into reconstructed images in milliseconds. The images are imperfect, but can reflect broad features of what a person sees, such as whether it is a train, an animal, or a human. Details, including exactly which animal appears, can be wrong.
How the system turns brain activity into images
The approach has three parts: an image encoder, a brain encoder, and an image decoder. Together, they connect patterns in visual data with patterns recorded from the brain, then use those brain-linked representations to generate a plausible image.
The image encoder first creates representations of an image without using brain data. The brain encoder then matches MEG signals to those representations. Finally, the image decoder uses the matched brain representations to produce an image.
This design builds on Meta AI's recently developed architecture for decoding speech perception from MEG signals. The image system was trained with a publicly available dataset of MEG recordings from healthy volunteers, published by the international academic research consortium Things.
Fast results, with limits on detail
The system's main advantage is speed. It can decode images within milliseconds, producing a continuous stream of visual predictions from brain activity and working almost in real time.
That speed comes with a trade-off in precision. The generated image can convey the general category or characteristics of what a person sees, but may miss or misrepresent finer details. The method is therefore a way to read broad visual content from brain signals, rather than to reproduce a person's view exactly.
Functional magnetic resonance imaging (fMRI) can produce more accurate image predictions from brain data, according to the source article, but it is slower. MEG offers a quicker stream of predictions, while fMRI offers greater accuracy. The two approaches illustrate a practical tension between how quickly brain activity can be decoded and how precisely an image can be reconstructed.
What the comparison with vision AI suggests
Meta AI researchers compared different pre-trained image modules to see which representations best matched brain signals. They found the strongest match with advanced vision AI systems such as DINOv2, a self-supervising architecture that learns visual representations without human guidance.
Meta AI says this finding supports the idea that self-supervised learning can lead AI systems to develop representations resembling those in the brain. When an algorithm and a brain are exposed to the same image, artificial neurons in the system can activate in ways similar to physical neurons, according to the researchers.
The work connects image decoding with a broader research question: how visual information is represented and used as a basis for human intelligence. The system offers researchers a way to study those representations as they unfold, with MEG's millisecond timing providing a continuous view of changing brain activity.
A research step toward future brain-computer interfaces
Meta researchers describe the work as part of a long-term effort to understand the foundations of human intelligence and build AI systems that learn and think more like people. The research could also, over time, contribute to non-invasive brain-computer interfaces in clinical settings.
One possible application mentioned by the researchers is helping people who have lost the ability to speak after a brain injury. The current image decoder does not itself demonstrate that kind of communication tool. Its contribution is an approach for translating recorded brain activity into interpretable predictions, and a foundation that further research may build on.
The project also fits into Meta's work on machine intelligence that uses more abstract representations. In early 2022, Yann LeCun, director of research at Meta AI, unveiled a new AI architecture intended to overcome limitations of current systems. Later, a team from several institutions, including Meta AI and New York University, demonstrated I-JEPA, a visual AI model based on the Vision Transformer.
I-JEPA was trained through self-supervised learning to predict details of parts of an image that are not visible. Rather than working at the pixel or token level, it learned abstract representations of objects. According to LeCun's theory, such representations could help create AI models that more closely resemble human learning and draw logical conclusions.