Ali Samir

← The writing map

Deep Diaries · · 3 min read

Unveiling the Mind's Canvas: Real-time Decoding of Visual Representations in the Brain

Exploring the Intersection of Neuroscience and AI in Understanding Brain Imaging and Perception

Originally published on Deep Diaries on Substack. Reproduced here as written.

In 1978, a movie titled "Eyes of Laura Mars" captivated audiences with its intriguing premise, centered around Laura Mars, a successful fashion photographer in the bustling streets of New York City. The unique twist in this story was Laura's extraordinary ability to perceive the world through the eyes of a serial killer.

Back then, it was a tale spun purely from the realm of fiction, a cinematic thrill designed to entertain and mystify. The concept of one person inhabiting another's vision was a fantastical notion we all readily acknowledged.

Fast forward to 2023, and our present reality unveils an altogether different narrative. In this era, a cadre of researchers is fervently dedicated to unraveling the enigma of how our brains perceive and represent the world around us. I had the pleasure of delving into one such cutting-edge study in my article titled "High-Resolution Image Reconstruction with Latent Diffusion Models from Human Brain Activity."

And yesterday, on October 18, 2023, the research landscape takes another leap forward. Meta's researchers have presented a remarkable new paper that delves into the real-time decoding of brain activity. This endeavor marks a significant stride toward deciphering the intricate workings of our most enigmatic organ.

Today, Meta is announcing an important milestone in the pursuit of that fundamental question. Using magnetoencephalography (MEG), a non-invasive neuroimaging technique in which thousands of brain activity measurements are taken per second, we showcase an AI system capable of decoding the unfolding of visual representations in the brain with an unprecedented temporal resolution.1

Using same architecture that was trained to decode speech perception from MEG signals they developed three components:

  1. Image Encoder: Constructs a collection of image representations (embeddings) without relying on the brain.

  2. Brain Encoder: It learns to align MEG signals to these image embeddings

  3. Image Decoder: It creates a convincing image based on these brain representations.

Figure 1 Shows the three components of the system

Figure 1

Should you find this intriguing, let me inform you that in accordance with Meta's findings and following a comparative analysis of the decoding performance across a range of pretrained image modules, it was revealed that the alignment of brain signals was most pronounced with contemporary computer vision AI systems such as DINOv2. DINOv2, a recent self-supervised architecture, has demonstrated its capacity to acquire intricate visual representations without the need for human annotations. This outcome substantiates the notion that self-supervised learning empowers AI systems to develop representations akin to those found in the human brain. To quote this directly:

The artificial neurons in the algorithm tend to be activated similarly to the physical neurons of the brain in response to the same image.

Check the following video shared by Meta for the decoded images from MEG

As they showed that the results can be better using functional Magnetic Resonance Imaging (fMRI) instead of MEG, but the MEG decoder can produces a continuous flow of decoded images as it can be used in real-time.

To compare the previous results with the fMRI results check the following video from Meta

Conclusion:

This pioneering research takes us one step closer to comprehending the remarkable complexity of the human brain, opening doors to unprecedented possibilities in neuroscience and artificial intelligence. As we venture further into this uncharted territory, the fusion of human cognition and AI technology promises a future filled with extraordinary breakthroughs, and maybe one day we can see through “Eyes of Laura Mars”.


You can also read:

High-Resolution Image Reconstruction with Latent Diffusion Models from Human Brain Activity