Deep Diaries · · 2 min read
High-Resolution Image Reconstruction with Latent Diffusion Models from Human Brain Activity
Reconstructing visual experiences from human brain activity with Stable Diffusion
Originally published on Deep Diaries on Substack. Reproduced here as written.
Reconstructing visual experiences from human brain activity offers a unique way to understand how the brain represents the world, and to interpret the connection between computer vision models and our visual system. 1

Recently, I came across a newly published paper that discusses how we can reconstruct the scenes we observe using fMRI signals and Latent Diffusion Models.
Let's understand what diffusion models are. Diffusion Models are a type of probabilistic generative models that can build or restore variables or samples from Gaussian noise. They take an image as input and repeatedly add Gaussian noise to it while simultaneously learning how to remove the noise at each time-step (this process is called de-noising). In the end, we have a model that can start with Gaussian noise and, through the de-noising process, generate a real image from the same distribution as the training data.
Up until now, the model has learned an unconditional distribution, which means we can't control the content of the generated image; it simply belongs to the same distribution as the training data.
How can we condition the generation process? By using the latent representation of text as a condition, we can learn a conditional distribution. This is the same idea used with fMRI signals, but with a slight modification. Instead of using Diffusion Models (DM), which were very expensive as they operated at the pixel level, we will use Latent Diffusion Models (LDM) and introduce an encoder between the input and the diffusion process:
The flow can be described as follows:
First, we predict a latent representation (z) for the fMRI signals within the early visual cortex of a person while they are viewing an image (X).
Second, we decode z using the decoder to obtain the decoded image (Xz), which is then passed to the encoder and subjected to the Gaussian noise adding process to produce a noisy version (Zt) (where t represents the number of times noise was added).
We repeat the first and second steps, but instead of using fMRI signals from the early visual cortex, we apply them to the signals of the higher visual cortex (c).
We use the decoded c and Zt as inputs to the denoising U-Net to generate Zc.
In the last step, we can decode Zc to obtain an image similar to image X.

The intriguing part is that all you need to train is the linear model that maps the fMRI signals to the latent representation, converting the early visual cortex fMRI signals into the latent representation of the image after encoding and the higher visual cortex into the text latent representation that describes the image.