With this post, the series moves into diffusion paper reading. The paper is Calvin Luo's Understanding Diffusion Models: A Unified Perspective (2022). In one line: it shows how diffusion arises as the end point of a chain of ideas that starts from the VAE. (It also happens to be the paper my senior labmate recommended.)
The paper is packed with equations. I'll follow its order as closely as I can and focus on why each equation is needed and what it actually says. The scope is the paper's treatment of the ELBO, VAEs, Markovian hierarchical VAEs, variational diffusion models, and score-based generative models.
I've deliberately stayed on the level of meaning rather than full derivations; other writers already cover the math well, and it's important. I've linked some resources for the math at the end. For the canonical DDPM paper, see Ho et al., 2020.
1. Where the paper starts: the ELBO
The author doesn't jump into diffusion. He first lays down the basics of latent variable models: behind the data we observe sits an unobserved latent , and a generative model wants to make the likelihood large.
This is Equation (1). The data we see could have come from any latent , so summing over all of them gives . The problem is that this integral is intractable for complex models, so the paper brings back the ELBO:
This is the key relationship in Equation (15). We want but can't reach it, so we maximize the ELBO instead. Since the KL divergence is never negative, raising the ELBO also pushes the approximate posterior toward the true posterior . The paper revisits the ELBO at length because reinterpreting diffusion through the ELBO later needs this foundation.
2. Why revisit the VAE
Next, the paper rewrites the VAE's ELBO, almost exactly as in the previous post:
In Equation (19), this splits into two terms:
- Reconstruction term: how well the original data is rebuilt from the latent .
- Prior matching term: how close the encoder's distribution is to the prior .
A VAE asks for two things at once: reconstruct well, and organize the latent space into a distribution that's good to sample from. We've seen that already. The next step is the question: what if we stack many latent layers instead of one?
3. From a VAE to a hierarchical VAE
The paper explains diffusion with a hierarchical VAE (HVAE), specifically a Markovian HVAE (MHVAE). The key is that there's no longer one latent but a sequence :
These are Equations (23) and (24). Instead of one latent explaining the data, a hierarchy of latents explains it. "Markovian" is the same Markov property you meet in reinforcement learning: each step depends only on the one before it. It looks complicated, but it's really just VAEs stacked in layers.
Here the author declares that he'll treat diffusion not as something new, but as a special case of this MHVAE. (The paper's Figure 3 shows the resulting chain from clean data to pure noise.)
4. How the paper defines diffusion: the VDM
The paper calls diffusion a variational diffusion model (VDM) and describes it as an MHVAE with three restrictions:
- The latent dimension equals the data dimension, so there's no compression.
- The encoder isn't learned. It's a fixed, linear Gaussian noising process, which keeps computation simple. (In code, this is usually implemented by a noise scheduler.)
- The noise schedule is designed so that the latent at the final timestep is a standard normal.
Those three lines are essentially the answer to "why do diffusion latents just look like noisy images?" In a VDM, the intermediate latents aren't meaningful abstractions; they're the original data with more and more noise mixed in.
These are Equations (30) and (31). The forward process mixes a little more Gaussian noise into the previous image to produce . The generative model learns the transitions the other way, from a noisier sample to a cleaner one:
These are Equations (32) and (33). Sampling starts from pure Gaussian noise and repeatedly applies down to .
In short, a VDM is a VAE with a hierarchy of latents, where every latent is one stage of a noisy image.
5. Equation (58), one of the most important in the paper
The paper then expands the ELBO of the VDM. The derivation is long, but the equation to hold onto is Equation (58):
- First term: reconstruction.
- Second term: whether the final noisy latent matches the Gaussian prior.
- Third term: whether, at every step, the model's reverse step follows the true posterior.
The third term is the one that really matters, because it's how the paper reads diffusion training:
"Get good at removing noise," stated precisely, means: make each model reverse transition resemble the ground-truth reverse transition.
Before moving on: the conditions on a diffusion model can feel oddly specific. They did to me, and I wondered why diffusion needs them. The latent dimension equals the data dimension; the encoder-like is a fixed Gaussian noising process rather than a network; and the final latent is designed to be close to a standard normal. My own take is that these conditions are exactly what make diffusion trainable.
- Because is a fixed linear Gaussian process, we can compute in closed form how much noise is mixed into at any timestep . In code, a noise scheduler handles this: it decides, for each timestep, how much of the original to keep and how much noise to add. Note that the scheduler doesn't fix the noise values themselves. It fixes the noise level; the actual noise is sampled from a Gaussian every time.
- Because the latent dimension equals the data dimension, all keep the same shape. The data's form stays put across timesteps; only the proportion of noise grows.
- The schedule makes nearly a standard normal for large enough . So generation doesn't have to start from some complicated distribution: sample Gaussian noise and denoise it step by step through the reverse process.
In other words, diffusion freezes the forward process into something simple and computable, and has a neural network learn the reverse.
6. Why the closed forms and matter
This is where the paper puts in real effort. Because the forward process is linear Gaussian, the distribution from the original image to any timestep has a closed form:
This is Equation (70). With only , we can sample any noisy state directly, without generating in order. That's why training can pick timesteps at random (read it together with the noise-scheduler point above).
With Bayes' rule, the true reverse posterior is tractable too:
This summarizes Equation (84). During training we know the clean image , so given a noisy we can compute the correct way to take one step toward clean. The model just has to learn to imitate that distribution. At sampling time we don't know , so takes its place. Training has a teacher; generation keeps only the student who learned from it.
7. So what does the network actually predict?
This was the part of the paper I enjoyed most. It shows the VDM can be read in three fully equivalent ways:
The prediction target can be:
- the clean image ,
- the source noise that was added,
- or the score function (a different line of thought from everything so far, which I'll come back to later).
The paper derives the noise-prediction form in Equation (125) and the score-prediction form in Equation (143). Equation (151) shows they differ only by a scale factor:
This makes it clear why "noise prediction" and "score prediction" are the same story. The score points toward higher probability, back toward the data manifold. The noise tells us how far the current image has strayed from the original. Both are signals that point in the denoising direction. Many implementations predict , and this paper shows that isn't just an empirical trick; it's mathematically tied to the score.
8. Moving to the score-based view
Score-based models come from a different lineage than the one this series has followed, so I'll cover their foundations separately later. For now, a light pass is enough.
If everything so far viewed diffusion from the VAE side, the paper now views it as a score-based model. It starts from an energy-based model:
This is Equation (152). is the energy function, and is a normalizing constant that's generally intractable. So instead of learning the distribution itself, we learn the gradient of its log probability, the score:
That's the gist of Equations (153) to (156). We learn a vector field that says, at every point in data space, which direction increases the likelihood. Think of the score as arrows pulling samples toward the modes. Sampling then follows the arrows:
This is Equation (158), Langevin dynamics: move a little in the direction the score points, and add a bit of Gaussian noise. The noise keeps samples from always collapsing into the same mode and gives them diversity.
9. The trouble with vanilla score matching
Following Song and Ermon, the paper lists the limits of vanilla score matching:
- First: real images lie on a low-dimensional manifold, so the score is poorly defined off it.
- Second: the learning signal is weak in low-density regions.
- Third: Langevin dynamics may not mix well between modes.
Learning the score of the original data distribution alone may not be enough for good generation. The fix is to learn scores at many noise levels together:
These are Equations (160) and (161). Learn the score for a lightly noised distribution, a heavily noised one, and everything in between. That gives smoother learning signals in low-density regions as well as high-density ones. And the paper shows this objective has almost the same form as the VDM objective derived earlier.
10. The unified perspective
Now the title makes sense. The connections the paper draws are roughly:
- A VAE explains generation with a latent variable model and the ELBO.
- An MHVAE extends that latent into many steps.
- A VDM is a special MHVAE whose steps are fixed to a chain of noisy images.
- Its training target can be read as prediction, noise prediction, or score prediction.
- Through score prediction, diffusion connects directly to score-based generative models.
Summary
Following the first part of the paper, we traced a line from VAEs all the way to score-based models. Diffusion isn't just "add noise, remove noise." It's a latent variable model with an ELBO, and at the same time a model that learns a score function. The paper shows these aren't separate views but two languages for the same model.
For me, the most important point is treating the VDM as a special case of the MHVAE and reading its ELBO as step-by-step denoising matching.
Next, I'll take these ideas to robotics with Diffusion Policy.
Further resources
For the math of diffusion models (the first two are the math-heavy ones):
- Video on the math, 1 (YouTube)
- Video on the math, 2 (YouTube)
And one on how diffusion models are actually trained. The math matters, but I think how the model is trained matters just as much:
- Video on training (YouTube)