So far we've covered distributions, Bayes' theorem, MLE, cross-entropy, and KL divergence. Move on to VAEs or diffusion models, and suddenly the page fills with probability expressions:
, , , , , , ...
My first reaction was: what are all these? What confused me more than any derivation was what each probability actually meant. So this post skips the derivations and goes through each expression with concrete examples.
I also made a one-page summary image with ChatGPT's image generator. It's worth a look before you read on (the labels are in Korean, but the layout follows this post).

First, what are , , and ?
Before the probabilities, let's meet the cast. Generative models mostly talk about and ; diffusion adds .
| Symbol | Meaning | Example |
|---|---|---|
| Real data | A cat photo, a handwritten digit | |
| The hidden cause that explains the data (latent variable) | Handwriting style, slant, the cat's pose, lighting | |
| Original data | The clean image before any noise | |
| Data at step | The original image with some noise mixed in | |
| Almost pure noise | Gaussian noise with the image information nearly gone |
VAEs mostly look at the relationship between and . Diffusion models look at : the same data with more and more noise over time. In short:
Keep that picture in mind and the probabilities get much less confusing.
: how probable the data itself is
is the probability of the data . For a generative model, it's the thing we ultimately want to know.
If is a cat photo, asks how plausible this photo is under the real distribution of images. A well-taken cat photo sits squarely on that distribution. An image of a cat with seven eyes and twisted legs should get a very low probability.
Training a generative model means making the model's distribution match the real data distribution . The big goal is: make realistic have high probability.
The catch is that is hard to compute directly. VAEs work around it with the ELBO; diffusion models work around it with a multi-step probabilistic process.
: how plausible a hidden cause is to begin with
is the prior over the latent variable : what we assume about before seeing any data. VAEs usually make it a standard normal:
Why? To make sampling easy later. After training, a generative model has to make new data, which means we should be able to pick almost any , feed it to the decoder, and get a plausible . If the latent space is scattered in strange ways, we won't know where to sample. So we encourage to be organized like a standard normal.
For handwritten digits, might hold stroke thickness, slant, and position. decides what overall distribution this hidden information should come from.
: the probability of given
is the probability of when is given. In a VAE, think of it as the decoder direction: look at a latent and produce the data .
Suppose contains:
- The digit is 7.
- It leans slightly to the right.
- The strokes are thin.
- There's a small hook at the top.
Then is the probability of a handwritten 7 that matches this description. A high means explains well. The VAE's reconstruction term pushes this value up: given , reconstruct the original well.
: seeing , which caused it?
is the probability that a particular produced the we're looking at. This is the encoder direction.
Look at a handwritten digit and you might say: "That's a 7, it leans right, and the strokes are thin." Those judgments are the hidden information . So infers the hidden causes backward from the image. In Bayes' form:
The meaning is intuitive. For a to be plausible, two things must hold: itself shouldn't be strange under the prior, ; and should actually produce well, .
The problem is that is hard to compute, mainly because in the denominator has to account for every possible . So a VAE doesn't use this true posterior directly; it uses an approximation.
: the encoder that approximates
is central to VAEs. Since the true posterior is intractable, we approximate it with a neural network; are the encoder's parameters. The encoder looks at and estimates: "this probably came from a like this."
Importantly, outputs a distribution, not a single value. It usually predicts a mean and a variance:
So even for one image , the candidates for the that could explain it are expressed as a distribution. For me, this is where a VAE starts to feel genuinely different from an autoencoder. An autoencoder compresses its input into one vector and rebuilds it; a VAE learns the distribution of latent variables that explain its input.
In diffusion, is usually the forward process
Diffusion papers use and constantly. By convention, is the forward process: the process that adds noise.
This is the probability of given . In words: if we add a little more noise to the previous image, what does the next one look like?
If is a clean cat photo, has a tiny bit of noise, a bit more, and after enough steps is almost pure noise. The forward process is a Gaussian process we fix in advance. It isn't learned; it's a noising procedure we design. So in diffusion, is usually the process we already know.
: jumping straight from the original to step
You'll also see
the probability of the noisy image given the original . Where takes one step at a time, jumps from the original to step in one go.
Suppose we need to add noise 100 times. adds it once per step. takes the clean photo and directly samples a version noised to step .
This matters because noising step by step from 1 to every training iteration would be wasteful. Thanks to properties of the Gaussian, we can sample directly from . That's what lets diffusion training pick a random and noise to that level in a single shot.
: the model that learns to denoise
Now the most important direction. Generation in a diffusion model starts from noise and ends at an image: from back through . That uses
the probability of the slightly cleaner image given the current :
Unlike the forward process, we don't know this one. Adding noise is easy; removing it is hard. So a neural network has to learn the reverse process, and are its parameters. Whenever you see in a diffusion paper, it's usually safe to read it as the reverse, generative process the model learns.
: the reverse step we can only compute during training
One more expression tends to confuse:
This is the distribution of when we know both and the original . It looks a lot like , but there's a difference.
is what we actually use at generation time, when we don't know . is a posterior we can compute during training, because then we do have . We have the original image, we noised it into ourselves, so we can work out what distribution the intermediate should follow. The model is trained to match that answer.
Very roughly, diffusion training is: make the we'll use at generation time imitate the reverse posterior of that we can compute during training.
Telling and apart
At this point and blur together, so here's the diffusion version in one table:
| Expression | Direction | What it means |
|---|---|---|
| Forward | Add one step of noise | |
| Forward | Sample the step- noisy image directly from the original | |
| Posterior | The correct reverse step, computable when the original is known | |
| Reverse | The denoising step the model has to learn |
The forward process is designed, so we know it. The reverse process is unknown, so we learn it. And during training we know , so we can give the model a target to follow. Diffusion looks complicated mostly because these three are constantly mixed together in the equations.
One example, start to finish
Let's walk through a single cat photo. is the clean photo.
In the forward process, noise is added bit by bit. adds a tiny bit to make ; adds a bit more to make . Keep going, and is noise in which you can't tell whether there was ever a cat.
Generation runs the other way. Draw a noise sample . The model applies to remove a little noise, then to remove a little more. Repeat, and you end up with an image close to .
So generation in a diffusion model is: start from pure noise and follow the learned reverse process, restoring the image a little at a time. That's why matters so much. At every step, the model answers one question: "what does this image look like with a little less noise?"
Quick reference
- : the probability of the data. What a generative model ultimately wants to model well.
- : the prior on the latent variable, usually a standard normal so it's easy to sample.
- : the probability of given . The decoder direction.
- : the probability of given . The true posterior, but intractable.
- : the encoder distribution that approximates .
- : the diffusion forward process. One step of added noise.
- : sample the step- noisy image straight from the original.
- : the correct reverse posterior, computable during training.
- : the reverse denoising process the model learns.
There's a lot of notation, but it boils down to a few questions:
- Does this probability point in the direction of generating data, or of inferring causes from data?
- In diffusion, does it add noise or remove noise?
- Is it a process we already know, or one we have to learn?
With those questions in hand, you can keep your bearings even when the equations get messier. Next, we get into model structure properly, starting with autoencoders.