HOYALABA PERSONAL RESEARCH NOTEBOOK

Foundations / 2026-04-22 / 10 MIN

What each probability actually means: a field guide to VAE and diffusion notation

p(x), p(z), p(x|z), p(z|x), q, and p-theta: what each expression in VAEs and diffusion models describes, explained with concrete examples instead of derivations.

So far we've covered distributions, Bayes' theorem, MLE, cross-entropy, and KL divergence. Move on to VAEs or diffusion models, and suddenly the page fills with probability expressions:

p(x)p(x), p(z)p(z), p(x∣z)p(x \mid z), p(z∣x)p(z \mid x), qϕ(z∣x)q_\phi(z \mid x), q(xt∣xt−1)q(x_t \mid x_{t-1}), pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t)...

My first reaction was: what are all these? What confused me more than any derivation was what each probability actually meant. So this post skips the derivations and goes through each expression with concrete examples.

I also made a one-page summary image with ChatGPT's image generator. It's worth a look before you read on (the labels are in Korean, but the layout follows this post).

One-page summary of the probability notation in VAEs and diffusion: the variables, the VAE probabilities, and the forward and reverse diffusion processes (labels in Korean).
FIG. 1 · One-page summary of the probability notation in VAEs and diffusion: the variables, the VAE probabilities, and the forward and reverse diffusion processes (labels in Korean).

First, what are xx, zz, and xtx_t?

Before the probabilities, let's meet the cast. Generative models mostly talk about xx and zz; diffusion adds xtx_t.

Symbol Meaning Example
xx Real data A cat photo, a handwritten digit
zz The hidden cause that explains the data (latent variable) Handwriting style, slant, the cat's pose, lighting
x0x_0 Original data The clean image before any noise
xtx_t Data at step tt The original image with some noise mixed in
xTx_T Almost pure noise Gaussian noise with the image information nearly gone

VAEs mostly look at the relationship between xx and zz. Diffusion models look at x0,x1,x2,…,xTx_0, x_1, x_2, \dots, x_T: the same data with more and more noise over time. In short:

Keep that picture in mind and the probabilities get much less confusing.

p(x)p(x): how probable the data itself is

p(x)p(x) is the probability of the data xx. For a generative model, it's the thing we ultimately want to know.

If xx is a cat photo, p(x)p(x) asks how plausible this photo is under the real distribution of images. A well-taken cat photo sits squarely on that distribution. An image of a cat with seven eyes and twisted legs should get a very low probability.

Training a generative model means making the model's distribution match the real data distribution pdata(x)p_{\text{data}}(x). The big goal is: make realistic xx have high probability.

The catch is that p(x)p(x) is hard to compute directly. VAEs work around it with the ELBO; diffusion models work around it with a multi-step probabilistic process.

p(z)p(z): how plausible a hidden cause is to begin with

p(z)p(z) is the prior over the latent variable zz: what we assume about zz before seeing any data. VAEs usually make it a standard normal:

p(z)=N(0,I)p(z)=\mathcal{N}(0,I)

Why? To make sampling easy later. After training, a generative model has to make new data, which means we should be able to pick almost any zz, feed it to the decoder, and get a plausible xx. If the latent space is scattered in strange ways, we won't know where to sample. So we encourage zz to be organized like a standard normal.

For handwritten digits, zz might hold stroke thickness, slant, and position. p(z)p(z) decides what overall distribution this hidden information should come from.

p(x∣z)p(x \mid z): the probability of xx given zz

p(x∣z)p(x \mid z) is the probability of xx when zz is given. In a VAE, think of it as the decoder direction: look at a latent zz and produce the data xx.

Suppose zz contains:

  • The digit is 7.
  • It leans slightly to the right.
  • The strokes are thin.
  • There's a small hook at the top.

Then p(x∣z)p(x \mid z) is the probability of a handwritten 7 that matches this description. A high p(x∣z)p(x \mid z) means zz explains xx well. The VAE's reconstruction term pushes this value up: given zz, reconstruct the original xx well.

p(z∣x)p(z \mid x): seeing xx, which zz caused it?

p(z∣x)p(z \mid x) is the probability that a particular zz produced the xx we're looking at. This is the encoder direction.

Look at a handwritten digit and you might say: "That's a 7, it leans right, and the strokes are thin." Those judgments are the hidden information zz. So p(z∣x)p(z \mid x) infers the hidden causes backward from the image. In Bayes' form:

p(z∣x)=p(x∣z) p(z)p(x)p(z \mid x)=\frac{p(x \mid z)\,p(z)}{p(x)}

The meaning is intuitive. For a zz to be plausible, two things must hold: zz itself shouldn't be strange under the prior, p(z)p(z); and zz should actually produce xx well, p(x∣z)p(x \mid z).

The problem is that p(z∣x)p(z \mid x) is hard to compute, mainly because p(x)p(x) in the denominator has to account for every possible zz. So a VAE doesn't use this true posterior directly; it uses an approximation.

qϕ(z∣x)q_\phi(z \mid x): the encoder that approximates p(z∣x)p(z \mid x)

qϕ(z∣x)q_\phi(z \mid x) is central to VAEs. Since the true posterior is intractable, we approximate it with a neural network; ϕ\phi are the encoder's parameters. The encoder looks at xx and estimates: "this probably came from a zz like this."

Importantly, qϕ(z∣x)q_\phi(z \mid x) outputs a distribution, not a single value. It usually predicts a mean and a variance:

qϕ(z∣x)=N(μϕ(x),σϕ2(x))q_\phi(z \mid x)=\mathcal{N}\big(\mu_\phi(x), \sigma_\phi^2(x)\big)

So even for one image xx, the candidates for the zz that could explain it are expressed as a distribution. For me, this is where a VAE starts to feel genuinely different from an autoencoder. An autoencoder compresses its input into one vector and rebuilds it; a VAE learns the distribution of latent variables that explain its input.

In diffusion, qq is usually the forward process

Diffusion papers use qq and pθp_\theta constantly. By convention, qq is the forward process: the process that adds noise.

q(xt∣xt−1)q(x_t \mid x_{t-1})

This is the probability of xtx_t given xt−1x_{t-1}. In words: if we add a little more noise to the previous image, what does the next one look like?

If x0x_0 is a clean cat photo, x1x_1 has a tiny bit of noise, x2x_2 a bit more, and after enough steps xTx_T is almost pure noise. The forward process is a Gaussian process we fix in advance. It isn't learned; it's a noising procedure we design. So in diffusion, qq is usually the process we already know.

q(xt∣x0)q(x_t \mid x_0): jumping straight from the original to step tt

You'll also see

q(xt∣x0),q(x_t \mid x_0),

the probability of the noisy image xtx_t given the original x0x_0. Where q(xt∣xt−1)q(x_t \mid x_{t-1}) takes one step at a time, q(xt∣x0)q(x_t \mid x_0) jumps from the original to step tt in one go.

Suppose we need to add noise 100 times. q(xt∣xt−1)q(x_t \mid x_{t-1}) adds it once per step. q(xt∣x0)q(x_t \mid x_0) takes the clean photo and directly samples a version noised to step tt.

This matters because noising step by step from 1 to tt every training iteration would be wasteful. Thanks to properties of the Gaussian, we can sample xtx_t directly from x0x_0. That's what lets diffusion training pick a random tt and noise x0x_0 to that level in a single shot.

pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t): the model that learns to denoise

Now the most important direction. Generation in a diffusion model starts from noise and ends at an image: from xTx_T back through xT−1,xT−2,…,x0x_{T-1}, x_{T-2}, \dots, x_0. That uses

pθ(xt−1∣xt),p_\theta(x_{t-1} \mid x_t),

the probability of the slightly cleaner image xt−1x_{t-1} given the current xtx_t:

Unlike the forward process, we don't know this one. Adding noise is easy; removing it is hard. So a neural network has to learn the reverse process, and θ\theta are its parameters. Whenever you see pθp_\theta in a diffusion paper, it's usually safe to read it as the reverse, generative process the model learns.

q(xt−1∣xt,x0)q(x_{t-1} \mid x_t, x_0): the reverse step we can only compute during training

One more expression tends to confuse:

q(xt−1∣xt,x0)q(x_{t-1} \mid x_t, x_0)

This is the distribution of xt−1x_{t-1} when we know both xtx_t and the original x0x_0. It looks a lot like pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t), but there's a difference.

pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) is what we actually use at generation time, when we don't know x0x_0. q(xt−1∣xt,x0)q(x_{t-1} \mid x_t, x_0) is a posterior we can compute during training, because then we do have x0x_0. We have the original image, we noised it into xtx_t ourselves, so we can work out what distribution the intermediate xt−1x_{t-1} should follow. The model pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) is trained to match that answer.

Very roughly, diffusion training is: make the pθp_\theta we'll use at generation time imitate the reverse posterior of qq that we can compute during training.

Telling qq and pθp_\theta apart

At this point qq and pp blur together, so here's the diffusion version in one table:

Expression Direction What it means
q(xt∣xt−1)q(x_t \mid x_{t-1}) Forward Add one step of noise
q(xt∣x0)q(x_t \mid x_0) Forward Sample the step-tt noisy image directly from the original
q(xt−1∣xt,x0)q(x_{t-1} \mid x_t, x_0) Posterior The correct reverse step, computable when the original is known
pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) Reverse The denoising step the model has to learn

The forward process is designed, so we know it. The reverse process is unknown, so we learn it. And during training we know x0x_0, so we can give the model a target to follow. Diffusion looks complicated mostly because these three are constantly mixed together in the equations.

One example, start to finish

Let's walk through a single cat photo. x0x_0 is the clean photo.

In the forward process, noise is added bit by bit. q(x1∣x0)q(x_1 \mid x_0) adds a tiny bit to make x1x_1; q(x2∣x1)q(x_2 \mid x_1) adds a bit more to make x2x_2. Keep going, and xTx_T is noise in which you can't tell whether there was ever a cat.

Generation runs the other way. Draw a noise sample xTx_T. The model applies pθ(xT−1∣xT)p_\theta(x_{T-1} \mid x_T) to remove a little noise, then pθ(xT−2∣xT−1)p_\theta(x_{T-2} \mid x_{T-1}) to remove a little more. Repeat, and you end up with an image close to x0x_0.

So generation in a diffusion model is: start from pure noise and follow the learned reverse process, restoring the image a little at a time. That's why pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) matters so much. At every step, the model answers one question: "what does this image look like with a little less noise?"

Quick reference

  • p(x)p(x): the probability of the data. What a generative model ultimately wants to model well.
  • p(z)p(z): the prior on the latent variable, usually a standard normal so it's easy to sample.
  • p(x∣z)p(x \mid z): the probability of xx given zz. The decoder direction.
  • p(z∣x)p(z \mid x): the probability of zz given xx. The true posterior, but intractable.
  • qϕ(z∣x)q_\phi(z \mid x): the encoder distribution that approximates p(z∣x)p(z \mid x).
  • q(xt∣xt−1)q(x_t \mid x_{t-1}): the diffusion forward process. One step of added noise.
  • q(xt∣x0)q(x_t \mid x_0): sample the step-tt noisy image straight from the original.
  • q(xt−1∣xt,x0)q(x_{t-1} \mid x_t, x_0): the correct reverse posterior, computable during training.
  • pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t): the reverse denoising process the model learns.

There's a lot of notation, but it boils down to a few questions:

  • Does this probability point in the direction of generating data, or of inferring causes from data?
  • In diffusion, does it add noise or remove noise?
  • Is it a process we already know, or one we have to learn?

With those questions in hand, you can keep your bearings even when the equations get messier. Next, we get into model structure properly, starting with autoencoders.