HOYALABA PERSONAL RESEARCH NOTEBOOK
/

Foundations / 2026-05-01 / 3 MIN

From an autoencoder to a VAE: why the ELBO has two terms

A latent point becomes a distribution. Reconstruction and prior matching explain why that change makes a generative model possible.

An autoencoder compresses an input and reconstructs it. That sounds close to generation, but there is an awkward gap: what happens if we choose a latent vector that the decoder never encountered during training?

This was the starting point of my Korean VAE note. Before moving on to diffusion papers, I wanted to understand how a generative model organizes the distribution it samples from.

1. A point becomes a distribution

A conventional deterministic encoder maps an input to one latent vector. Good reconstruction does not, by itself, ensure that arbitrary points between those vectors decode into useful samples.

A VAE instead learns an approximate posterior, often a diagonal Gaussian:

qϕ(z∣x)=N ⁣(μϕ(x),diag⁡(σϕ2(x))).q_\phi(z\mid x)=\mathcal{N}\!\left(\mu_\phi(x),\operatorname{diag}(\sigma_\phi^2(x))\right).

Here the encoder predicts the distribution's mean and variance. We sample a latent variable from that distribution and let the decoder explain the input from the sample. I find this the first useful shift in intuition: the representation is no longer just one point.

2. Why a lower bound appears

The generative model combines a prior with a decoder:

pθ(x)=∫pθ(x∣z)p(z) dz.p_\theta(x)=\int p_\theta(x\mid z)p(z)\,dz.

This expression adds up the possible ways latent variables could produce the observation. Computing that integral directly is generally difficult for a neural decoder. The true posterior is therefore difficult too, and we introduce the tractable approximation qϕ(z∣x)q_\phi(z\mid x).

The relationship I found most helpful is

log⁡pθ(x)=ELBO⁡(x)+DKL ⁣(qϕ(z∣x) ∥ pθ(z∣x)).\log p_\theta(x)=\operatorname{ELBO}(x)+D_{\mathrm{KL}}\!\left(q_\phi(z\mid x)\,\|\,p_\theta(z\mid x)\right).

The KL term cannot be negative. That makes the ELBO a lower bound on the log evidence. With the generative model fixed, making the approximate posterior closer to the true posterior closes the gap. During training, both the model and the approximation change, so I keep that qualification in mind.

3. Two terms, two jobs

The form used for training makes the intuition easier to see:

ELBO⁡(x)=Eqϕ(z∣x)[log⁡pθ(x∣z)]−DKL ⁣(qϕ(z∣x) ∥ p(z)).\operatorname{ELBO}(x)=\mathbb{E}_{q_\phi(z\mid x)}[\log p_\theta(x\mid z)]-D_{\mathrm{KL}}\!\left(q_\phi(z\mid x)\,\|\,p(z)\right).

The first term rewards explaining the input from the latent sample. It is the reconstruction side. Depending on the decoder's observation model, its negative can take a form related to squared error or binary cross entropy.

The second term discourages the approximate posterior from drifting too far from the prior, often a standard Gaussian. It gives the latent distributions a shared reference rather than letting reconstruction alone arrange them arbitrarily.

This is why dropping either term changes the problem. Strong reconstruction alone does not establish a convenient sampling distribution. Matching the prior without preserving information about the input does not give useful reconstructions.

4. The sampling step and gradients

The reparameterization trick writes a latent sample as

z=μϕ(x)+σϕ(x)⊙ϵ,ϵ∼N(0,I).z=\mu_\phi(x)+\sigma_\phi(x)\odot\epsilon,\qquad\epsilon\sim\mathcal{N}(0,I).

The random variable is now separate from the encoder's outputs. For a sampled ϵ\epsilon, the expression is differentiable with respect to the predicted mean and scale. That lets the reconstruction signal reach the encoder through the sampled latent variable.

The notation looks like a small rearrangement. In the implementation, it is the bridge that makes this training procedure practical.

5. Generating after training

To generate a new example, we sample from the prior and pass the latent variable to the decoder. The encoder is not needed for this unconditional generation step.

A well-trained latent model can support useful sampling and interpolation, but neither the prior choice nor the training objective guarantees perfect samples. The approximation family, latent dimension, and decoder assumptions still limit what the model can represent.

For me, the VAE is a useful stop on the way to diffusion because it makes the probabilistic modeling choices explicit: what is sampled, what is learned, and what objective connects them. Those are questions worth carrying into the next paper.