I'm a senior in electrical engineering. I never took a probability and statistics course, and I'm not in an AI department, so deep learning did not come easily to me. Generative models were the hardest part: the notation piles up quickly, and most explanations assume you already know why each symbol is there.
This series is the path I took from the ground up, from autoencoders (AE) to variational autoencoders (VAE) and finally to diffusion models. I'm writing it down in case the same path helps someone else.
How the series is organized
- Part 1. Introduction (this post)
- Part 2. Distributions, Bayes' theorem, expectation, and Jensen's inequality
- Part 3. Maximum likelihood, cross-entropy, and KL divergence
- Part 4. What each probability actually means
- Part 5. Autoencoders
- Part 6. VAEs and the ELBO
- Part 7. Reading Understanding Diffusion Models: A Unified Perspective
- Parts 9 and 10. Reading Diffusion Policy: part 1 and part 2
The numbers follow the original Korean series, which has no published part 8.
I'll explain things from the point of view of a fourth-year undergraduate rather than a specialist. Where the math gets heavy, I'll try to say what each expression means, and I'll spend extra time on the pairs that are easy to mix up, such as and .
Resources that helped me
I learned this material the hard way. I read the main diffusion paper about five times before it started to make sense, and I leaned on a lot of other people's explanations along the way. These were the most useful:
- "Everything about autoencoders" by Hwalseok Lee (YouTube, in Korean). An outstanding lecture. It gave me the overall structure, and I filled in the equations and finer details myself afterward.
- Hyeongmin Lee's post on Bayes' rule (blog, in Korean). Bayes' theorem turned out to be more slippery than I expected, and this post helped the most.
- Calvin Luo, Understanding Diffusion Models: A Unified Perspective. My senior labmate recommended it as the main paper to study. Once I had the background, it turned out to be one of the kindest papers on the subject.
I'll cover the same ideas in later posts, but these sources are worth keeping open alongside the series.
