ML 101
M11 · L03
Deep Learning in Practice

Generative Models

VAEs, GANs, diffusion models, and normalizing flows — learning to model data distributions and generate new samples from scratch.

01 / 13
ML 101
M11 · L03
The Core Distinction

Discriminative vs. Generative

Discriminative models learn p(y|x) — given input, predict label. Generative models learn p(x) — the data distribution itself — enabling synthesis of new samples.

The Challenge
A 256×256 image lives in a 196K-dimensional space. Natural images occupy an astronomically tiny, structured manifold within it. Generative models must find this manifold.
02 / 13
ML 101
M11 · L03
VAE

Variational Autoencoders

Encoder maps x → latent (μ, σ²). Sample z ~ N(μ, σ²) via reparameterization. Decoder reconstructs x from z. Optimize the ELBO.

ELBO Objective
\mathcal{L}=\mathbb{E}[\log p_\theta(x|z)]-D_{\mathrm{KL}}(q_\phi(z|x)\|p(z))
03 / 13
ML 101
M11 · L03
VAE Trick

Reparameterization Trick

Sampling z directly blocks gradients. Instead, sample ε ~ N(0, I) and compute z = μ + σ·ε. Stochasticity moves outside the computational graph.

Smooth
Structured latent space — interpolation works
Blurry
MSE reconstruction → averaging over uncertainty
04 / 13
ML 101
M11 · L03
GANs

Adversarial Training

Generator G maps noise z → synthetic samples. Discriminator D distinguishes real from fake. They play a minimax game.

GAN Objective
\min_G\max_D\;\mathbb{E}[\log D(x)]+\mathbb{E}[\log(1-D(G(z)))]
05 / 13
ML 101
M11 · L03
GAN Failure Modes

Instability & Mode Collapse

  • Mode collapse — G produces only a few samples; ignores diversity
  • Vanishing gradients — strong D → near-zero G gradients → G stops learning
  • WGAN-GP — Wasserstein distance + gradient penalty solves vanishing gradients
  • Spectral norm — enforces Lipschitz constraint on D without explicit penalty
  • Progressive growing — start at 4×4, scale up → stable 1024×1024 faces
06 / 13
ML 101
M11 · L03
Diffusion Models

Denoise to Generate

Gradually add Gaussian noise over T steps (forward). Train a U-Net to predict and remove noise at each step (reverse). Sample by starting from pure noise.

Training Loss
\mathcal{L}=\mathbb{E}_{t,x_0,\varepsilon}\|\varepsilon-\varepsilon_\theta(x_t,t)\|^2
07 / 13
ML 101
M11 · L03
Guidance

Classifier-Free Guidance

Randomly drop conditioning during training. At inference, extrapolate between conditional and unconditional predictions with guidance scale w.

Guidance Scale w
w=1: standard conditional. w=7–15: sharper adherence to prompt, less diversity. w→∞: maximally conditioned but may lose realism.
08 / 13
ML 101
M11 · L03
Stable Diffusion

Latent Diffusion

64×
Dimensionality reduction via VAE encoder
CLIP
Text embeddings injected via cross-attention

Diffusion runs in latent space, not pixel space. Decoder maps back to image. Consumer-GPU friendly. Powers Stable Diffusion, DALL-E 2, Imagen.

09 / 13
ML 101
M11 · L03
Normalizing Flows

Exact Likelihood

  • Invertible bijection f maps Gaussian noise ↔ data
  • Exact log-likelihood via change-of-variables + Jacobian log-det
  • Architectures: RealNVP, Glow, FFJORD (neural ODE)
  • Best for: density estimation, audio synthesis (WaveGlow)
  • Limitation: invertibility constraint limits expressiveness per layer
10 / 13
ML 101
M11 · L03
Evaluation

Measuring Generation Quality

  • FID — Fréchet distance between real & generated Inception features; lower = better
  • IS — sharpness × diversity via Inception; ignores distribution match
  • Precision / Recall — quality vs. diversity on real manifold
  • CLIP Score — text-image semantic alignment for T2I models
  • State-of-the-art diffusion: FID < 2 on ImageNet 256×256
11 / 13
ML 101
Generative Models

Check what generated

Four questions on VAEs, GANs and diffusion — a skimmer will pick the plausible wrong answer.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • VAEs: ELBO + reparameterization — stable, smooth latents, blurry outputs
  • GANs: minimax game — sharp samples, unstable training, mode collapse risk
  • Diffusion: reverse denoising — SOTA quality and diversity, slower sampling
  • Latent diffusion (Stable Diffusion): run in VAE latent space → consumer-GPU friendly
  • Classifier-free guidance: quality-diversity tradeoff controlled at inference
  • Flows: invertible bijections → exact likelihood; best for density estimation
13 / 13