Generative Models
VAEs, GANs, diffusion models, and normalizing flows — learning to model data distributions and generate new samples from scratch.
Discriminative vs. Generative
Discriminative models learn p(y|x) — given input, predict label. Generative models learn p(x) — the data distribution itself — enabling synthesis of new samples.
Variational Autoencoders
Encoder maps x → latent (μ, σ²). Sample z ~ N(μ, σ²) via reparameterization. Decoder reconstructs x from z. Optimize the ELBO.
Reparameterization Trick
Sampling z directly blocks gradients. Instead, sample ε ~ N(0, I) and compute z = μ + σ·ε. Stochasticity moves outside the computational graph.
Adversarial Training
Generator G maps noise z → synthetic samples. Discriminator D distinguishes real from fake. They play a minimax game.
Instability & Mode Collapse
- Mode collapse — G produces only a few samples; ignores diversity
- Vanishing gradients — strong D → near-zero G gradients → G stops learning
- WGAN-GP — Wasserstein distance + gradient penalty solves vanishing gradients
- Spectral norm — enforces Lipschitz constraint on D without explicit penalty
- Progressive growing — start at 4×4, scale up → stable 1024×1024 faces
Denoise to Generate
Gradually add Gaussian noise over T steps (forward). Train a U-Net to predict and remove noise at each step (reverse). Sample by starting from pure noise.
Classifier-Free Guidance
Randomly drop conditioning during training. At inference, extrapolate between conditional and unconditional predictions with guidance scale w.
Latent Diffusion
Diffusion runs in latent space, not pixel space. Decoder maps back to image. Consumer-GPU friendly. Powers Stable Diffusion, DALL-E 2, Imagen.
Exact Likelihood
- Invertible bijection f maps Gaussian noise ↔ data
- Exact log-likelihood via change-of-variables + Jacobian log-det
- Architectures: RealNVP, Glow, FFJORD (neural ODE)
- Best for: density estimation, audio synthesis (WaveGlow)
- Limitation: invertibility constraint limits expressiveness per layer
Measuring Generation Quality
- FID — Fréchet distance between real & generated Inception features; lower = better
- IS — sharpness × diversity via Inception; ignores distribution match
- Precision / Recall — quality vs. diversity on real manifold
- CLIP Score — text-image semantic alignment for T2I models
- State-of-the-art diffusion: FID < 2 on ImageNet 256×256
Key Takeaways
- VAEs: ELBO + reparameterization — stable, smooth latents, blurry outputs
- GANs: minimax game — sharp samples, unstable training, mode collapse risk
- Diffusion: reverse denoising — SOTA quality and diversity, slower sampling
- Latent diffusion (Stable Diffusion): run in VAE latent space → consumer-GPU friendly
- Classifier-free guidance: quality-diversity tradeoff controlled at inference
- Flows: invertible bijections → exact likelihood; best for density estimation