Home / ML 101 / Module 11 / Lesson 3

Generative Models

VAEs, GANs, diffusion models, and normalizing flows — learning to model and sample from data distributions to generate new images, text, audio, and more.

~20 min read M11 · L3 Advanced

Generative vs. Discriminative Models

Most supervised learning models are discriminative: they learn a conditional distribution p(y|x) — given an input x, predict a label y. Generative models take a fundamentally different approach: they learn the joint distribution p(x) or p(x, y) of the data itself. Once learned, this distribution can be sampled to produce new data points that resemble the training set.

Why does this matter? Generative models unlock capabilities beyond classification: synthesizing realistic images, completing missing data, augmenting training sets, learning compressed representations, and enabling creative AI applications in art, music, drug discovery, and scientific simulation. The central challenge is that high-dimensional data (images, audio, text) lives on a complex manifold — learning its distribution requires powerful inductive biases.

The Core Challenge

A 256×256 RGB image is a point in a 196,608-dimensional space. The set of natural images is an astronomically tiny, highly structured subset of this space. Generative models must learn to concentrate probability mass precisely on this natural image manifold — without seeing it explicitly defined.

Variational Autoencoders (VAEs)

A Variational Autoencoder learns a latent variable model: data x is assumed to be generated from a low-dimensional latent vector z drawn from a prior p(z) = N(0, I), passed through a decoder p(x|z). The encoder q(z|x) approximates the intractable posterior p(z|x) with a Gaussian parameterized by mean μ and variance σ².

The ELBO Objective

Because the exact marginal log-likelihood log p(x) is intractable, VAEs optimize the Evidence Lower Bound (ELBO):

VAE ELBO
\mathcal{L}(\theta,\phi) = \mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{\mathrm{KL}}(q_\phi(z|x)\,\|\,p(z))
The reconstruction term (first) encourages the decoder to recover x from z; the KL term regularizes the encoder to stay close to the standard Gaussian prior. Maximizing the ELBO tightens the lower bound on log p(x).

The Reparameterization Trick

Sampling z ~ q(z|x) blocks gradient flow — the sample operation is non-differentiable. The reparameterization trick sidesteps this: instead of sampling z directly, sample ε ~ N(0, I) and compute z = μ + σ·ε. The stochasticity is moved outside the computational graph, so gradients flow through μ and σ to the encoder parameters.

Reparameterization
z = \mu_\phi(x) + \sigma_\phi(x) \cdot \varepsilon, \quad \varepsilon \sim \mathcal{N}(0, I)
ε is sampled from a fixed standard Gaussian; μ and σ are deterministic functions of x output by the encoder network. Gradients flow through μ and σ normally.

VAEs produce smooth, structured latent spaces — interpolating between two latent codes produces semantically meaningful intermediate samples. However, VAE samples are often blurry because the pixel-wise reconstruction loss (MSE or BCE) encourages averaging over uncertainty. This is the main quality limitation compared to GANs and diffusion models.

Generative Adversarial Networks (GANs)

Introduced by Goodfellow et al. in 2014, GANs frame generation as a two-player minimax game between a generator G and a discriminator D. The generator maps noise z ~ p(z) to synthetic samples G(z); the discriminator learns to distinguish real data from generated fakes. They are trained alternately in opposition.

GAN Minimax Objective
\min_G \max_D\; \mathbb{E}_{x\sim p_{\text{data}}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))]
D maximizes this expression (distinguishing real from fake); G minimizes it (fooling D). At Nash equilibrium, G produces the true data distribution and D outputs 0.5 everywhere.

Training Dynamics and Failure Modes

GAN training is notoriously unstable. Key failure modes:

Stabilization Techniques

Wasserstein GAN
Earth Mover Distance
Replaces JS divergence with Wasserstein-1 distance. Provides meaningful gradients everywhere; requires weight clipping or gradient penalty (WGAN-GP).
Spectral Norm
Lipschitz constraint
Normalizes discriminator weights by their spectral norm, enforcing a Lipschitz constraint without explicit gradient penalty. Stabilizes training significantly.
Progressive Growing
Curriculum training
Starts with low-resolution images (4×4) and progressively adds higher-resolution layers. Used by ProGAN to produce photorealistic faces at 1024×1024.
Conditional GAN
Conditional generation
Condition both G and D on a class label or text embedding. Enables controlled generation — synthesizing specific classes or text-guided images.

StyleGAN

StyleGAN (NVIDIA, 2019) fundamentally redesigned the GAN generator. Instead of feeding noise directly into the input, an 8-layer MLP maps z to an intermediate latent space W, which then modulates each layer of the synthesis network via adaptive instance normalization (AdaIN). This disentangles style control — coarse features (pose, shape) are controlled by early layers, fine details (texture, color) by later layers. StyleGAN2/3 further removed artifacts and improved alias-free synthesis.

Diffusion Models

Denoising Diffusion Probabilistic Models (DDPMs), introduced by Ho et al. (2020), have become the dominant generative paradigm for images and audio, powering Stable Diffusion, DALL-E 2, Imagen, and Sora. They define a forward process that gradually adds Gaussian noise to data over T steps, then train a neural network to reverse this process.

Forward Process

The forward process q(x_t | x_{t-1}) adds Gaussian noise at each step according to a variance schedule β_1, ..., β_T. By using the closed-form marginal, any noisy version x_t can be sampled directly from x_0 without stepping through all t-1 intermediate steps:

Forward Process (closed form)
q(x_t|x_0) = \mathcal{N}\!\left(x_t;\,\sqrt{\bar\alpha_t}\,x_0,\,(1-\bar\alpha_t)I\right)
ᾱ_t = ∏_{s=1}^{t}(1−β_s) is the cumulative noise product. As t → T, x_t converges to pure Gaussian noise. This allows training on arbitrary noise levels.

Reverse Process and Training

The reverse process p_θ(x_{t-1}|x_t) is approximated by a U-Net (or transformer) that takes noisy x_t and timestep t as input and predicts the noise ε added during the forward step. The simplified training objective is:

Diffusion Training Loss
\mathcal{L} = \mathbb{E}_{t,x_0,\varepsilon}\!\left[\|\varepsilon - \varepsilon_\theta(x_t, t)\|^2\right]
The network ε_θ predicts the noise ε from noisy x_t. MSE between predicted and actual noise. This simple objective has proven highly effective across image, audio, and video domains.

Guidance and Conditioning

Classifier-free guidance (CFG) is the dominant technique for conditioning diffusion models on text or class labels without a separate classifier. During training, the conditioning signal (e.g., text embedding) is randomly dropped with probability p, training the model as both conditional and unconditional. At inference, the predicted noise is extrapolated between the conditional and unconditional predictions:

Classifier-Free Guidance
\tilde\varepsilon_\theta(x_t, c) = \varepsilon_\theta(x_t) + w\cdot(\varepsilon_\theta(x_t, c) - \varepsilon_\theta(x_t))
w is the guidance scale (typically 7–15). Higher w sharpens adherence to the condition but reduces diversity. w=1 recovers standard conditional generation.

Latent Diffusion Models

Latent Diffusion Models (LDMs), the architecture behind Stable Diffusion, reduce computational cost by operating in the latent space of a pretrained VAE rather than pixel space. The diffusion process runs on compressed 64×64 latents instead of 512×512 pixels — a 64× reduction in dimensionality. A cross-attention mechanism in the denoising U-Net injects text conditioning at every layer via CLIP text embeddings.

Normalizing Flows

Normalizing flows define a generative model via an invertible, differentiable mapping f between a simple base distribution (e.g., Gaussian) and the complex data distribution. By stacking multiple invertible transformations, flows can model arbitrarily complex distributions while maintaining exact likelihood computation — a unique advantage over VAEs and GANs.

Change of Variables
\log p_X(x) = \log p_Z(f(x)) + \log\left|\det\frac{\partial f}{\partial x}\right|
The log-determinant of the Jacobian accounts for the volume change under transformation f. Flow architectures (RealNVP, Glow, FFJORD) are designed so this determinant is cheap to compute.

The constraint that f must be invertible and have a tractable Jacobian limits the expressive power of individual flow layers, requiring many layers to model complex distributions. Modern continuous normalizing flows (CNFs) defined via neural ODEs remove this constraint but are expensive to integrate numerically. Flows are most successful in audio synthesis (WaveGlow, WaveFlow) and density estimation tasks.

Evaluation Metrics

Evaluating generative models is notoriously difficult — perceptual quality, diversity, and distribution coverage are all important, but no single metric captures all three.

MetricWhat It MeasuresLimitation
FID (Fréchet Inception Distance)Distance between real and generated feature distributions (Inception-v3)Sensitive to sample size; may not correlate with human perception
IS (Inception Score)Sharpness × diversity of generated samplesIgnores whether generated samples match the training distribution
Precision / RecallPrecision = quality; Recall = diversity relative to real data manifoldComputationally expensive; depends on manifold estimation
CLIP ScoreSemantic alignment between generated image and text promptOnly relevant for text-conditional generation
Log-likelihoodExact for flows; ELBO bound for VAEs; inapplicable to GANsHigh likelihood ≠ high quality or diversity

FID is the de facto standard for unconditional and class-conditional image generation benchmarks. Lower is better — state-of-the-art diffusion models achieve FID < 2 on ImageNet 256×256.

Comparison of Generative Families

Model FamilySample QualityDiversityTraining StabilityExact Likelihood
VAEModerate (blurry)HighStableELBO lower bound
GANHighMode collapse riskUnstableNo
DiffusionVery highHighStableELBO lower bound
Normalizing FlowModerateHighStableYes

Applications

Image Synthesis
Text-to-image
Stable Diffusion, DALL-E 3, Midjourney — photorealistic and artistic image generation from text prompts using LDMs.
Audio Generation
Speech & music
WaveNet, AudioLM, MusicGen — generating realistic speech and music via autoregressive or diffusion models in the waveform or spectrogram domain.
Video Generation
Temporal coherence
Sora, Stable Video Diffusion — extending spatial diffusion to temporal sequences with 3D U-Nets or video transformers. Coherence across frames is the key challenge.
Drug Discovery
Molecular design
RFDiffusion, DiffSBDD — diffusion models over 3D protein and molecular structures. Generating novel molecules with desired binding properties or protein folds.
Data Augmentation
Synthetic training data
Generating synthetic labeled examples for rare classes or privacy-sensitive domains (medical imaging). Can improve downstream classifier performance.
Inpainting & Editing
Conditional generation
Filling in masked regions, style transfer, and semantic editing via conditional diffusion models (SDEdit, InstructPix2Pix).

Key Takeaways

Generative models learn the distribution of data to synthesize new samples. VAEs use an ELBO objective with the reparameterization trick — stable training but blurry outputs. GANs formulate a minimax game — sharp samples but unstable training and mode collapse risk. Diffusion models learn to reverse a gradual noising process — state-of-the-art quality and diversity, but slower sampling. Normalizing flows are invertible bijections enabling exact likelihood. FID is the standard evaluation metric. Latent diffusion (Stable Diffusion) brings diffusion-quality generation to consumer hardware by operating in VAE latent space. Classifier-free guidance controls the quality-diversity tradeoff at inference time.

Previous M11-L2: Transfer Learning Module Overview Next Lesson M12-L1: Responsible ML