Generative vs. Discriminative Models
Most supervised learning models are discriminative: they learn a conditional distribution p(y|x) — given an input x, predict a label y. Generative models take a fundamentally different approach: they learn the joint distribution p(x) or p(x, y) of the data itself. Once learned, this distribution can be sampled to produce new data points that resemble the training set.
Why does this matter? Generative models unlock capabilities beyond classification: synthesizing realistic images, completing missing data, augmenting training sets, learning compressed representations, and enabling creative AI applications in art, music, drug discovery, and scientific simulation. The central challenge is that high-dimensional data (images, audio, text) lives on a complex manifold — learning its distribution requires powerful inductive biases.
A 256×256 RGB image is a point in a 196,608-dimensional space. The set of natural images is an astronomically tiny, highly structured subset of this space. Generative models must learn to concentrate probability mass precisely on this natural image manifold — without seeing it explicitly defined.
Variational Autoencoders (VAEs)
A Variational Autoencoder learns a latent variable model: data x is assumed to be generated from a low-dimensional latent vector z drawn from a prior p(z) = N(0, I), passed through a decoder p(x|z). The encoder q(z|x) approximates the intractable posterior p(z|x) with a Gaussian parameterized by mean μ and variance σ².
The ELBO Objective
Because the exact marginal log-likelihood log p(x) is intractable, VAEs optimize the Evidence Lower Bound (ELBO):
The Reparameterization Trick
Sampling z ~ q(z|x) blocks gradient flow — the sample operation is non-differentiable. The reparameterization trick sidesteps this: instead of sampling z directly, sample ε ~ N(0, I) and compute z = μ + σ·ε. The stochasticity is moved outside the computational graph, so gradients flow through μ and σ to the encoder parameters.
VAEs produce smooth, structured latent spaces — interpolating between two latent codes produces semantically meaningful intermediate samples. However, VAE samples are often blurry because the pixel-wise reconstruction loss (MSE or BCE) encourages averaging over uncertainty. This is the main quality limitation compared to GANs and diffusion models.
Generative Adversarial Networks (GANs)
Introduced by Goodfellow et al. in 2014, GANs frame generation as a two-player minimax game between a generator G and a discriminator D. The generator maps noise z ~ p(z) to synthetic samples G(z); the discriminator learns to distinguish real data from generated fakes. They are trained alternately in opposition.
Training Dynamics and Failure Modes
GAN training is notoriously unstable. Key failure modes:
- Mode collapse — the generator learns to produce only a few high-quality samples (or even one) that fool the discriminator, ignoring the full data distribution
- Vanishing gradients — if D becomes too strong early, G receives near-zero gradients and stops learning
- Oscillation — neither G nor D converges; losses cycle without reaching equilibrium
Stabilization Techniques
StyleGAN
StyleGAN (NVIDIA, 2019) fundamentally redesigned the GAN generator. Instead of feeding noise directly into the input, an 8-layer MLP maps z to an intermediate latent space W, which then modulates each layer of the synthesis network via adaptive instance normalization (AdaIN). This disentangles style control — coarse features (pose, shape) are controlled by early layers, fine details (texture, color) by later layers. StyleGAN2/3 further removed artifacts and improved alias-free synthesis.
Diffusion Models
Denoising Diffusion Probabilistic Models (DDPMs), introduced by Ho et al. (2020), have become the dominant generative paradigm for images and audio, powering Stable Diffusion, DALL-E 2, Imagen, and Sora. They define a forward process that gradually adds Gaussian noise to data over T steps, then train a neural network to reverse this process.
Forward Process
The forward process q(x_t | x_{t-1}) adds Gaussian noise at each step according to a variance schedule β_1, ..., β_T. By using the closed-form marginal, any noisy version x_t can be sampled directly from x_0 without stepping through all t-1 intermediate steps:
Reverse Process and Training
The reverse process p_θ(x_{t-1}|x_t) is approximated by a U-Net (or transformer) that takes noisy x_t and timestep t as input and predicts the noise ε added during the forward step. The simplified training objective is:
Guidance and Conditioning
Classifier-free guidance (CFG) is the dominant technique for conditioning diffusion models on text or class labels without a separate classifier. During training, the conditioning signal (e.g., text embedding) is randomly dropped with probability p, training the model as both conditional and unconditional. At inference, the predicted noise is extrapolated between the conditional and unconditional predictions:
Latent Diffusion Models
Latent Diffusion Models (LDMs), the architecture behind Stable Diffusion, reduce computational cost by operating in the latent space of a pretrained VAE rather than pixel space. The diffusion process runs on compressed 64×64 latents instead of 512×512 pixels — a 64× reduction in dimensionality. A cross-attention mechanism in the denoising U-Net injects text conditioning at every layer via CLIP text embeddings.
Normalizing Flows
Normalizing flows define a generative model via an invertible, differentiable mapping f between a simple base distribution (e.g., Gaussian) and the complex data distribution. By stacking multiple invertible transformations, flows can model arbitrarily complex distributions while maintaining exact likelihood computation — a unique advantage over VAEs and GANs.
The constraint that f must be invertible and have a tractable Jacobian limits the expressive power of individual flow layers, requiring many layers to model complex distributions. Modern continuous normalizing flows (CNFs) defined via neural ODEs remove this constraint but are expensive to integrate numerically. Flows are most successful in audio synthesis (WaveGlow, WaveFlow) and density estimation tasks.
Evaluation Metrics
Evaluating generative models is notoriously difficult — perceptual quality, diversity, and distribution coverage are all important, but no single metric captures all three.
| Metric | What It Measures | Limitation |
|---|---|---|
| FID (Fréchet Inception Distance) | Distance between real and generated feature distributions (Inception-v3) | Sensitive to sample size; may not correlate with human perception |
| IS (Inception Score) | Sharpness × diversity of generated samples | Ignores whether generated samples match the training distribution |
| Precision / Recall | Precision = quality; Recall = diversity relative to real data manifold | Computationally expensive; depends on manifold estimation |
| CLIP Score | Semantic alignment between generated image and text prompt | Only relevant for text-conditional generation |
| Log-likelihood | Exact for flows; ELBO bound for VAEs; inapplicable to GANs | High likelihood ≠ high quality or diversity |
FID is the de facto standard for unconditional and class-conditional image generation benchmarks. Lower is better — state-of-the-art diffusion models achieve FID < 2 on ImageNet 256×256.
Comparison of Generative Families
| Model Family | Sample Quality | Diversity | Training Stability | Exact Likelihood |
|---|---|---|---|---|
| VAE | Moderate (blurry) | High | Stable | ELBO lower bound |
| GAN | High | Mode collapse risk | Unstable | No |
| Diffusion | Very high | High | Stable | ELBO lower bound |
| Normalizing Flow | Moderate | High | Stable | Yes |
Applications
Generative models learn the distribution of data to synthesize new samples. VAEs use an ELBO objective with the reparameterization trick — stable training but blurry outputs. GANs formulate a minimax game — sharp samples but unstable training and mode collapse risk. Diffusion models learn to reverse a gradual noising process — state-of-the-art quality and diversity, but slower sampling. Normalizing flows are invertible bijections enabling exact likelihood. FID is the standard evaluation metric. Latent diffusion (Stable Diffusion) brings diffusion-quality generation to consumer hardware by operating in VAE latent space. Classifier-free guidance controls the quality-diversity tradeoff at inference time.