AI & MLDeep
Advanced

Generative Adversarial Networks (GANs)

10 min read

Learn
Deep Reading
Estimated 10 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Static Playbook
Static guide & reference tables

Two networks playing an adversarial game

A Generative Adversarial Network (GAN), introduced by Ian Goodfellow in 2014, trains two neural networks against each other in a competitive game. The generator tries to produce fake data (typically images) realistic enough to fool a critic. The discriminator tries to correctly tell real training data apart from the generator's fakes. Each network improves by trying to beat the other, and — in theory — this adversarial pressure drives the generator toward producing increasingly realistic output.

This was a genuinely novel training paradigm: instead of a fixed loss function measuring pixel-level similarity to a target, the "judge" of quality is itself a trained neural network that keeps getting better at spotting fakes, forcing the generator to keep improving to keep up.

The minimax game

Formally, the generator $G$ takes random noise $z$ and outputs a fake sample $G(z)$. The discriminator $D$ takes a sample (real or fake) and outputs the probability it's real. Training alternates between two objectives:

The discriminator is trained to maximize $\log D(x) + \log(1 - D(G(z)))$ — correctly classify real samples $x$ as real and fake samples $G(z)$ as fake.

The generator is trained to minimize $\log(1 - D(G(z)))$ — in other words, to fool the discriminator into outputting a high probability that its fakes are real.

This is a minimax game: the generator's win is the discriminator's loss, and vice versa. At the theoretical equilibrium, the generator produces samples indistinguishable from real data, and the discriminator can do no better than random guessing (50/50) on any input.

python

Note

Unlike a standard supervised loss that monotonically improves, GAN training is a dynamic equilibrium between two competing networks — and it doesn't always converge cleanly. If the discriminator gets too good too fast, the generator's gradients vanish (it gets no useful signal on how to improve). If the generator gets ahead, the discriminator's feedback becomes meaningless. Getting GAN training to actually stabilize was, for years, as much an empirical art as a science, involving tricks like gradient penalties, spectral normalization, and carefully balanced learning rates.

Mode collapse: the signature GAN failure

Mode collapse is the most characteristic GAN failure mode: instead of learning to generate the full diversity of the training distribution, the generator discovers a small number of outputs (sometimes just one) that reliably fool the current discriminator, and collapses to producing near-identical variations of those outputs regardless of the input noise.

This happens because the generator's only objective is to fool the discriminator — it has no explicit incentive to cover the full diversity of real data. If one type of fake consistently works, gradient descent will happily exploit it repeatedly rather than exploring other modes of the data distribution. Detecting mode collapse in practice usually means noticing that a large fraction of generated samples look suspiciously similar despite different input noise vectors.

Notable GAN architectures

DCGAN (Deep Convolutional GAN) established the standard convolutional architecture patterns that made GAN training more reliable for images. StyleGAN (and StyleGAN2/3) introduced a style-based generator architecture that separates high-level attributes (pose, identity) from fine details (skin texture, hair), enabling remarkably realistic and controllable face generation — it's the technology behind sites like "this person does not exist." CycleGAN enabled image-to-image translation without paired training data (e.g. turning horse photos into zebra photos) using a cycle-consistency loss.

GANs vs. diffusion models

PropertyGANsDiffusion models
Training stabilityUnstable — adversarial equilibriumStable — simple denoising regression loss
Output diversityProne to mode collapseHigh diversity, covers the data distribution well
Generation speedFast — single forward passSlow — many iterative denoising steps
Sample quality (current SOTA)Sharp but less consistentState of the art for most image tasks
Training data efficiencyCan work with less dataTypically needs large datasets

Why diffusion models displaced GANs for image generation

Through the early 2020s, diffusion models overtook GANs as the dominant approach for state-of-the-art image generation (Stable Diffusion, DALL-E, Midjourney's newer versions all use diffusion). The main reasons: diffusion training is far more stable — there's no adversarial dynamic to destabilize, just a straightforward denoising prediction loss — and diffusion models tend not to suffer from mode collapse, since they're explicitly trained to reconstruct the true data distribution rather than merely to fool a discriminator.

GANs remain relevant where their strengths matter most: fast single-pass generation (useful for real-time applications), and tasks like super-resolution or specific style transfer where their sharper output characteristics are an advantage.

What's next

For the approach that has largely superseded GANs in image generation, see Diffusion Models & Image Generation. To go deeper on the underlying neural network building blocks, see Convolutional Neural Networks (CNNs), which both GANs and diffusion models rely on heavily for image processing.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.