Generative Adversarial Networks (GANs)
Two networks playing an adversarial game
A Generative Adversarial Network (GAN), introduced by Ian Goodfellow in 2014, trains two neural networks against each other in a competitive game. The generator tries to produce fake data (typically images) realistic enough to fool a critic. The discriminator tries to correctly tell real training data apart from the generator's fakes. Each network improves by trying to beat the other, and — in theory — this adversarial pressure drives the generator toward producing increasingly realistic output.
This was a genuinely novel training paradigm: instead of a fixed loss function measuring pixel-level similarity to a target, the "judge" of quality is itself a trained neural network that keeps getting better at spotting fakes, forcing the generator to keep improving to keep up.
The minimax game
Formally, the generator $G$ takes random noise $z$ and outputs a fake sample $G(z)$. The discriminator $D$ takes a sample (real or fake) and outputs the probability it's real. Training alternates between two objectives:
The discriminator is trained to maximize $\log D(x) + \log(1 - D(G(z)))$ — correctly classify real samples $x$ as real and fake samples $G(z)$ as fake.
The generator is trained to minimize $\log(1 - D(G(z)))$ — in other words, to fool the discriminator into outputting a high probability that its fakes are real.
This is a minimax game: the generator's win is the discriminator's loss, and vice versa. At the theoretical equilibrium, the generator produces samples indistinguishable from real data, and the discriminator can do no better than random guessing (50/50) on any input.
Note
Mode collapse: the signature GAN failure
Mode collapse is the most characteristic GAN failure mode: instead of learning to generate the full diversity of the training distribution, the generator discovers a small number of outputs (sometimes just one) that reliably fool the current discriminator, and collapses to producing near-identical variations of those outputs regardless of the input noise.
This happens because the generator's only objective is to fool the discriminator — it has no explicit incentive to cover the full diversity of real data. If one type of fake consistently works, gradient descent will happily exploit it repeatedly rather than exploring other modes of the data distribution. Detecting mode collapse in practice usually means noticing that a large fraction of generated samples look suspiciously similar despite different input noise vectors.
Notable GAN architectures
DCGAN (Deep Convolutional GAN) established the standard convolutional architecture patterns that made GAN training more reliable for images. StyleGAN (and StyleGAN2/3) introduced a style-based generator architecture that separates high-level attributes (pose, identity) from fine details (skin texture, hair), enabling remarkably realistic and controllable face generation — it's the technology behind sites like "this person does not exist." CycleGAN enabled image-to-image translation without paired training data (e.g. turning horse photos into zebra photos) using a cycle-consistency loss.
GANs vs. diffusion models
| Property | GANs | Diffusion models |
|---|---|---|
| Training stability | Unstable — adversarial equilibrium | Stable — simple denoising regression loss |
| Output diversity | Prone to mode collapse | High diversity, covers the data distribution well |
| Generation speed | Fast — single forward pass | Slow — many iterative denoising steps |
| Sample quality (current SOTA) | Sharp but less consistent | State of the art for most image tasks |
| Training data efficiency | Can work with less data | Typically needs large datasets |
Why diffusion models displaced GANs for image generation
Through the early 2020s, diffusion models overtook GANs as the dominant approach for state-of-the-art image generation (Stable Diffusion, DALL-E, Midjourney's newer versions all use diffusion). The main reasons: diffusion training is far more stable — there's no adversarial dynamic to destabilize, just a straightforward denoising prediction loss — and diffusion models tend not to suffer from mode collapse, since they're explicitly trained to reconstruct the true data distribution rather than merely to fool a discriminator.
GANs remain relevant where their strengths matter most: fast single-pass generation (useful for real-time applications), and tasks like super-resolution or specific style transfer where their sharper output characteristics are an advantage.
What's next
For the approach that has largely superseded GANs in image generation, see Diffusion Models & Image Generation. To go deeper on the underlying neural network building blocks, see Convolutional Neural Networks (CNNs), which both GANs and diffusion models rely on heavily for image processing.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.