Diffusion Models & Image Generation
Generating images by learning to undo noise
Diffusion models are the technology behind Stable Diffusion, DALL-E, Midjourney, and most modern text-to-image and text-to-video systems. Their core idea sounds almost backwards: to learn how to *create* an image, first teach a model how to systematically *destroy* one — then train it to reverse that destruction, step by step, until it can conjure a coherent image out of pure random noise.
Forward diffusion: destroying an image
Forward diffusion is a fixed, non-learned process. Starting from a real training image, you add a small amount of Gaussian noise, then add a bit more, then more again, across many steps (often 1,000 in the original formulation) — until, at the final step, the image has been reduced to pure random noise with no trace of the original content left.
Mathematically, at each step $t$ you sample $x_t = \sqrt{1-\beta_t}\, x_{t-1} + \sqrt{\beta_t}\, \epsilon$, where $\epsilon$ is random Gaussian noise and $\beta_t$ is a small noise-schedule constant that increases over the course of the process. This gives you a dataset of (mostly clean image, slightly noisier image) pairs at every intermediate step — for free, without needing any labels, since you generated the noise yourself.
Reverse diffusion: learning to denoise
The actual neural network — usually a U-Net or, in newer models, a transformer-based architecture — is trained to do the opposite: given a noisy image at step $t$, predict the noise that was added, so it can be subtracted to recover a slightly cleaner image at step $t-1$.
Crucially, the model isn't trained to jump straight from noise to a finished image in one shot — that's too large a leap to learn reliably. It's trained on the much easier task of removing *one small increment* of noise at a time. Chain that denoising step a thousand times, starting from pure random noise, and you end up with a coherent image — one the model effectively "hallucinated" into existence by repeatedly guessing what noise to remove.
Note
Why Stable Diffusion works in a compressed 'latent' space
Running the full denoising process directly on full-resolution pixel images (say, 512×512×3 values) at every one of hundreds of steps is computationally brutal. Stable Diffusion's key efficiency trick is to run diffusion in a much smaller latent space instead: a separate autoencoder first compresses the image into a compact representation (e.g. 64×64×4), diffusion happens entirely in that compressed space, and only at the very end does a decoder expand the final denoised latent back into a full-resolution pixel image.
This is why Stable Diffusion can run on a consumer GPU while comparable pixel-space diffusion models often can't — the expensive iterative part of the process operates on a representation dozens of times smaller than the final image.
DALL-E and the broader diffusion family
DALL-E 2 and 3 use a related but distinct pipeline: a text encoder produces an embedding of the prompt, a "prior" model maps that text embedding to an image embedding (in DALL-E 2's case, using CLIP's shared embedding space), and a diffusion decoder generates the final image conditioned on that image embedding. DALL-E 3 additionally leans heavily on highly detailed, model-generated captions during training to improve prompt adherence.
Across the whole family — Stable Diffusion, DALL-E, Imagen, Midjourney's underlying architecture — the shared foundation is the same forward/reverse diffusion process; the differences are mostly in what space diffusion happens in, how text conditioning is injected, and what data the model was trained on.
Diffusion vs. earlier generative approaches
| Approach | How it generates | Training stability | Typical output quality |
|---|---|---|---|
| Diffusion models | Iteratively denoise from random noise | High — simple regression loss | State of the art, high diversity |
| GANs | Generator competes against a discriminator | Low — prone to mode collapse | Sharp but less diverse |
| VAEs | Encode to latent space, decode back | High | Blurrier, less detailed |
What's next
Diffusion models are one branch of the generative modeling family tree; GANs are the other major branch and took a very different, adversarial approach to the same problem. See the GANs lesson next to understand why diffusion models largely overtook them for image generation.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.