AI & MLDeep
Advanced

Diffusion Models & Image Generation

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Static Playbook
Static guide & reference tables

Generating images by learning to undo noise

Diffusion models are the technology behind Stable Diffusion, DALL-E, Midjourney, and most modern text-to-image and text-to-video systems. Their core idea sounds almost backwards: to learn how to *create* an image, first teach a model how to systematically *destroy* one — then train it to reverse that destruction, step by step, until it can conjure a coherent image out of pure random noise.

Forward diffusion: destroying an image

Forward diffusion is a fixed, non-learned process. Starting from a real training image, you add a small amount of Gaussian noise, then add a bit more, then more again, across many steps (often 1,000 in the original formulation) — until, at the final step, the image has been reduced to pure random noise with no trace of the original content left.

Mathematically, at each step $t$ you sample $x_t = \sqrt{1-\beta_t}\, x_{t-1} + \sqrt{\beta_t}\, \epsilon$, where $\epsilon$ is random Gaussian noise and $\beta_t$ is a small noise-schedule constant that increases over the course of the process. This gives you a dataset of (mostly clean image, slightly noisier image) pairs at every intermediate step — for free, without needing any labels, since you generated the noise yourself.

Reverse diffusion: learning to denoise

The actual neural network — usually a U-Net or, in newer models, a transformer-based architecture — is trained to do the opposite: given a noisy image at step $t$, predict the noise that was added, so it can be subtracted to recover a slightly cleaner image at step $t-1$.

Crucially, the model isn't trained to jump straight from noise to a finished image in one shot — that's too large a leap to learn reliably. It's trained on the much easier task of removing *one small increment* of noise at a time. Chain that denoising step a thousand times, starting from pure random noise, and you end up with a coherent image — one the model effectively "hallucinated" into existence by repeatedly guessing what noise to remove.

text

Note

In text-to-image diffusion, the random starting noise is unconditioned — it's just static. What makes the output match your prompt is that the text embedding (usually from a model like CLIP) is fed into the denoising network at every step, biasing each noise-removal prediction toward content consistent with the prompt. Change the prompt mid-generation and you're literally steering which image emerges from the same starting noise.

Why Stable Diffusion works in a compressed 'latent' space

Running the full denoising process directly on full-resolution pixel images (say, 512×512×3 values) at every one of hundreds of steps is computationally brutal. Stable Diffusion's key efficiency trick is to run diffusion in a much smaller latent space instead: a separate autoencoder first compresses the image into a compact representation (e.g. 64×64×4), diffusion happens entirely in that compressed space, and only at the very end does a decoder expand the final denoised latent back into a full-resolution pixel image.

This is why Stable Diffusion can run on a consumer GPU while comparable pixel-space diffusion models often can't — the expensive iterative part of the process operates on a representation dozens of times smaller than the final image.

DALL-E and the broader diffusion family

DALL-E 2 and 3 use a related but distinct pipeline: a text encoder produces an embedding of the prompt, a "prior" model maps that text embedding to an image embedding (in DALL-E 2's case, using CLIP's shared embedding space), and a diffusion decoder generates the final image conditioned on that image embedding. DALL-E 3 additionally leans heavily on highly detailed, model-generated captions during training to improve prompt adherence.

Across the whole family — Stable Diffusion, DALL-E, Imagen, Midjourney's underlying architecture — the shared foundation is the same forward/reverse diffusion process; the differences are mostly in what space diffusion happens in, how text conditioning is injected, and what data the model was trained on.

Diffusion vs. earlier generative approaches

ApproachHow it generatesTraining stabilityTypical output quality
Diffusion modelsIteratively denoise from random noiseHigh — simple regression lossState of the art, high diversity
GANsGenerator competes against a discriminatorLow — prone to mode collapseSharp but less diverse
VAEsEncode to latent space, decode backHighBlurrier, less detailed

What's next

Diffusion models are one branch of the generative modeling family tree; GANs are the other major branch and took a very different, adversarial approach to the same problem. See the GANs lesson next to understand why diffusion models largely overtook them for image generation.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.