№ 01 · diffusion models

Why adding noise can create images

Noise is usually the damage. Diffusion models use it as the raw material — and the thing they learn is almost embarrassingly modest.

Interactive Drag the noise level to add or remove noise, and pause to hold a frame.

t = 0.00 · clean signal
Each dot is a piece of the picture. Reading right, small Gaussian steps destroy the order until only static is left; reading back, the network only ever removes a little noise at a time, and the structure returns. Nobody told it to draw rings — that is the accumulated effect of many small denoising decisions. Drag to scrub.

The trick, in one sentence

Take a picture and add a little noise. Then a little more. Keep going until nothing is left but static. Now train a network on one small job only: given a slightly noisier picture, produce the slightly cleaner one. Run that backwards, starting from pure static, and you get a picture that never existed.

That is the whole idea. The model is never asked to imagine a cat from nothing. It is asked, thousands of times, to remove a tiny amount of grain — a problem so easy that a network can be mediocre at each step and still be excellent at all of them.

Why the easy version works

If every step adds only a small amount of Gaussian noise, you can make an assumption that pays for everything else: the way back is also Gaussian. So the reverse step does not need to be some unimaginable function — it just needs to predict a mean and a spread, and the network can be an ordinary one. This is why the authors could make the reverse process simple by construction rather than clever by architecture.

The algebra trick that makes training cheap

You might expect training to require simulating the whole noising chain, step by step, for every image. It does not. Because each step is Gaussian, the chain composes in closed form: you can jump straight to any noise level in a single expression, from one random draw.

So training looks like this: pick a random timestep, destroy the image directly to that level, and ask the network to predict the noise you just added. One sample, one shot, no simulation. The forward direction is free; all the work goes into the reverse direction.

Two minutes twenty-eight. The same idea, moving, drawn frame by frame in code — the figure above, told as a story with a beginning and an end. Nothing sets that length but the script: the film is sized to the narration it carries, so an idea that needs more room gets more room rather than being compressed.

What the network is really learning

Predicting the noise turns out to be the same as estimating the gradient of the log-density — the direction in which the data gets more likely. Sampling then looks like walking uphill on that landscape with a bit of shake to keep you honest. Diffusion and score matching are two descriptions of one object, and that equivalence is what let later work unify the whole family with stochastic differential equations.

Why nobody noticed for five years

The idea appeared in 2015 and mostly sat there. It was straightforward, it trained fine, and it produced samples that were not competitive with the generative models of the day. Ho, Jain and Abbeel showed in 2020 that a particular weighting of the training objective — with a connection to denoising score matching — made the samples excellent rather than merely valid. The architecture had not been the problem. The objective had.

One consequence is worth sitting with. Because these models spend much of their capacity on imperceptible detail, their sampling procedure doubles as a compression scheme: coarse structure first, fine grain last. Denoising turns out to be a way of ordering information — which is the part that makes it interesting far beyond pictures.

Where this comes from

  1. Denoising Diffusion Probabilistic Models linked only, not reproduced
    Jonathan Ho, Ajay Jain, Pieter Abbeel · arXiv:2006.11239 · 2020
    arxiv.org/abs/2006.11239
  2. Score-Based Generative Modeling through Stochastic Differential Equations linked only, not reproduced
    Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon and Ben Poole · arXiv:2011.13456 · 2020
    arxiv.org/abs/2011.13456
  3. Elucidating the Design Space of Diffusion-Based Generative Models linked only, not reproduced
    Tero Karras, Miika Aittala, Timo Aila, Samuli Laine · arXiv:2206.00364 · 2022
    arxiv.org/abs/2206.00364