Why adding noise can create images
Noise is usually the damage. Diffusion models use it as the raw material — and the thing they learn is almost embarrassingly modest.
Interactive Drag the noise level to add or remove noise, and pause to hold a frame.
The trick, in one sentence
Take a picture and add a little noise. Then a little more. Keep going until nothing is left but static. Now train a network on one small job only: given a slightly noisier picture, produce the slightly cleaner one. Run that backwards, starting from pure static, and you get a picture that never existed.
That is the whole idea. The model is never asked to imagine a cat from nothing. It is asked, thousands of times, to remove a tiny amount of grain — a problem so easy that a network can be mediocre at each step and still be excellent at all of them.
Why the easy version works
If every step adds only a small amount of Gaussian noise, you can make an assumption that pays for everything else: the way back is also Gaussian. So the reverse step does not need to be some unimaginable function — it just needs to predict a mean and a spread, and the network can be an ordinary one. This is why the authors could make the reverse process simple by construction rather than clever by architecture.
The algebra trick that makes training cheap
You might expect training to require simulating the whole noising chain, step by step, for every image. It does not. Because each step is Gaussian, the chain composes in closed form: you can jump straight to any noise level in a single expression, from one random draw.
So training looks like this: pick a random timestep, destroy the image directly to that level, and ask the network to predict the noise you just added. One sample, one shot, no simulation. The forward direction is free; all the work goes into the reverse direction.
What the network is really learning
Predicting the noise turns out to be the same as estimating the gradient of the log-density — the direction in which the data gets more likely. Sampling then looks like walking uphill on that landscape with a bit of shake to keep you honest. Diffusion and score matching are two descriptions of one object, and that equivalence is what let later work unify the whole family with stochastic differential equations.
Why nobody noticed for five years
The idea appeared in 2015 and mostly sat there. It was straightforward, it trained fine, and it produced samples that were not competitive with the generative models of the day. Ho, Jain and Abbeel showed in 2020 that a particular weighting of the training objective — with a connection to denoising score matching — made the samples excellent rather than merely valid. The architecture had not been the problem. The objective had.
One consequence is worth sitting with. Because these models spend much of their capacity on imperceptible detail, their sampling procedure doubles as a compression scheme: coarse structure first, fine grain last. Denoising turns out to be a way of ordering information — which is the part that makes it interesting far beyond pictures.
Where this comes from
- Denoising Diffusion Probabilistic Models linked only, not reproduced
arxiv.org/abs/2006.11239 - Score-Based Generative Modeling through Stochastic Differential Equations linked only, not reproduced
arxiv.org/abs/2011.13456 - Elucidating the Design Space of Diffusion-Based Generative Models linked only, not reproduced
arxiv.org/abs/2206.00364