Why a noisy bottleneck can invent new data
An ordinary autoencoder squeezes data to a short code and learns to rebuild it. A variational autoencoder does the same with one change: the code is noisy on purpose. That change lets you draw a random code afterwards and get plausible new data back out.
A bottleneck with a spread
An autoencoder is two networks in a row. The encoder turns an input — say an image — into a short list of numbers, a code. The decoder turns it back into an image. Trained together to make output match input, the code becomes a compressed summary. In a variational autoencoder the encoder outputs not a single code but a mean and a standard deviation — the centre and width of a bell curve — and the code fed to the decoder is a random draw from that curve. Kingma and Welling call this a probabilistic encoder: given a datapoint, a distribution over the codes it could have come from.
Why it matters
A plain autoencoder compresses well and cannot invent. Its codes land wherever training put them, with empty space between; a code from that empty space decodes to junk, because nothing was ever decoded from there. The variational version fills it. Kingma and Welling's 2013 paper made that trainable with ordinary gradient descent, and the paper's samples — digits and faces decoded from codes drawn from the prior — showed it working.
Interactive Slide the KL weight from 0 upward and watch four code clouds pack toward the origin and widen; then click or tap anywhere on the plane to decode that code and see whether it lands in covered territory or in a hole.
Two terms and a trick
The training loss has two parts, both from one inequality in the paper. The first is reconstruction: how well decoded output matches input. The second pulls on the codes: every encoded bell curve is compared with a fixed reference, a standard normal — a bell curve centred at zero with width one. The comparison is the Kullback–Leibler divergence, zero when two distributions agree, growing as they differ. For a code of J numbers the paper's appendix gives a closed form: KL = ½ Σj ( μj² + σj² − 1 − log σj² ). It is zero when every mean is 0 and every standard deviation 1.
The two terms fight. Reconstruction wants each input its own sharp code, far from its neighbours, noise-free. The KL term wants every code a wide cloud on the origin. Training settles between: clusters packed near the origin, each blurred enough to overlap. Overlap is the point: trained on noisy draws, the decoder must produce something sensible for every code in each cloud, and because the clouds touch, it covers the whole neighbourhood of the origin. A fresh random code from the standard normal lands in covered territory and decodes plausibly. Switch the term off and you are back to a plain autoencoder, holes included.
One obstacle remains: gradient descent needs how the loss changes with the encoder's mean and width, but the code is a random draw — you cannot differentiate through a dice roll. The paper's reparameterisation trick moves the dice: instead of drawing a code directly, draw a standard normal number ε, then set z = μ + σ · ε. The randomness now lives in ε, untouched by any parameter; z is an ordinary differentiable function of μ and σ. Kingma and Welling report one draw per datapoint was enough when the minibatch held around a hundred examples.
In short
A variational autoencoder compresses data through a bottleneck that is a bell curve, not a point. The loss rewards faithful reconstruction and pulls every code cloud toward a standard normal at the origin. Reparameterising the draw as μ + σ·ε makes that loss differentiable. The result is a latent space with no holes, so a random code decodes to something that looks like data.
Where this comes from
- Auto-Encoding Variational Bayes linked only, not reproduced
arxiv.org/abs/1312.6114