№ 43 · machine learning

How attention lets a model look back at everything

Attention lets a model look at every piece of its input at once and decide, for the job in front of it right now, how much each piece should count. Nothing is thrown away or memorised in bulk. Each step asks a question; the answer is a weighted mix of everything it can see.

The bottleneck it replaced

Early neural translators squeezed a whole sentence into one fixed-length vector, then generated the translation from that vector alone. Bahdanau, Cho and Bengio argued in 2014 that the single vector was a bottleneck: a long sentence is compressed just as hard as a short one, and earlier measurements showed the design losing quality quickly as sentences grew longer. Their fix: keep one vector per input word and let the decoder, at every output word, take its own weighted sum over all of them — a soft search for what matters now. Their model held up on 50-word sentences where the fixed-vector model fell apart, and its weights looked like an alignment: translating the man into l'homme, it read both English words to choose l'.

Queries, keys, values

Vaswani and colleagues in 2017 wrote the same idea in the form that has stuck. Every position offers three vectors: a key saying what it contains, a value it hands over if chosen, and a query — what it is looking for when it does the asking. The score between a query and a key is their dot product — multiply matching components, add them up — so a query pointing the same way as a key scores high; one at right angles scores zero.

The scores then go through a softmax: e to each score, divided by the total. This turns any list of numbers into positive weights summing to one, so the output is an honest average — every value times its weight, added up. Bahdanau's paper calls this an expected annotation: the mix you would get picking one source word at random with those odds. The softmax is also smooth — a slightly better score earns a slightly larger share — so it trains by gradient descent.

Interactive Click any word to make it the query; the bars are the softmax weights it puts on every word's key. Then drag the score multiplier to see the same scores sharpen into one winner or flatten into an even spread.

Eight words, each with a hand-made 4-number key (animate, thing, action, function-word) and a hand-made query — illustrative vectors chosen to make the point, not a trained model. Score = query · key, then softmax(multiplier × score); the bars and percentages are computed live. The multiplier stands in for how large raw dot products get: at 0 every word gets 12.5 %, and as it grows the biggest score — or the tied biggest scores, as when dog looks at saw and chased — takes nearly everything — the saturation that dividing by √dk is meant to prevent. The output row is the weighted sum of the keys used as values (a shortcut for the toy: a real head projects a separate value vector for each word).

Why divide by √d

The full formula is softmax(QKᵀ/√dk)V; the √dk does real work. A dot product is a sum of dk terms. If each query and key component is random with mean zero and variance one, the sum has variance dk, so typical scores grow like √dk. Push large scores through a softmax and one weight goes to nearly 1, the rest to nearly 0, and the slope goes nearly flat — the authors suspected this was why unscaled dot-product attention performed worse than the alternative at large dk. Dividing by √dk puts scores back on the scale the softmax handles well; their model used dk = 64 per head.

What the weights do not tell you

An attention weight says how much of a value went into a mix. It is tempting to read them as the model's reasoning, but the Transformer paper only says attention could make models more interpretable and that some heads appear to track sentence structure. Treat the weights as one mixing step in one head of one layer, not an account of why the model answered as it did.

In short

Attention keeps every input and lets each step ask its own question of them. The question is a query, each input answers with a key, and a dot product scores the match. A softmax turns the scores into shares summing to one, and the output is the values mixed in those shares. The √d keeps the scores small enough that the softmax stays soft.

Where this comes from

  1. Neural Machine Translation by Jointly Learning to Align and Translate linked only, not reproduced
    Dzmitry Bahdanau and Kyunghyun Cho and Yoshua Bengio · arXiv:1409.0473 · 2014
    arxiv.org/abs/1409.0473
  2. Attention Is All You Need linked only, not reproduced
    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin · arXiv:1706.03762 · 2017
    arxiv.org/abs/1706.03762