Why bigger language models are predictably better — until they are not
Make a language model ten times bigger and its loss drops by a fixed fraction. Do it again: same fraction. On the right axes that is a straight line across six orders of magnitude — until something gives.
A rule that fits a page
A language model is trained to guess the next piece of text. Its loss scores how surprised it is by the real next piece, averaged over a test set: lower is better; the units are nats, a cousin of bits. In 2020 Kaplan and colleagues trained many Transformers on one web-text corpus and asked how well model size, training-set size or compute alone predicts the loss. The answer was a power law — one exponent and one constant each.
Why it matters
The paper reported smooth trends with no sign of bending at the top end, and drew a practical conclusion: for a fixed compute budget, loss is minimised by training a very large model and stopping well before it has finished learning, not by training a small one to completion. That finding is a large part of why models got big.
Interactive Drag parameters and tokens across six orders of magnitude and read the loss each fit predicts; flip axes between log-log and linear to see the same curve turn from a straight line into a hockey stick.
What a power law is
A power law says loss is proportional to N raised to a negative power. The paper writes the parameter law as L(N) = (Nc/N)αN, with fitted values αN ≈ 0.076 and Nc ≈ 8.8 × 1013, where N counts parameters excluding vocabulary and position embeddings. Because the loss depends on a power of N, multiplying N by a fixed factor always multiplies the loss by the same factor: every doubling cuts it to 0.95 of what it was; every tenfold step to about 0.84. Take logarithms and log L is a straight line in log N with slope −αN. That is why the plots use log-log axes: a straight line there means "each 10× buys the same fraction", whatever the starting size.
Data has its own law, L(D) = (Dc/D)αD with αD ≈ 0.095 and Dc ≈ 5.4 × 1013 tokens, measured with a large model and training stopped once test loss stopped falling; and compute has a third, with exponent about 0.050 when model and batch size are chosen well for the budget. Each law holds when that quantity alone is the bottleneck.
Two cautions from the authors. First, the constants are fits for their setup — WebText2, a 50,257-token vocabulary, a 1,024-token context. Change the tokenizer and the loss rescales, so Nc and Dc shift; the paper says they carry no fundamental meaning. Second, the lines cannot go on forever. Text is not fully predictable, so loss has a floor above zero: the entropy of language. The paper also found its own laws contradict each other far above the tested range — the compute law predicts a loss lower than the available data would allow — and read the crossing, near a loss of roughly 1.7 nats per token, as an estimate of the point at or before which the trend must break, uncertain by an order of magnitude either way. Two years later, Hoffmann and colleagues re-measured the split: where Kaplan had advised spending a bigger budget mostly on parameters, they found model size and training tokens should grow in equal proportion. The split moved; the power-law form stayed.
In short
Loss falls as a small negative power of model size, data and compute. A power is a straight line on log-log axes, so every tenfold step buys the same fraction. The constants belong to one corpus and tokenizer; the shape is the finding. And a line that falls forever would eventually predict a model less surprised by text than text allows — so it must bend.
Where this comes from
- Scaling Laws for Neural Language Models linked only, not reproduced
arxiv.org/abs/2001.08361 - Training Compute-Optimal Large Language Models linked only, not reproduced
arxiv.org/abs/2203.15556