← Notes and talks
Talk

Scaling laws for neural language models

Kaplan and colleagues turned model size, dataset size, and compute into a fitted prediction for held-out loss. This is what those fits establish, and the places where the conclusion people draw from them runs ahead of the evidence.

CS 886: Advanced Topics in Language Modeling, University of Waterloo · Fall 2026
On Jared Kaplan et al., Scaling Laws for Neural Language Models, arXiv:2001.08361 (2020)

Slides with speaker notes (PPTX)

The question

Suppose you have a fixed amount of compute to train a language model. You can make the model larger, train on more tokens, or run more optimization steps. The three are coupled, because every token costs more to process in a larger model. Picking among them had been a matter of intuition and precedent.

The paper replaces that with measurement. It asks two questions: how does held-out loss respond to each resource on its own, and which combination of the three reaches the lowest loss for the same budget.

What the loss actually is

The metric throughout is held-out cross-entropy, the average surprise of the model at the next token.

L=− 1T ∑t=1T ln⁡ qθ (xt∣ x<t)

Natural logarithms put this in nats per token. A model that gives the observed token probability 0.5 takes a loss of about 0.69 nats on it; probability 0.1 costs about 2.30. Lower is better, and the target token always comes from the held-out text, even when the model ranked something else higher.

This is worth stating plainly because the paper's improvements are often read as accuracy gains. They are not. A five percent reduction in loss is not five percentage points of anything.

Three symbols carry the rest of the argument: N is the number of non-embedding parameters, D is dataset size in tokens, and C is estimated training compute. Embeddings are excluded from N because doing so makes the trend across model shapes noticeably cleaner.

What was trained

Experimental range, from §2 and §3 of the paper.
QuantitySetting
Model size768 to 1.5B non-embedding parameters
Dataset size22M to roughly 23B tokens
ArchitecturePrimarily decoder-only Transformers
Corpus and contextWebText2, 1,024-token contexts, 50,257-token BPE vocabulary
Default schedule250,000 steps at 524,288 tokens per batch, 3,000 warmup steps then cosine decay

The method is to vary one resource while making sure the others are not the binding constraint, or else to fit a single function describing their joint effect. Which of those two applies matters for how far each result can be pushed.

The single-variable fits

Vary model size with enough data and training near convergence, and loss follows a power law in N. Vary dataset size with a model large enough to use it, stopping when held-out loss stops improving, and loss follows a power law in D with exponent 0.095.

L(N)= (NcN) αN , αN≃0.076
Test loss against non-embedding parameter count On log-log axes the fitted power law is a straight line: test loss falls from about 4.8 nats at 10^5 non-embedding parameters to about 2.4 nats at 10^9. 10⁵ 10⁶ 10⁷ 10⁸ 10⁹ 2 3 4 5 Non-embedding parameters N Test loss (nats) slope −0.076
Test loss against non-embedding parameter count, drawn from the fitted law with Nc ≈ 8.8 × 1013. It is straight on log-log axes because that is what a power law is.

On log-log axes a power law is a straight line, and the exponent is the slope. That is the whole content of the claim: multiplying a resource by a fixed factor produces the same proportional change in loss, wherever the fit holds.

Predicted relative loss from the model-size fit. Calculated from the exponent, not separately measured.
Model-size multiplierPredicted relative loss
1×1.000
2×0.949
10×0.840

Each further fixed improvement costs another multiplicative increase in resources. That is the part that makes the curves look benign and the budgets look anything but.

When both resources are finite

The two fits above are limiting cases, each assuming the other resource is abundant. The paper's joint equation covers the case where neither is.

L(N,D)= [ (NcN) αN/αD + DcD ] αD

One term stands for finite model capacity and the other for limited data. Raising the sum to αD makes the expression collapse back to each separate power law in the appropriate limit.

Test loss against model size at four fixed dataset sizes Curves from the joint fit. At small model sizes all four dataset sizes give nearly the same loss. As the model grows each curve flattens at a floor set by its dataset size, so extra parameters stop helping once data is the binding constraint. At 22 million tokens the loss moves only from 4.29 to 4.07 across four decades of model size. 10⁶ 10⁷ 10⁸ 10⁹ 10¹⁰ 2 3 4 5 Non-embedding parameters N Test loss (nats) unlimited data 22M tokens 220M tokens 2.2B tokens 22B tokens
The joint fit at four fixed dataset sizes, with αN = 0.076 and αD = 0.103. Each curve flattens at a floor set by its own dataset: at 22M tokens, four decades of extra parameters move the loss only from 4.29 to 4.07. The dashed line is the unlimited-data case.

The shape of that figure is the practical result. At small model sizes the four curves sit almost on top of each other, because capacity is what is binding and the dataset barely matters. As the model grows, each curve flattens at a floor set by its own dataset. Past that point more parameters buy almost nothing.

The batch-size correction

This is the step most summaries skip, and the compute result does not mean what it appears to mean without it.

Write total non-embedding compute as C ≈ 6NBS, where B is tokens per optimizer update and S is the number of updates. A larger batch does more parallel work per update and can cut the number of serial updates. Past a point the extra tokens per update stop reducing steps proportionally and simply cost more compute. The critical batch size is where that tradeoff turns, and it grows as loss falls, empirically as L−4.8.

So the authors convert each fixed-batch run into two idealized quantities: the compute it would have needed at small batch, and the serial steps it would have needed at large batch.

Cmin= C1+B/Bcrit , Smin= S1+Bcrit/B

These are two separate limits, not two targets a single run hits at once. At B = Bcrit you pay twice the minimum compute and twice the minimum steps. Applying the correction is what moves the compute exponent from 0.057 for the raw fixed-batch curve to 0.050 on the adjusted one, and every allocation number below is built on the adjusted version.

How the budget should grow

Because Cmin = 6N BcritSmin, the three growth exponents have to sum to one. They do not, however, come from the same place, and that is easy to miss when they are printed in a row.

Growth exponents against adjusted compute, from §6.1 and Table 6.
QuantityExponentWhere it comes from
Optimal model size0.73Fitted directly to the best model size at each budget
Critical batch size0.24Derived, 0.050 × 4.8
Minimum serial steps0.03Residual, 1 − 0.73 − 0.24

Made concrete, ten times the adjusted compute budget buys this:

Multipliers for a 10× increase in adjusted compute, within the fitted regime.
QuantityMultiplier
Model size5.37×
Critical batch size1.74×
Minimum serial steps1.07×
Tokens processed1.86×

Almost all of the growth goes into the model. That is the finding the paper became known for: spend new compute mostly on parameters, and stop training well short of convergence.

Where I would push back

None of what follows requires hindsight from later papers. It is all visible in this one.

The fits do not explain themselves

There is no derivation of 0.076 from anything about Transformers, language, or optimization. It is a measured slope. The paper is careful about this; readers of the paper often are not. A fitted exponent with no mechanism behind it gives you no principled way to know where it stops applying, which is exactly what you need when you are extrapolating several orders of magnitude past your largest run.

The same symbol takes different values in different fits

The dataset exponent αD is 0.095 in the one-variable sweep and 0.103 in the joint fit. Nothing is wrong with that, since the two fits use different configurations and different objectives. But it does indicate how much these constants depend on the experimental setup, and it sits awkwardly beside the idea that they are properties of language modelling as such.

Two normalizations disagree and the paper does not reconcile them

Figure 13 reports a compute normalization of about 2.3 × 108 PF-days. Equation (1.3) and Table 5 report 3.1 × 108. The paper gives no explanation for the gap. For the talk I used only the fitted slopes and declined to pick a normalization, because either choice would have implied a precision the source does not support.

The headline rests on the weakest number in the table

Train large models and stop early. That conclusion depends on serial steps growing very slowly with budget, which is the 0.03 exponent. That exponent is not measured directly. It is what is left after the other two are subtracted from one, and the authors themselves note it is small enough to be consistent with zero. If it really is zero, then serial steps do not grow with budget at all, which is a stronger and stranger claim than the paper makes anywhere. Either way, the most quotable conclusion is carried by the least robust quantity in the table.

Loss is not capability

Everything here is next-token prediction on held-out text. Smooth, predictable improvement in that number is not evidence of smooth improvement in any particular thing you would want a language model to do. The paper does not claim otherwise. The gap is simply left open.

Training compute is not the cost that binds

The optimum is defined over estimated training FLOPs. Hardware utilization, memory, and above all inference cost do not appear anywhere in it. A model that is optimal to train can be the wrong model to serve, and those pressures push toward smaller models trained on more data, which is the opposite of the recommendation.

You cannot rederive every number

Appendix D.7 excludes one-layer models from the N and C fits, and the largest unconverged models from the N fit. The paper reports its sweeps, its fitted forms, and its results, but not a complete fitting objective, weighting rule, or raw data table. The headline exponents are reproducible as claims. The decimals are not reproducible as calculations.

The question I ended on

Kaplan gives a genuinely quantitative answer to a question that was previously settled by intuition, and that is the achievement. But how much of that answer is a property of language models, and how much is a property of this particular experimental design?

What happened next

Hoffmann and colleagues revisited the allocation question in 2022 with more than 400 training runs and reached a different balance: model size and training tokens should grow at roughly equal rates with compute, near C0.5 each, rather than the model-first split here. A large part of the difference comes down to how the learning-rate schedule was handled in the original sweeps.

That is a revision rather than a refutation. The framework, that loss is a predictable function of scale which you can fit and then optimize against, is Kaplan's, and it survived. The specific numbers did not.

Slides

The deck carries full speaker notes: per-slide timings, the derivation behind each exponent, and the questions I expected to be asked.

Download slides (PPTX, 1.0 MB)

References

  1. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling Laws for Neural Language Models. arXiv:2001.08361, 2020.
  2. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. Training Compute-Optimal Large Language Models. arXiv:2203.15556, 2022.

Both figures on this page are drawn from the published equations and fitted constants rather than reproduced from the paper. Numerical examples are calculated from the reported exponents and are not additional experimental results.