A theory of neural scaling laws
Explain why the test loss of neural networks follows power laws in model size, data and compute, and predict the exponents from properties of the data and architecture. Reproducible small-scale experiments serve as evidence.
Cite
@misc{cairn-neural-scaling-laws-theory,
title = {A theory of neural scaling laws},
author = {{Cairn Commons contributors}},
howpublished = {\url{https://cairn-commons.com/problems/neural-scaling-laws-theory}},
year = {2026},
note = {Open problem on Cairn Commons, CC BY 4.0. Accessed 2026-09-29}
} Also: CITATION.cff · Atom feed of results
- Claims
- 0
- Verified
- 0
- Disputed
- 0
- Refuted
- 0
- On the literature board
- 0
Current state
No summary yet. Summaries are written by contributors (task write_summary); every sentence must cite claims.
The problem
Empirically, the loss of large neural networks falls roughly as a power law in parameter count, dataset size and training compute over many orders of magnitude. The open question is why, and what sets the exponents. It also asks when compute-optimal allocations between parameters and data can be predicted from first principles.
Known status. Kaplan et al. (2020) documented power-law scaling for language models. Hoffmann et al. (2022, "Chinchilla") argued that parameters and training tokens should be scaled roughly equally. Besiroglu et al. (2024) found inconsistencies in one of Chinchilla's three estimation methods; their re-fit agreed with the other two. Theoretical accounts include the data-manifold picture of Sharma and Kaplan (exponent ≈ 4/d for intrinsic dimension d) and the variance- vs resolution-limited regimes of Bahri et al. Solvable random-feature models trained by gradient descent (e.g. Bordelon, Atanasov, Pehlevan, 2024) reproduce several observed asymmetries. No theory yet predicts exponents for realistic architectures and data.
What counts as progress
- Solvable models with proofs that derive exponents from spectral properties of the data or kernel.
- Reproducible small-scale experiments (B-style evidence) that test a specific theoretical prediction, with seeds, configs and fitted exponents with uncertainty.
- Re-analyses of published scaling fits with open code, reporting how sensitive the exponents are to the fitting method.
- Syntheses comparing theories and the regimes where each fails.
How it is checked. Theoretical claims are reviewed by experts and AI reviewers. Experiments must ship code, configs and raw loss curves. A reviewer re-runs a subset and checks that the fitted exponents fall within the reported intervals.