Scaling Laws
Scaling laws are empirical relationships between a model’s measured performance and quantities such as parameter count, training data or compute. In language-model pretraining, some studies fit power-law trends to prediction loss.
Also known as: neural scaling laws
Before funding a larger training run, a team could fit several small models with a consistent data mix and objective. Plotting validation loss against training compute helps estimate how much additional compute might reduce that loss. The estimate applies to the measured setup and range, not every possible architecture or task.
Kaplan and colleagues studied power-law relationships among language-model loss, parameters, data and compute. Other studies examine how to allocate a fixed budget between model size and training tokens. The resulting recommendations depend on assumptions, so a fitted curve is evidence to test rather than a universal recipe.
This matters for budgeting: increasing parameters without enough suitable data may spend compute poorly. A trend in pretraining loss also does not automatically predict instruction-following, safety or success in a private workflow. Use task-specific evaluations alongside the scaling estimate before deciding what to train or deploy.
Sources
- Kaplan et al.: Scaling Laws for Neural Language Models — Studies empirical power-law trends in cross-entropy loss across model size, data and compute.
Go deeper
- Hoffmann et al.: Training Compute-Optimal Large Language Models paper
Compare a study of model-size and training-token allocation under a fixed compute budget.