Lil’Log
Read post

Scaling Laws, Carefully

A thorough technical deep-dive into neural scaling laws for large language models, tracing their history from early ML loss predictability work (Amari 1992, Hestness 2017) through the landmark Kaplan et al. and Chinchilla papers. The post explains the core power-law relationships between model size, dataset size, and compute, then reconciles the disagreements between Kaplan and Chinchilla (embedding parameter counting, small-model extrapolation errors). It extends into data-limited regimes, covering how data repetition affects training and how to model effective token counts with exponential decay penalties. A dedicated section on the trickiness of fitting scaling laws in practice highlights how rounding precision, loss averaging, and fit-region selection can dramatically shift predictions. Includes a toy simulation demonstrating these failure modes.

    #llm#deep-learning
Jun 25•15m read time•From lilianweng.github.io
Post cover image
Table of contents
Early days: ML loss predictability #Scaling Laws in Data-Infinite Region #Scaling Laws in Data-Limited Region #Trickiness of Fitting Scaling Laws in Reality #Citation #References #
274 Impressions
Lil’Log's image
Lil’Log

Lilian Weng is a machine learning researcher and writer who shares insights, research findings, and ...

13 Followers

•

45 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard