sean goedecke
Read post

Overtraining as the path to human-like AI

Gwern's 13,000-word essay 'Human-like Neural Nets by Catapulting' proposes that LLMs fail to achieve human-like generalization because they haven't 'grokked' their training data. Grokking — a phenomenon where continued overtraining past apparent convergence causes a sudden leap in generalization — requires training a very large, over-parameterized model on a relatively small dataset. This is the opposite of what frontier AI labs currently do (training smaller models on massive datasets). Gwern argues that training a ~100-trillion-parameter model on constrained data, at a cost of $3–10B, could force the model to discover deeper generalizations rather than relying on memorization. The author summarizes and evaluates this argument, noting the political and engineering obstacles, and expresses cautious optimism that a major lab should attempt it.

    #machine-learning#llm#neural-networks
Jul 18•11m read time•From seangoedecke.com
Post cover image
Table of contents
What is grokking?Are LLMs bad because they can’t grok?AI labs train small-ish models on oceans of dataGrokking requires training a huge model on a small datasetConclusion
38.6K Impressions5 Comments
sean goedecke's image
sean goedecke

104 Followers

•

1.3K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard