Gwern's 13,000-word essay 'Human-like Neural Nets by Catapulting' proposes that LLMs fail to achieve human-like generalization because they haven't 'grokked' their training data. Grokking — a phenomenon where continued overtraining past apparent convergence causes a sudden leap in generalization — requires training a very large, over-parameterized model on a relatively small dataset. This is the opposite of what frontier AI labs currently do (training smaller models on massive datasets). Gwern argues that training a ~100-trillion-parameter model on constrained data, at a cost of $3–10B, could force the model to discover deeper generalizations rather than relying on memorization. The author summarizes and evaluates this argument, noting the political and engineering obstacles, and expresses cautious optimism that a major lab should attempt it.
Table of contents
What is grokking?Are LLMs bad because they can’t grok?AI labs train small-ish models on oceans of dataGrokking requires training a huge model on a small datasetConclusion38.6K Impressions5 Comments