Lightning Jar - Web Studio Lightning Jar Wordmark
Skip to main content
Reading List

Overtraining as the Path to Human-like AI

Summary:

Sean Goedecke argues that today's LLMs generalize worse than humans because they have never been pushed through grokking — the phenomenon where a model trained far past the point of memorizing a constrained dataset eventually snaps onto the deeper, simpler representation underneath it. Feeding a model all the data in the world lets it improve by memorizing more, which relieves the very pressure grokking requires, so the frontier labs' data-maximizing recipe may be steering around the interesting outcome. His proposed experiment inverts the recipe: a massively over-parameterized model overtrained on a deliberately small corpus, at a cost he guesses in the billions and with a long, funding-hostile plateau before any payoff. Speculative by its own admission, but a crisply argued case that one of the most interesting training regimes remains untried.

Excerpt:

"LLMs may fall short of human-like generalization not for lack of data but for lack of grokking — and the untried experiment is a huge model overtrained on a small corpus."
#LLMs#Machine Learning#Grokking#AI Research
Read Full Source