Four years after ChatGPT's release, the most linguistically capable machines ever built still need roughly 100,000 times more words than a child to learn a language. Meta's open-weight Llama 3.1, released in 2024, pretrained on 15 trillion tokens. A preteen raised in a linguistically rich home has heard something in the vicinity of 100 million words — and, unlike the model, actually understands them.
Cognitive scientists call this the data efficiency gap, and it is now the central puzzle sitting between AI research and developmental psychology. Frontier models from OpenAI, Anthropic, and DeepSeek are pretraining on an estimated 10 times more data than Llama 3.1 did two years ago. Add literacy to a human's diet and the lifetime count reaches maybe 300 million words by age 20 — still four orders of magnitude short of what a modern LLM eats before breakfast.
“Claude has seen the amount of language that an entire city will experience in one generation.”— Ethan Gotlieb Wilcox, cognitive scientist and linguist at Georgetown University
The scale disparity is difficult to hold in the mind. Printed out, the words used to train a modern LLM would stack past the International Space Station. A child's 100 million words would reach 20 meters. Toddlers, meanwhile, typically start producing grammatically correct sentences after hearing 10 million words, or 30 million on the high end.
Key facts
- 01Meta's Llama 3.1 pretrained on 15 trillion tokens; frontier models now use roughly 10x more.
- 02A preteen in a linguistically rich home has heard about 100 million words — LLMs consume 100,000x more.
- 03Toddlers produce grammatically correct sentences after hearing 10 to 30 million words.
- 04Available internet training data could run dry as early as the 2030s, forcing a rethink of scale-driven AI.
- 05Printed LLM training data would stack past the International Space Station; a child's 100M words fits in 20 meters.
Michael C. Frank, a cognitive scientist at Stanford University, frames the mismatch bluntly. The progress in LLMs has been genuine, he says, but the field is still burning down forests and scraping the entire sum of human knowledge to reproduce a milestone that happens in living rooms over the course of a year.
The stakes are practical as well as scientific. Available high-quality internet text could be exhausted as early as the 2030s, which means the scaling recipe that got the industry from GPT-2 in 2019 to ChatGPT in 2022 to today's frontier models has a visible end. Reverse-engineering how babies do it could produce more data-efficient models useful for training on video, or for building chatbots that serve minority language communities where trillions of tokens simply do not exist.
It could also settle a decades-old fight inside linguistics. In the 1950s, MIT's Noam Chomsky argued that language cannot be learned from statistics alone — that children must be born with hardwired knowledge of grammar because the linguistic input they receive is too impoverished to explain what they end up knowing. B.F. Skinner had argued the opposite: that language is learned through conditioning, like any other behavior. Chomsky won the argument for a generation.
That view, known as generative grammar, dominated US linguistics for decades and shaped the first wave of AI. Pentagon funding in the 1950s and 1960s poured into rule-based symbolic AI systems that tried to hand-code grammar into programs. The approach mostly failed. Natural-language processing froze through the AI winter that began in the 1970s. Neural networks returned in the 2010s, and by 2018 and 2019 the transformer-based models BERT and GPT-2 made clear that brute statistical learning from billions of tokens could, in fact, produce fluent language.
That result is a problem for the Chomskyan view — LLMs are naïve pattern-learners with none of the evolved biological machinery of the human cortex, and they nonetheless learn syntax. Alison Gopnik, a developmental psychologist at the University of California, Berkeley, says most researchers, including skeptics of AI, did not expect this to work. Richard Futrell, a linguist at the University of California, Irvine, notes that Chomsky's signature argument was precisely that language could not be learned from statistics. LLMs appear to falsify it — at scale.
The unresolved question is whether that falsification holds at human scale. Alex Warstadt, a linguist and data scientist at the University of California, San Diego, was a PhD student at New York University when BERT and GPT-2 shipped. The pushback he heard from linguistic colleagues was consistent: no one had ever been impressed by a language model trained on the amount of data a child actually encounters. Train GPT-2 on 30 million words, Frank notes, and you do not get a kid. You get a nonsense generator.
Ethan Gotlieb Wilcox, a cognitive scientist at Georgetown University, puts the mismatch in demographic terms — a modern LLM has seen the language an entire city will experience in one generation. Whatever children are doing, they are doing it with a rounding error's worth of input by comparison, and they are doing it in about a year.
The commercial implication is what makes this an AI story and not just a linguistics one. The scaling curve that has driven every major capability jump since 2018 assumes an effectively infinite supply of text. It does not have one. Labs are already turning to synthetic data, reinforcement learning from verifiable rewards, and longer test-time reasoning to squeeze more performance out of the same corpus — all of them workarounds for a data ceiling that human toddlers apparently do not have.
If a research group figures out what children are doing that transformers are not, the payoff is not a better chatbot. It is a training regime that gets frontier-level performance from a fraction of today's corpus, on a fraction of today's compute, in languages that currently have no path to a competitive model. That is the prize sitting behind the data efficiency gap, and it is why cognitive science has quietly become one of the most strategically interesting adjacent fields to AI.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




