Frontier language models went from solving 18% of New York Times Connections puzzles in late 2024 to near-perfect scores by early 2025, according to a Columbia University team cited in an MIT Technology Review analysis published August 26, 2026. That jump captures the pace of progress on one narrow puzzle benchmark. It also obscures a set of tests where the best models still lose to humans, sometimes badly.
The gaps show up in seven puzzle categories the review walks through: mental rotation, Knights and Knaves, SimpleBench, ARC-AGI, intuition traps, Tower of Hanoi, and river-crossing plus logic grids. Each isolates a different failure mode. Together they map what current LLMs are actually doing when they reason, and where memorization ends and generalization begins.
Spatial reasoning is the clearest human win. Mental rotation problems, standard IQ-test material asking whether two images show the same 3D object from different angles, remain a weak spot even for models with vision inputs. World-model claims from major labs have not translated into the kind of object manipulation architects and mechanical engineers do intuitively.
“Though today's language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles.”— Grace Huckins, AI reporter at MIT Technology Review
Key facts
- 01Best models solved only 18% of NYT Connections puzzles in late 2024, then reached near-perfect scores by early 2025.
- 02Apple researchers found LLMs handle Tower of Hanoi and river-crossing puzzles well until disks or people hit six, when performance collapses.
- 03A 2024 Google and University of Illinois Urbana-Champaign study showed models fail Knights and Knaves variants by pattern-matching training data.
- 04Researchers at the University of Washington, Stanford, and the Allen Institute for AI found LLMs struggle with logic grid puzzles as clue count grows.
- 05The term 'machine learning' was popularized in a 1959 IBM article by Arthur Samuel about a checkers-playing algorithm.
Memory can work against a model too. A 2024 study from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of Knights and Knaves puzzles, where knights always tell the truth and knaves always lie. Models frequently answered as if they were solving the canonical version they had seen in training, missing the tweaks. SimpleBench exploits the same pattern with math and physics questions that look like standard textbook problems but contain a twist. Humans catch the trick; top-tier models often don't.
ARC-AGI, the abstract visual reasoning benchmark, has become the most watched of these tests. Models improved sharply on ARC over the past year, but the review notes they perform better when grids are fed as strings of numbers encoding cell colors rather than as images. That is a tell.
The same pattern shows up in the puzzles Apple studied earlier this year. Apple researchers found LLMs can solve small Tower of Hanoi instances and river-crossing puzzles cleanly, but performance falls apart once the number of disks or people reaches six or higher. A separate study from the University of Washington, Stanford University, and the Allen Institute for AI observed similar breakdowns on logic grid puzzles as the count of houses, attributes, and clues grows.
The Apple paper went viral. Commentators split on what it actually proved: a specific limitation of transformer-based reasoning, or the unremarkable observation that error rates compound as any solver, human or machine, faces more state to track. That debate is unsettled and matters for how labs pitch frontier reasoning capabilities to enterprise buyers.
Humans are not exempt from puzzle traps either. The review flags a class of intuition problems where people give fast, wrong answers while models, given time to deliberate, get them right. One example: a bat colony that doubles daily fills a cave in 60 days, so the cave is half full on day 59, not day 30. Humans routinely miss it. Modern LLMs, running chain-of-thought, generally do not.
Puzzles have been a load-bearing part of AI development since Arthur Samuel's 1959 IBM article on a checkers-playing algorithm popularized the term machine learning. Chess and Go followed. The current puzzle set differs in one important way: the tests are chosen specifically because they resist the training-set memorization that lets models look smarter than they are on standardized benchmarks.
That is what makes the mixed scorecard useful. Connections went from 18% to near-solved in a matter of months, which suggests puzzle categories fall quickly once a lab targets them. Mental rotation and six-disk Tower of Hanoi have not fallen. The distance between those two facts is roughly the distance between pattern retrieval and general reasoning.
For enterprises buying agentic products, the practical read is narrower than the headline benchmarks suggest. Long-horizon planning, spatial manipulation, and problems whose complexity scales past the training distribution remain the failure modes worth stress-testing before deployment. A model that aces a coding benchmark can still lose to a four-house logic grid, and the ARC-AGI internal-reasoning research suggests that even correct answers may not come from the reasoning path a buyer assumes.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




