Seven Puzzles That Stump AI: A Human Challenge
Puzzles and games have long been benchmarks for AI progress, from checkers to modern language models. While AI has rapidly improved at some tasks, such as solving NYT Connections puzzles, it still struggles with subtle variations and spatial reasoning. This article presents seven tests that highlight where human cognition still outperforms machines, inviting readers to try them.
The article traces AI's puzzle-solving lineage back to Arthur Samuel's 1959 checkers program, which popularized the term "machine learning." Progress has been uneven: Columbia University researchers found that top models solved only 18% of NYT Connections puzzles in late 2024, yet by early 2025 some achieved near-perfect scores. However, this rapid improvement masks persistent blind spots.
Spatial reasoning remains a significant weakness, with language models failing at mental rotation tasks that humans handle intuitively. Similarly, models trained on familiar puzzle formats like Knights and Knaves often default to memorized responses when subtle variations appear, revealing a reliance on pattern matching rather than genuine logical adaptation.
These findings could influence how AI is deployed in fields requiring spatial judgment, such as architecture or engineering, where human oversight may remain essential. The gap between AI's rapid progress on some tasks and its struggles with others may also shape public expectations, tempering both enthusiasm and skepticism. As models improve, the boundary between memorization and genuine reasoning could become increasingly important for users who rely on AI for problem-solving.