MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-08-26 · via MIT Technology Review

Seven Puzzles That Stump AI: A Human Challenge

Puzzles and games have long been benchmarks for AI progress, from checkers to modern language models. While AI has rapidly improved at some tasks, such as solving NYT Connections puzzles, it still struggles with subtle variations and spatial reasoning. This article presents seven tests that highlight where human cognition still outperforms machines, inviting readers to try them.

Expanded Detail

The article traces AI's puzzle-solving lineage back to Arthur Samuel's 1959 checkers program, which popularized the term "machine learning." Progress has been uneven: Columbia University researchers found that top models solved only 18% of NYT Connections puzzles in late 2024, yet by early 2025 some achieved near-perfect scores. However, this rapid improvement masks persistent blind spots.

Spatial reasoning remains a significant weakness, with language models failing at mental rotation tasks that humans handle intuitively. Similarly, models trained on familiar puzzle formats like Knights and Knaves often default to memorized responses when subtle variations appear, revealing a reliance on pattern matching rather than genuine logical adaptation.

Context

These findings could influence how AI is deployed in fields requiring spatial judgment, such as architecture or engineering, where human oversight may remain essential. The gap between AI's rapid progress on some tasks and its struggles with others may also shape public expectations, tempering both enthusiasm and skepticism. As models improve, the boundary between memorization and genuine reasoning could become increasingly important for users who rely on AI for problem-solving.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at MIT Technology Review →
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “AI models flub these intelligence tests. Can you fare any better?.” Browse more stories.