'The illusion of thinking': Apple research finds AI models collapse and give up with hard puzzles
reasoningllm-limitationsai-researchappletower-of-hanoi
Abstraction: Apple study showing large reasoning models collapse on hard logic puzzles
Key points:
- Apple researchers tested OpenAI o1/o3, DeepSeek R1, Claude 3.7 Sonnet Thinking, and Google Gemini Flash Thinking on Tower of Hanoi, checker jumping, river-crossing, and block stacking puzzles
- All models show progressive accuracy decline as complexity increases, reaching complete collapse (zero accuracy) beyond a model-specific threshold; Claude 3.7 Sonnet and DeepSeek R1 fail at 5-disc Tower of Hanoi
- Counterintuitive finding: models reduce reasoning token usage (think less) as difficulty approaches and exceeds their collapse threshold, even though problems become harder
- Providing the algorithm directly in the prompt does not fix the collapse; models still fail to follow it correctly at high complexity
- Authors interpret results as "The Illusion of Thinking" - models perform well on math/coding but lack genuine generalized reasoning for complex compositional problems
- AI expert Gary Marcus noted ordinary humans also fail Tower of Hanoi with 8 discs, and the study lacks human comparison baselines
Connections: Apple · Openai · Deepseek · Large Language Models · Reasoning
Source: https://mashable.com/article/apple-research-ai-reasoning-models-collapse-logic-puzzles