
Thore Graepel exits DeepMind over LLM reasoning limits
Graepel's departure highlights a structural gap between statistical pattern matching in LLMs and the search-based reasoning of prior systems like AlphaGo. Evidence from ARC and math benchmarks shows persistent failure modes on out-of-distribution problems. Next steps require hybrid architectures integrating explicit planning modules rather than scale alone.
Graepel, a core AlphaGo team member, cited the absence of tree search and value estimation mechanisms in transformer-based systems as the core barrier to reliable reasoning. AlphaGo evaluated millions of positions per move using Monte Carlo tree search; today's LLMs generate tokens via next-token prediction without maintaining an explicit world model or backtracking capability. This distinction appears in benchmark gaps: o1-preview scores 83% on AIME math but drops below 30% on ARC-AGI tasks requiring novel abstraction.
Graepel: Hybrid search-LLM system reaches 60% on ARC-AGI by Q4 2027
Sources (3)
- [1]Opinion: Don’t be fooled—LLMs don’t reason(https://www.technologyreview.com/2026/10/02/1145666/the-download-biological-de-aging-ai-reasoning/)
- [2]Mastering the game of Go with deep neural networks and tree search(https://www.nature.com/articles/nature16961)
- [3]On the Measure of Intelligence(https://arxiv.org/abs/1911.01547)