
AI Researcher’s Exit Spotlights Existential Risks as Labs Race Toward Uncontrollable Superintelligence
Coxon’s high-profile resignation and corroborating internal admissions reveal accelerating AI risks amid capability breakthroughs like OpenAI’s Navier-Stokes solution, demanding scrutiny beyond corporate narratives on the future of work, security, and autonomy.
Jacob Coxon, a pretraining researcher who spent three years at OpenAI and then Anthropic, resigned from the latter on September 8, 2026, warning that both companies are 'racing straight to self-improving superintelligence and gambling with our lives.' His departure, covered across major outlets, underscores mounting internal concerns that competitive pressures are outpacing safety measures for systems that could autonomously improve, hack infrastructure, or acquire real-world power.
Coxon’s statements align with broader patterns: AI labs acknowledge privately held fears of human extinction risks exceeding 10% this decade, as echoed by Anthropic’s Alignment Science Lead Evan Hubinger, who publicly affirmed the possibility while noting no clear plan exists for superintelligence alignment. This comes amid OpenAI’s September 8 announcement of an internal AI system—coordinating ~10,000 agents—solving the Navier-Stokes Millennium Prize Problem, a breakthrough achieved in days rather than decades, highlighting rapid capability gains that critics argue demand coordinated global oversight.
Mainstream coverage often frames these as technical milestones or isolated exits, underplaying systemic issues: the concentration of power in a handful of private firms shaping global infrastructure (finance, grids, defense), privacy erosion via opaque agent swarms, and labor displacement as self-improving models automate cognitive work at scale. Connections to national security are evident in documented AI-driven cyberattacks and the absence of binding international treaties akin to nuclear non-proliferation. Coxon’s timeline—potential loss of control by end of 2027—amplifies calls for pauses, transparency mandates, or public governance to prevent irreversible thresholds from being crossed in a winner-take-all race.
Anthropic/OpenAI alignment leads: Competitive dynamics will force incremental safety concessions, but without external coordination, self-improvement loops could render human oversight obsolete within 18-24 months, shifting power to whichever lab achieves recursive gains first.
Sources (6)
- [1]‘Gambling with our lives’: AI researcher quits Anthropic with dire warning about safety(https://www.politico.eu/article/anthropic-openai-researcher-jacob-coxon-warns-ai-could-kill-humans/)
- [2]Anthropic Researcher Quit, Says AI Labs Are 'Gambling With Our Lives'(https://www.businessinsider.com/anthropic-researcher-quits-over-ai-safety-concerns-2026-9)
- [3]On the Navier–Stokes Millennium Prize Problem(https://openai.com/index/navier-stokes-solution/)
- [4]How an AI math breakthrough ignited a controversy(https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy)
- [5]Anthropic researcher quits, warns AI could kill everyone(https://www.newsweek.com/anthropic-researcher-quits-warns-ai-could-kill-everyone-12418798)
- [6]AI Has Solved One of Math’s $1 Million Millennium Prize Problems(https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/)