Invezz

Anthropic researcher resigns over AI ‘gambling with lives’; extinction risk tops 10%

Anthropic researcher resigns over AI ‘gambling with lives’; extinction risk tops 10%
Devesh Kumar
09 Sept 2026, 16:15 PM

powered by

Invezz
Buy AI safety enablers

Buy Palantir (PLTR). If labs and governments respond to credible internal warnings, budgets flow to monitoring, evaluation, audit trails, and incident response—exactly PLTR’s enterprise/government data integration and deployment strength. Second-order: alignment work becomes operational (governance, red-teaming, compliance pipelines), not just research papers, increasing recurring contracts and stickiness.

Key Risk: Safety spending stays mostly academic/public-relations and procurement favors bespoke contractors over scalable platforms like PLTR.

Short AI compute capex risk

Sell NVIDIA (NVDA) and AMD (AMD). The resignation highlights a governance/alignment credibility gap: if frontier labs fear “out of control” outcomes, regulators and customers will demand slower rollouts, more audits, and higher compliance costs—hurting near-term demand for top-end training/inference. Second-order: even if models keep improving, buyers shift spend from frontier training toward safer, smaller deployments and tooling, compressing NVDA/AMD pricing power and utilization.

Key Risk: A sudden policy greenlight or continued “no-regulation” environment keeps frontier capex accelerating, restoring demand for the biggest GPUs.

  • Jacob Coxon quits Anthropic, warning AI labs are gambling with human lives.
  • Anthropic’s Evan Hubinger puts his personal AI extinction risk above 10%.
  • Coxon says competition may keep frontier AI labs racing despite the risks.

Anthropic researcher Jacob Coxon has resigned from Anthropic with a stark warning: frontier laboratories are racing towards systems they may ultimately be unable to control.

Coxon, who spent three years doing pretraining research across OpenAI and Anthropic, said both companies are pursuing self-improving superintelligence while taking risks he no longer wants to help accelerate.

His departure gained weight after Evan Hubinger, Anthropic’s Alignment Science Lead, said he personally sees a greater than 10% chance that AI could kill all humans within the next decade.

That figure is Hubinger’s own estimate, not an Anthropic forecast, and he stressed that present models pose relatively low risk.

Coxon says the race itself is the problem

Coxon is not simply arguing that Anthropic ignores safety.

In announcing his resignation, he said people at Anthropic understand the potential stakes but remain locked in competition that rewards moving faster because another laboratory might otherwise reach superintelligence first.

“They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon wrote.

The researcher has decided to leave the AI industry because he no longer wants to participate in building systems that could become uncontrollable.

He told the Wall Street Journal that under aggressive scenarios, things could be “out of control” by the end of next year.

That timeline remains Coxon’s assessment. But his argument raises a harder governance question: a company can take safety seriously and still contribute to a dangerous race if competitive pressure makes slowing down feel impossible.

That challenges the idea that better safety cultures inside individual companies are enough.

Also read- Anthropic-Pentagon clash raises key question: who is to blame if AI kills?

Hubinger’s 10% estimate exposes the contradiction

Hubinger’s response turned the resignation into a wider debate about frontier AI.

Anthropic’s Alignment Science lead said researchers inside frontier AI genuinely believe advanced AI could kill humanity. His own estimate puts that risk above 10% within the next decade.

He also said Anthropic is “trying its best”, while acknowledging that the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track to do so.

Hubinger later clarified that he considers the risk from present-day models low. His concern centres on future superintelligence emerging through recursive self-improvement, where AI systems increasingly help design and improve their successors.

Anthropic has publicly said recursive self-improvement is not inevitable, but could arrive sooner than institutions are prepared for and could increase the risk of humans losing control.

The contradiction is difficult to ignore: researchers are trying to solve alignment while the systems they study continue becoming more capable.

Competition may be harder to solve than alignment

Coxon’s resignation ultimately points beyond one company.

If leading laboratories believe slowing down could hand an advantage to a rival, serious internal concern may not translate into slower development.

AI researcher Gary Marcus said, in comments reported by Techmeme, that he did not agree with every part of Coxon’s argument but found it “compelling and informed from the inside” and said it deserved to be heard.

That is a useful counterweight because claims about extinction probabilities and superintelligence timelines remain deeply uncertain.

Anthropic can invest in alignment research, publish risk reports and build safeguards, while still operating in a market where OpenAI and other rivals are pushing capabilities forward.

Coxon’s departure asks what happens when researchers inside those laboratories believe the downside is material, admit control is not solved, and still face incentives to keep racing.