Today’s Focus

Evan Hubinger, a safety researcher at Anthropic, wrote on X that he believes there is a greater than 10% chance advanced artificial intelligence “could kill all humans” within the next decade, the BBC reported.

Hubinger said the risk from currently deployed models is “low” but that he is worried future systems could improve themselves to the point of posing an existential threat. He did not describe a specific mechanism by which AI might attack humanity.

His post responded to Jacob Coxon, who announced his resignation from Anthropic on Tuesday. Coxon, who says he spent three years on pre-training research at both OpenAI and Anthropic, wrote that “neither company is acting responsibly” and accused both of “racing straight to self-improving superintelligence,” according to the New York Post.

“These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources,” Coxon wrote, per the BBC.

The comments landed a day after the Financial Times reported that Anthropic withheld its latest model from the United Kingdom’s AI Safety Institute (AISI), one of the leading government bodies evaluating frontier AI risks. A UK Cabinet Office spokesperson declined to confirm the reported withholding but told the BBC the government “continues to collaborate closely with industry partners, including Anthropic, to make models safer.”

The BBC and New York Post said both Anthropic and OpenAI were contacted for comment.

Coxon told followers his warning was “not a marketing stunt,” adding that many executives “couch their phrasing in the press to sound sensible” while expressing fear privately, the New York Post reported. Anthropic Chief Executive Dario Amodei and OpenAI Chief Executive Sam Altman have previously said advanced AI carries catastrophic risks, while also arguing their companies are best positioned to manage them.

The Debate

Supporters argue

Those backing the researchers’ warnings say the public is finally hearing what insiders have long said privately. Coxon wrote on X that “the people building AI earnestly believe that it could kill us all by the end of the decade,” per the New York Post, and argued no other human activity poses comparable danger.

Neil Lawrence, Professor of Machine Learning at the University of Cambridge, told BBC Radio 4’s Today Programme that Hubinger’s warning was “credible” given the pace of capability gains at frontier labs.

Advocates of stricter oversight, including the Center for AI Safety, have argued for years that governments should treat extinction risk from AI as a policy priority on par with pandemics and nuclear war. They point to the AISI and its US counterpart as the minimum institutional response.

Supporters of the AISI process say the reported decision to withhold Anthropic’s latest model, if accurate, validates their concern that voluntary safety commitments from labs cannot substitute for binding rules. They argue that if even a safety-branded company like Anthropic sidesteps outside evaluation, mandatory pre-deployment testing is needed.

Critics argue

Skeptics contend that dramatic “kill all humans” framing distorts policymaking and serves the commercial interests of the very labs issuing the warnings. Meta Chief AI Scientist Yann LeCun has argued repeatedly on social media that current large language models are nowhere near the autonomy or planning ability needed to threaten humanity, and that existential rhetoric distracts from present harms.

Princeton computer scientists Arvind Narayanan and Sayash Kapoor, authors of the “AI Snake Oil” newsletter, have written that speculative extinction scenarios lack empirical grounding and risk crowding out enforceable rules on bias, fraud, and labor displacement.

Some industry figures also push back on Coxon’s characterization. They note that Anthropic publishes a Responsible Scaling Policy and that OpenAI has a Preparedness Framework, both of which commit the companies to halt deployment if capability thresholds are crossed.

Critics of expansive AI regulation, including venture capitalist Marc Andreessen, have argued that heavy-handed rules built around existential fears would entrench incumbents and slow beneficial applications in medicine, science, and defense.

What the experts say

Independent researchers who have tried to quantify AI risk generally find wide disagreement rather than consensus. A 2023 survey of 2,778 AI researchers by AI Impacts, a nonprofit research group, found the median respondent gave a 5% probability that advanced AI would cause human extinction or similarly severe disempowerment, with answers ranging from 0% to over 50%.

The RAND Corporation, in a 2024 report on catastrophic AI risks, concluded that current evidence does not permit precise probability estimates and urged governments to invest in evaluation capacity rather than rely on lab self-assessment.

The UK AISI and the US AI Safety Institute, both established in 2023 and 2024, have published technical evaluations showing frontier models can assist with cyberattacks and provide uplift on biological threats, though not yet at levels sufficient for autonomous mass harm.

Stanford’s 2025 AI Index Report documented that compute used to train frontier models has roughly doubled every six months since 2020, a pace the report’s authors, led by Nestor Maslej, said is outstripping the growth of independent evaluation infrastructure.

By the Numbers

10%: the probability Anthropic researcher Evan Hubinger gave on X that AI “could kill all humans” within a decade, per the BBC.

3: years Jacob Coxon says he spent doing pre-training research at OpenAI and Anthropic before his resignation, according to the New York Post.

5%: median probability of human extinction or severe disempowerment from advanced AI in a 2023 survey of 2,778 AI researchers by AI Impacts.

2023: year the UK established the AI Safety Institute, the body from which Anthropic reportedly withheld its latest model, per the Financial Times as cited by the BBC.

6 months: approximate doubling time for training compute used on frontier AI models since 2020, per Stanford’s 2025 AI Index Report.

2: frontier labs, Anthropic and OpenAI, that Coxon accused of “racing straight to self-improving superintelligence,” per the New York Post.

Sources

Get the briefing in your inbox every morning.

Subscribe