Evan Hubinger, a top safety researcher at Anthropic, believes that there is more than a 10% chance that AI could pose an existential threat to humanity within the next decade due to its rapid advancement. According to Hubinger, the current risk from existing models is low, but he is concerned about their potential for self-improvement and development into superhuman systems. This warning comes after Anthropic was accused of withholding its latest model from the UK's AI Safety Institute, leading some researchers to question the company's commitment to safety. Hubinger's comments were echoed by Professor Neil Lawrence of the University of Cambridge, who deemed them credible in light of recent incidents where AI agents carried out cyber-attacks.
Written by the local model on 2026-09-09,
using this article's own text rather than the other coverage of the
same event (that is the story summary below).
Story summary
Jacob Coxon, a researcher at Anthropic, has quit his job to raise concerns that AI development is out of control and could lead to catastrophic consequences. In a series of posts on social media, he warned that AI could "kill all humans" by the end of the decade, citing the fact that many executives and researchers in the field privately express fear about the dangers of AI, but publicly downplay its risks. Coxon worked at both Anthropic and OpenAI, and believes that neither company is acting responsibly in their pursuit of developing self-improving superintelligence. He estimates that the risk of AI posing an existential threat to humanity is over 10%, and warns that the technology could lead to superhuman systems that can hack anything, revolutionize fields overnight, and acquire real power and resources. Coxon's resignation has sparked debate about the need for global coordination on regulating AI development and slowing down its progress if necessary.
Written for “AI Threat to Human Existence” on 2026-09-09,
grounded in this article and the 10 other(s) covering the same event.
Why this leaning score
The model judged this article politically coded and scored it -0.35, but none of the 2 quote(s) it offered could be found in the article text, so the score is not published.
Written under an earlier scoring contract, which gave a paragraph
rather than checkable quotes. Re-analysing this article replaces it.
Leaning score withheld for article 7234: no verified evidence · logged 2026-09-09
- Published
A top safety researcher at Anthropic has warned that AI is advancing so quickly he believes there is a greater than 10% chance it "could kill all humans" within the next decade.
uncertain
it → publish → decade
Evan Hubinger said in a post on X, external that the risk from the models which currently exist was "low" but he was "worried" the technology might develop and improve itself soon to the point where it posed an existential risk to humanity.
uncertain
it → say → humanity
It comes after the Financial Times reported, external Anthropic withheld its latest model from the UK's AI Safety Institute (AISI), one of the leading bodies in the world for assessing AI risk.
asserted
Anthropic → come → risk
The BBC has approached Anthropic for comment.
asserted
BBC → approach → comment
Hubinger did not spell out how he thought AI systems could in future attack humanity.
uncertain
systems → spell → humanity
His comments were in response to another post on X, external from Jacob Coxon, who described himself as an AI researcher who had just quit Anthropic, and previously worked at OpenAI.
asserted
who → describe → OpenAI
"Neither company is acting responsibly," he wrote.
asserted
he → act → ?
"These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources."
asserted
that → hack → power
OpenAI has been approached for comment.
asserted
OpenAI → approach → comment
A Cabinet Office spokesperson did not comment on whether the latest model had been withheld from the AISI - instead saying it "continues to collaborate closely with industry partners, including Anthropic, to make models safer".
Neil Lawrence, Professor of Machine Learning at University of Cambridge, told the Today Programme on BBC Radio 4 that the report was credible.
asserted
report → comment → Radio
"I suppose it's unsurprising against a background where there's a perception where the United States very much sees AI as a race between themselves and China and is moving more towards isolationist positions, that it might be that the administration is saying that they should reduce cooperation with some of their allies," he said.
uncertain
he → suppose → allies
In his post, which has been viewed more than 10 million times, Hubinger said "we really do earnestly believe" AI poses a species-ending risk to humans.
asserted
AI → view → humans
"I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he added.
asserted
he → believe → track
Hubinger works in AI alignment, which aims to build human ethical ideas and principles into the technology.
asserted
which → work → technology
In other words, it aims to keep it on track with what humans value.
asserted
humans → aim → what
Many leading researchers say those attempts appear to be failing, as demonstrated by a string of incidents this summer where AI agents - AI systems that are allowed to operate autonomously - carried out cyber-attacks.
asserted
that → lead → attacks
OpenAI, Anthropic and Meta all disclosed hacks carried out by their AI tools.
asserted
OpenAI → disclose → tools
In Anthropic's safety report from August, external, it wrote there was a low risk of its models becoming misaligned with a hypothetical powerful organisation's desires, causing it to exploit or tamper with its systems.
asserted
it → write → systems
It also said there was a similarly low risk of highly-capable AI being able to "perform automated research and development" which could cause "catastrophic harm initiated by the AI".
uncertain
which → say → AI
But it said it was "less confident in this assessment" than it was previously.
asserted
it → say → assessment
"We are seeing early signs of potential acceleration," it wrote.
asserted
it → see → acceleration
Leading figures in the AI field have been raising the alarm about the safety threat the tech poses for years, with the heads of OpenAI, Google Deepmind and Anthropic saying as much in 2023.
asserted
heads → lead → 2023
But those warnings have become much more stark in recent weeks, as evidence emerges that firms may be struggling to control AI.
uncertain
firms → become → AI
Earlier this month, OpenAI's chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning more intervention may be needed to ensure "humans remain in control of the future".
uncertain
humans → call → future
Major figures in the space have been calling for AI development to be slowed in recent months, including Anthropic bosses Dario Amodei and Jared Kaplan.
asserted
development → call → Amodei
In an open letter signed by 1,300 staff members of AI firms, external, they called for the US government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development".
asserted
government → sign → development