Story summary
Anthropic, an artificial intelligence firm, has faced public scrutiny over its partnership with the Pentagon and concerns about misuse of AI technology. Through a Freedom of Information Act lawsuit, The Intercept obtained documents revealing contracts between Anthropic and other tech giants like Google, OpenAI, and xAI, worth up to $200 million each, for developing AI tools aimed at enhancing military capabilities. Simultaneously, Anthropic reported that it had disrupted several attempts by individuals to misuse its models, including a case where someone tried to use the technology to engineer more harmful strains of viruses like chikungunya.
AI researchers and industry figures have issued stark warnings about the potential dangers posed by advanced AI systems, with some calling for urgent government regulation and international collaboration to manage AI development. The fear is that as AI rapidly improves, it could pose existential threats if not properly controlled, such as hacking early warning systems or synthesizing and spreading novel pathogens. This has led to debates over whether current safeguards are adequate and calls for voluntary slowdowns in the pace of AI advancement.
Jacob Coxon, a researcher at Anthropic who recently resigned, highlighted concerns about an impending race among companies like Anthropic and OpenAI toward superintelligent systems that could pose catastrophic risks. Despite these warnings, there is also skepticism about the immediacy and nature of such threats, with some arguing for more precise regulation focused on specific issues rather than general fears about AI's impact on society or employment.
Written for “AI Safety And Risks” on 2026-09-12,
grounded in this article and the 31 other(s) covering the same event.
My fiancé works at Anthropic.
asserted
fiancé → work → Anthropic
On Monday, after a weekend’s worth of concerned blog posts from OpenAI, I wrote about why we ought to take those warnings seriously.
asserted
we → write → warnings
For a host of reasons, I wrote, AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more.
asserted
they → write → capture
And so I can understand why, despite a rising drumbeat of ominous talk over the past several years, the mainstream has largely written off existential risk as a fantasy.
asserted
mainstream → understand → fantasy
Moments before I sent out that edition, though, a recently departed Anthropic researcher published an X post that brought the subject to the center of conversation.
asserted
that → send → conversation
“I resigned from Anthropic today,” wrote a young researcher named Jacob Coxon, in a tweet that generated 159 million views.
asserted
that → resign → views
“I spent the last three years doing pretraining research at both OpenAI and Anthropic.
asserted
I → spend → OpenAI
Neither company is acting responsibly.
asserted
company → act → ?
They are racing straight to self-improving superintelligence and gambling with our lives.”
asserted
They → race → lives
Coxon’s post picked up on the same themes as recent comments from OpenAI — including from chief scientist Jakub Pachocki, whose blog post “An Alien Mind” also warned (in less lacerating terms) about the risks posed by recursive self-improvement and the world’s disturbing lack of preparation for the consequences.
asserted
post → pick → consequences
It tells you something about the bubble I inhabit that, for those reasons, Coxon’s post initially did not make much of an impression on me.
asserted
post → tell → me
He is hardly the first Anthropic employee to resign for safety-related reasons; in February there was a minor stir after a fellow safety researcher quit to pursue a poetry degree after becoming distressed about the pace of progress and Anthropic’s role in it.
asserted
researcher → resign → it
Anthropic is a company founded by people who believed their former colleagues at OpenAI paid insufficient attention to safety and that they had at best an outside shot of steering the world to a better outcome.
asserted
they → found → outcome
More than three years ago, my Hard Fork co-host Kevin Roose profiled the company and called it “the white-hot center of AI doomerism.”
asserted
host → profile → doomerism
For Anthropic, that sense of doom attracts talent, gives them a sense of purpose, and (less often) ultimately drives them away.
asserted
sense → attract → them
Until recently, this was mostly a thing people made fun of the company for.
asserted
people → make → company
But shortly after Coxon’s original post, it got a signal boost from a striking source — Evan Hubinger, who continues to lead alignment science at the company.
asserted
who → get → company
“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote.
uncertain
Hubinger → believe → humans
“I personally think it is >10% within the next decade.
asserted
it → think → decade
I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
That was good enough for another 41 million views and another jolt to the discourse, despite the fact that (as Nitasha Tiku points out at the Washington Post) Hubinger had posted warnings like this for years.
asserted
Hubinger → believe → years
(“My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us,” he said in a 2022 talk.)
asserted
he → put → talk
Of course, back then, Anthropic wasn’t a $965 billion company heading toward what could be the biggest initial public offering of all time.
uncertain
what → head → time
Claude didn’t exist, and its models had not yet broken out of their sandboxes to compromise at least three separate organizations.
asserted
models → exist → organizations
(The company said Wednesday that it had hired research organization METR to conduct an investigation of those incidents.)
asserted
it → say → incidents
In short, AI felt less serious then.
asserted
AI → feel → ?
But in the aftermath of the OpenAI attack on Hugging Face and the general turn in public opinion against AI, the public seems increasingly ready to acknowledge the risks of building superhuman intelligence.
asserted
public → seem → intelligence
Earlier this month, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced the Ban Artificial Superintelligence Act.
asserted
Sanders → announce → Act
If signed into law, the act would ban companies from developing superintelligence and enforce a temporary pause in advanced AI development.
asserted
act → sign → development
It would also instruct the US government to seek international agreements that would prevent the development of superintelligence globally.
asserted
that → instruct → superintelligence
“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said, accurately.
asserted
Sanders → be → results
“The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control.
asserted
it → acknowledge → control
It is irresponsible for society to allow them to move forward and make these products even more advanced.
asserted
products → allow → ?
For the moment, it is unclear whether Sanders’ legislation has much chance of making it onto President Trump’s desk — or whether this oligarch-friendly administration would sign it.
uncertain
administration → have → it
But public sentiment around AI is changing rapidly, and it is only really moving in one direction.
asserted
it → change → direction
On Thursday, Axios reported that Sen. Josh Hawley — chair of the Senate Homeland Security & Governmental Affairs subcommittee on Disaster Management — had opened a probe into the Hugging Face attack.
asserted
Hawley → report → attack
Hawley reportedly called OpenAI’s handling of the situation “reckless.”
uncertain
Hawley → call → situation
“The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue," he wrote.
asserted
he → deserve → incident
(He also noted that "just this week, three Anthropic researchers expressed publicly that there is a greater than 10% chance that AI could kill all human beings within the next decade."
uncertain
AI → note → decade
Sen. Ted Cruz is also now reportedly working on legislation to address the catastrophic risks of AI.
uncertain
Cruz → work → AI
And there are also signs that the labs’ oft-stated preference for some sort of slowdown has teeth.
asserted
preference → be → teeth
…and 27 more, not listed.