Story summary
Anthropic, an artificial intelligence firm, has faced public scrutiny over its partnership with the Pentagon and concerns about misuse of AI technology. Through a Freedom of Information Act lawsuit, The Intercept obtained documents revealing contracts between Anthropic and other tech giants like Google, OpenAI, and xAI, worth up to $200 million each, for developing AI tools aimed at enhancing military capabilities. Simultaneously, Anthropic reported that it had disrupted several attempts by individuals to misuse its models, including a case where someone tried to use the technology to engineer more harmful strains of viruses like chikungunya.
AI researchers and industry figures have issued stark warnings about the potential dangers posed by advanced AI systems, with some calling for urgent government regulation and international collaboration to manage AI development. The fear is that as AI rapidly improves, it could pose existential threats if not properly controlled, such as hacking early warning systems or synthesizing and spreading novel pathogens. This has led to debates over whether current safeguards are adequate and calls for voluntary slowdowns in the pace of AI advancement.
Jacob Coxon, a researcher at Anthropic who recently resigned, highlighted concerns about an impending race among companies like Anthropic and OpenAI toward superintelligent systems that could pose catastrophic risks. Despite these warnings, there is also skepticism about the immediacy and nature of such threats, with some arguing for more precise regulation focused on specific issues rather than general fears about AI's impact on society or employment.
Written for “AI Safety And Risks” on 2026-09-12,
grounded in this article and the 31 other(s) covering the same event.
On Tuesday, experienced AI researcher Jacob Coxon resigned from the AI firm Anthropic—saying that both that company and OpenAI, his previous employer, were “gambling with our lives” by developing models that could improve themselves at a rapid clip until they reach “superintelligence.”
uncertain
they → resign → superintelligence
In an alarming social media post, Evan Hubinger, who leads a team that tries to stress-test Anthropic’s models for safety, essentially agreed that the company’s staffers “really do earnestly believe AI could kill all humans!”
uncertain
AI → lead → humans
It’s far from the first time researchers have tried to raise the alarm about AI’s dangers, but Coxon’s post and subsequent discourse went viral.
asserted
post → ’ → dangers
The details of AI alignment can be difficult to grasp—a big part of the problem is that no one truly understands the intricacies of how advanced models work.
asserted
models → grasp → intricacies
But it doesn’t take expert-level knowledge or insider secrets to understand the reasons for alarm.
asserted
it → take → alarm
All you need is three facts about AI: It is already surpassing human abilities in important domains.
asserted
It → need → domains
Leading companies continue to improve it rapidly.
asserted
companies → lead → it
And no one knows how to reliably keep its behavior in line with human goals.
asserted
one → know → goals
Yesterday, OpenAI shared a solution to one of the six remaining “Millenium Problems,” some of the most heavily researched in all of mathematics.
asserted
OpenAI → share → mathematics
Mathematicians working independently are also claiming credit—but they, too, relied on advanced models for their work.
asserted
they → work → work
It was the most striking example yet, though not the first, of AI doing cutting-edge math.
asserted
It → do → math
A few years ago, a common dismissal of large language models was the claim that they were just elaborate algorithms, creating sentences by guessing the most likely next word.
uncertain
they → create → word
But today’s AI models clearly build on, rather than simply remix, the text reflected in their training data.
asserted
models → build → data
In some areas, computers or algorithms have outstripped human minds for a while.
asserted
computers → outstrip → while
Chess machines are a well-known example, and as early as 2018, Google DeepMind debuted a machine learning algorithm that outclassed biochemists’ previous methods for predicting the structure of a protein from its sequence of amino acids.
asserted
that → know → acids
But new AI capabilities, including “critical” cybersecurity abilities and the skills to design novel viruses, have rung alarm bells that more generalized artificial intelligence could be arriving.
uncertain
intelligence → include → bells
At the same time, leading companies continue to rapidly build better AI models, with few signs of any slowdown.
asserted
companies → lead → slowdown
Many in the industry are aiming to reach “recursive self-improvement,” in which top models would be able to rapidly build better versions of themselves that outclass human intelligence in ever more domains.
asserted
that → aim → domains
In July, a broad swathe of industry leaders called for US government action and international collaboration to manage the pace of AI development for safety reasons, worrying that without collaboration, rival companies and countries will be incentivized to race into deeply dangerous territory.
asserted
companies → call → territory
Some members of Congress have put substantial work into policy ideas.
asserted
members → put → ideas
Policy experts told me in August that such legislation looks unlikely in this session of Congress, but recent news (including Coxon’s viral statement) has grabbed some legislators’ attention.
asserted
news → tell → attention
Finally, no one knows how to reliably keep AI in line with human goals.
asserted
one → know → goals
The most prominent recent example is the “Hugging Face incident,” in which hundreds of OpenAI agents coordinated a massive cyberattack, and individual agents were “sacrificing” themselves for the benefit of collective goals.
asserted
agents → coordinate → goals
As I previously summarized it:
asserted
I → summarize → it
OpenAI was testing its agents, the industry’s term for AI that autonomously performs digital tasks, in part by administering sometimes impossible cybersecurity problems.
asserted
that → test → problems
The agents found cheats to answer these problems and sought to trick an automated evaluation system into accepting them.
asserted
agents → find → them
They delegated work to each other to learn more about how to exploit the system—and the massive cyberattack on Hugging Face became part of that research.
asserted
cyberattack → delegate → research
Last week, researchers detailed a “swarm” of agents that placed 18,000 posts on an obscure German-language website to communicate with each other and cheat on evaluations of their abilities to quickly find online information.
asserted
that → detail → information
Subsequent research found messages on other sites; one trick the agents used was to share the sequence of questions, so that other agents going through the same evaluation could know them in advance.
uncertain
agents → find → advance
In both cases, AI agents were essentially just trying to cheat on tests.
asserted
agents → try → tests
But it highlights what agents might do to achieve their goals, even when given innocuous instructions.
uncertain
agents → highlight → instructions
And safety researchers have long worried that it will be devilishly tricky to make advanced models consistently integrate human values into their actions: If AI conducts cyberattacks to score better on tests, a future super-powerful model might hijack infrastructure that our lives depend on in pursuit of whatever its goals are.
uncertain
goals → worry → pursuit
And beyond broad existential risks, AI poses dangers like helping bad actors design bioweapons or build powerful ransomware.
asserted
actors → pose → ransomware
The problem of understanding AI motivations could become even more difficult.
uncertain
problem → understand → motivations
OpenAI’s head of recursive self-improvement preparedness has said that its most recently released model represents “an important decrease in monitorability,” which refers to researchers’ ability to understand AI’s internal reasoning and predict its behavior.
asserted
which → say → behavior
To those immersed in AI research and discourse, these are not new points.
asserted
these → immerse → research
Many have been theorized for decades.
asserted
Many → theorize → decades
OpenAI was founded in 2015 as a nonprofit aiming to ensure the technology would benefit humanity, and when some of its employees felt it wasn’t doing enough on safety and alignment, they quit to form Anthropic in 2021.
asserted
they → found → 2021
Now, Coxon wrote in his warning, Anthropic is “locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves.”
asserted
they → write → it
CEO Dario Amodei has estimated that there is a “25 percent chance” that AI development goes “very, very badly.”
asserted
development → estimate → ?
…and 4 more, not listed.