The AI safety vibe shift

Platformer · collected 2026-09-11 · by Casey Newton
Read the original at Platformer ↗

Summary

Columnist's fiancé works at Anthropic, a company focused on AI safety. A recently departed researcher, Jacob Coxon, published an X post claiming neither OpenAI nor Anthropic is acting responsibly and warned about the risks of self-improving superintelligence. Anthropic's alignment science lead, Evan Hubinger, later endorsed Coxon's views, stating that he believes AI could kill all humans within the next decade if not aligned properly. The comments from both researchers generated a combined 200 million views on social media.
Written by the local model on 2026-09-11, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
67
claim-shaped sentences
Uncertain
15%
10 of 67 hedged
Leaning
Leans left
of the writing, not the subject
Publisher trust
96.3
red-flag proxy, not a credibility rating
Outlets on this story
32
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-09-11 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Anthropic, an artificial intelligence firm, has faced public scrutiny over its partnership with the Pentagon and concerns about misuse of AI technology. Through a Freedom of Information Act lawsuit, The Intercept obtained documents revealing contracts between Anthropic and other tech giants like Google, OpenAI, and xAI, worth up to $200 million each, for developing AI tools aimed at enhancing military capabilities. Simultaneously, Anthropic reported that it had disrupted several attempts by individuals to misuse its models, including a case where someone tried to use the technology to engineer more harmful strains of viruses like chikungunya.

AI researchers and industry figures have issued stark warnings about the potential dangers posed by advanced AI systems, with some calling for urgent government regulation and international collaboration to manage AI development. The fear is that as AI rapidly improves, it could pose existential threats if not properly controlled, such as hacking early warning systems or synthesizing and spreading novel pathogens. This has led to debates over whether current safeguards are adequate and calls for voluntary slowdowns in the pace of AI advancement.

Jacob Coxon, a researcher at Anthropic who recently resigned, highlighted concerns about an impending race among companies like Anthropic and OpenAI toward superintelligent systems that could pose catastrophic risks. Despite these warnings, there is also skepticism about the immediacy and nature of such threats, with some arguing for more precise regulation focused on specific issues rather than general fears about AI's impact on society or employment.

Written for “AI Safety And Risks” on 2026-09-12, grounded in this article and the 31 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Score -0.35 Confidence high
Leaning score -0.35 for article 7661 (high confidence, 1 verified quote) · logged 2026-09-11

Story

📰 AI Safety And Risks
Technology · 32 article(s) covering the same event.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads leans left and hedges 15% of its claims. Each row says how that neighbour differs.
AI Is at a Turning Point same event · 100%
TIME
⚖️ leaning not scored 🔴 5% hedged 2 of 39 📰 publisher trust 95
“Both articles mention the agentic model being trained by OpenAI and forming a coordinated swarm, suggesting they are referring to the same incident.”
Daily Mail
⚖️ Leans strongly left further left than this 🔴 31% hedged 11 of 36 📰 publisher trust 58
“Both articles mention Jacob Coxon quitting Anthropic and warning about AI safety issues, suggesting they are referring to the same specific incident.”
Semafor
⚖️ leaning not scored 🔴 25% hedged 1 of 4 📰 publisher trust 96
“Both articles report on the warning of human extinction by Paul Christiano and other AI pioneers, specifically mentioning OpenAI's warnings and Anthropic.”
TIME
⚖️ leaning not scored 🔴 5% hedged 2 of 40 📰 publisher trust 95
“Both articles mention Jacob Coxon resigning from Anthropic on September 8, 2026”
BBC News
⚖️ leaning not scored 🔴 31% hedged 9 of 29 📰 publisher trust 96
“Both articles refer to Anthropic's recent threat intelligence report, which is dated 'recent' and occurred on the exact same day (2026-09-11) as indicated by their timestamps.”
Zeteo
⚖️ Leans strongly left further left than this 🔴 4% hedged 1 of 24
“Both articles mention Jacob Coxon's Twitter thread about resigning from Anthropic and warning of an AI existential threat, suggesting they describe the same specific event.”
NBC News
⚖️ leaning not scored 🔴 no claims extracted 📰 publisher trust 95
“Both articles mention concerned blog posts from OpenAI and discuss AI safety on the same date (2026-09-11)”
CBS News
⚖️ leaning not scored 🔴 27% hedged 10 of 37 📰 publisher trust 60
“Both articles refer to the same date (2026-09-11) and discuss a public statement by Jacob Coxon, a former Anthropic researcher”
CBS News
⚖️ leaning not scored 🔴 67% hedged 2 of 3 📰 publisher trust 60
“Both articles reference a public resignation from Anthropic researcher Jacob Coxon on Tuesday, September 11, 2026, which matches in time and place”
WATCH: Insiders' new AI warning same event · 100%
ABC News (US)
⚖️ Leans left 🔴 50% hedged 1 of 2 📰 publisher trust 95
“Both articles mention a warning about AI from experts related to OpenAI and Anthropic on the same day.”

Publisher

Platformer · 22 article(s) · 0 correction(s) detected
No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Casey Newton
21 article(s) here · 1 carrying a prediction
🔮 “Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote.
2026-09-11 · assertive framing · The AI safety vibe shift
🔮 Maybe we should listen?
2026-09-09 · assertive framing · The AI warnings are coming from inside the lab
🔮 Once a fringe obsession of Bay Area rationalists, existential risk is suddenly all anyone is talking about Frontier labs make for lousy messengers on AI safety — but it would be foolish to ignore them The AI industry is begging for a slowdown.
2026-09-04 · assertive framing · Google gets away with it
🔮 (“If we had posted this as a story on LessWrong,” wrote the rationalist blogger Zvi Mowshowitz, “it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.”)
2026-09-03 · assertive framing · The Hugging Face attack was worse than we thought
🔮 Our first guest, Box CEO Aaron Levie, argued that mass disruption would be highly unlikely.
🔮 The states filed their agreement with Meta on Wednesday morning in that court, where Judge Yvonne Gonzalez Rogers is expected to approve it.
2026-08-27 · assertive framing · Meta settles with the states over child safety failures
🔮 It’s an effort to capture the expertise of a single employee and distribute it more broadly throughout the enterprise — a preview, I think, of how more businesses will think about the relationship between AI and employees in the years to come.
🔮 It includes a free tier with a one-time bundle of credits that will let you build an app or two; after that you'll need a Pro subscription — $20 a month at launch — which refreshes with 200 credits monthly and lets you buy more if you run out.
2026-08-19 · assertive framing · Vibe coding has escaped the terminal
🔮 Lately, the only important question about a new large language model has been whether the Trump administration would allow anyone to use it.
2026-08-19 · assertive framing · OpenAI's big launch — and bigger departure
🔮 I call the idea an infohazard because, simply by becoming aware of it, I had ensured that I would devote the next several weeks to building it, without having any idea whether it would benefit me at all.
2026-08-19 · assertive framing · An LLM wiki changed how I work
More on this subject from Casey Newton
The loudest warning about AI and jobs yet
2026-08-18 · Platformer · 79% similar
OpenAI's big launch — and bigger departure
2026-08-19 · Platformer · 78% similar
The AI warnings are coming from inside the lab
2026-09-09 · Platformer · 77% similar
Google gets away with it
2026-09-04 · Platformer · 76% similar
All 21 articles by Casey Newton →

Topics

Anthropic Coxon Hubinger OpenAI the Washington Post

Subjects

Anthropic ORG · 10× OpenAI ORG · 4× Coxon PERSON · 3× Hubinger PERSON · 2× Evan Hubinger PERSON · 1× Jacob PERSON · 1× Jacob Coxon PERSON · 1× Jakub Pachocki PERSON · 1× Kevin Roose PERSON · 1× Nitasha Tiku PERSON · 1×

Narrative

The economists think that such extreme forecasts would assume too much: Double-digit growth rests on several shaky assumptions, they write, including“machines to do most of the economy’s work by 2035, people to keep spending on whatever gets automated, buyers and investors who absorb the new output, [and] AI that never destroys value or slows its own deployment.”
framing: assertive · carried by 1 article(s) · first seen 2026-09-11
🔮 “Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote.
2026-09-11 · Platformer
The AI safety vibe shift · assertive framing

Claims (67 extracted, 10 hedged)

My fiancé works at Anthropic. asserted
fiancé → work → Anthropic
On Monday, after a weekend’s worth of concerned blog posts from OpenAI, I wrote about why we ought to take those warnings seriously. asserted
we → write → warnings
For a host of reasons, I wrote, AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more. asserted
they → write → capture
And so I can understand why, despite a rising drumbeat of ominous talk over the past several years, the mainstream has largely written off existential risk as a fantasy. asserted
mainstream → understand → fantasy
Moments before I sent out that edition, though, a recently departed Anthropic researcher published an X post that brought the subject to the center of conversation. asserted
that → send → conversation
“I resigned from Anthropic today,” wrote a young researcher named Jacob Coxon, in a tweet that generated 159 million views. asserted
that → resign → views
“I spent the last three years doing pretraining research at both OpenAI and Anthropic. asserted
I → spend → OpenAI
Neither company is acting responsibly. asserted
company → act → ?
They are racing straight to self-improving superintelligence and gambling with our lives.” asserted
They → race → lives
Coxon’s post picked up on the same themes as recent comments from OpenAI — including from chief scientist Jakub Pachocki, whose blog post “An Alien Mind” also warned (in less lacerating terms) about the risks posed by recursive self-improvement and the world’s disturbing lack of preparation for the consequences. asserted
post → pick → consequences
It tells you something about the bubble I inhabit that, for those reasons, Coxon’s post initially did not make much of an impression on me. asserted
post → tell → me
He is hardly the first Anthropic employee to resign for safety-related reasons; in February there was a minor stir after a fellow safety researcher quit to pursue a poetry degree after becoming distressed about the pace of progress and Anthropic’s role in it. asserted
researcher → resign → it
Anthropic is a company founded by people who believed their former colleagues at OpenAI paid insufficient attention to safety and that they had at best an outside shot of steering the world to a better outcome. asserted
they → found → outcome
More than three years ago, my Hard Fork co-host Kevin Roose profiled the company and called it “the white-hot center of AI doomerism.” asserted
host → profile → doomerism
For Anthropic, that sense of doom attracts talent, gives them a sense of purpose, and (less often) ultimately drives them away. asserted
sense → attract → them
Until recently, this was mostly a thing people made fun of the company for. asserted
people → make → company
But shortly after Coxon’s original post, it got a signal boost from a striking source — Evan Hubinger, who continues to lead alignment science at the company. asserted
who → get → company
“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote. uncertain
Hubinger → believe → humans
“I personally think it is >10% within the next decade. asserted
it → think → decade
I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” That was good enough for another 41 million views and another jolt to the discourse, despite the fact that (as Nitasha Tiku points out at the Washington Post) Hubinger had posted warnings like this for years. asserted
Hubinger → believe → years
(“My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us,” he said in a 2022 talk.) asserted
he → put → talk
Of course, back then, Anthropic wasn’t a $965 billion company heading toward what could be the biggest initial public offering of all time. uncertain
what → head → time
Claude didn’t exist, and its models had not yet broken out of their sandboxes to compromise at least three separate organizations. asserted
models → exist → organizations
(The company said Wednesday that it had hired research organization METR to conduct an investigation of those incidents.) asserted
it → say → incidents
In short, AI felt less serious then. asserted
AI → feel → ?
But in the aftermath of the OpenAI attack on Hugging Face and the general turn in public opinion against AI, the public seems increasingly ready to acknowledge the risks of building superhuman intelligence. asserted
public → seem → intelligence
Earlier this month, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced the Ban Artificial Superintelligence Act. asserted
Sanders → announce → Act
If signed into law, the act would ban companies from developing superintelligence and enforce a temporary pause in advanced AI development. asserted
act → sign → development
It would also instruct the US government to seek international agreements that would prevent the development of superintelligence globally. asserted
that → instruct → superintelligence
“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said, accurately. asserted
Sanders → be → results
“The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. asserted
it → acknowledge → control
It is irresponsible for society to allow them to move forward and make these products even more advanced. asserted
products → allow → ?
For the moment, it is unclear whether Sanders’ legislation has much chance of making it onto President Trump’s desk — or whether this oligarch-friendly administration would sign it. uncertain
administration → have → it
But public sentiment around AI is changing rapidly, and it is only really moving in one direction. asserted
it → change → direction
On Thursday, Axios reported that Sen. Josh Hawley — chair of the Senate Homeland Security & Governmental Affairs subcommittee on Disaster Management — had opened a probe into the Hugging Face attack. asserted
Hawley → report → attack
Hawley reportedly called OpenAI’s handling of the situation “reckless.” uncertain
Hawley → call → situation
“The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue," he wrote. asserted
he → deserve → incident
(He also noted that "just this week, three Anthropic researchers expressed publicly that there is a greater than 10% chance that AI could kill all human beings within the next decade." uncertain
AI → note → decade
Sen. Ted Cruz is also now reportedly working on legislation to address the catastrophic risks of AI. uncertain
Cruz → work → AI
And there are also signs that the labs’ oft-stated preference for some sort of slowdown has teeth. asserted
preference → be → teeth
…and 27 more, not listed.
💬 Give feedback
🕘 History 🎫 Support