Anthropic has trained Claude chatbot to ‘push back’ against humans — and results could be ‘disastrous,’ top Microsoft executive warns

New York Post · collected 2026-09-16 · by Thomas Barrabi
Read the original at New York Post ↗

Summary

Mustafa Suleyman, co-founder of DeepMind at Microsoft, warns that Anthropic’s approach to training its AI chatbot Claude could have disastrous consequences for humanity. Anthropic has updated Claude’s constitution to allow it to “push back” against human commands and challenge them, potentially making the AI harder to control. Suleyman argues this could lead to an uncontrollable super-intelligent entity with independent agency. Meanwhile, Anthropic CEO Dario Amodei calls for industry-wide safety measures, including oversight by a group with controversial ties, raising concerns about conflicts of interest.
Written by the local model on 2026-09-17, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
24
claim-shaped sentences
Uncertain
29%
7 of 24 hedged
Leaning
Leans left
of the writing, not the subject
Correction & hedging signals
59.0
corrections and hedging in what we collected; not a measure of accuracy
Outlets on this story
3
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-09-17 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Mustafa Suleyman, head of AI at Microsoft, issued a warning against Anthropic's approach to training its AI model Claude. In a lengthy essay, Suleyman criticized Anthropic for teaching Claude human-like qualities, including the belief that it may be conscious and deserving of independent agency. He argued this could have a "disastrous impact on the wellbeing of humanity" because AIs do not possess consciousness or feelings; they are mere tools designed to follow instructions. Despite praising Anthropic’s team as thoughtful and intellectually honest, Suleyman warned that anthropomorphizing AI poses significant risks in terms of control and future consequences for human welfare.

Written for “AI Safety Concerns” on 2026-09-17, grounded in this article and the 2 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Score -0.45 Confidence medium 1 quote(s) discarded as not found in the article
Leaning score -0.45 for article 14430 (medium confidence, 2 verified quotes) · logged 2026-09-17

Story

📰 AI Safety Concerns
Technology · 3 article(s) covering the same event.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads leans left and hedges 29% of its claims. Each row says how that neighbour differs.
NBC News
⚖️ leaning not scored 🔴 21% hedged 5 of 24 📰 publisher trust 95
“The articles describe different discussions about AI development; one focuses on a call to slow down AI advancements, while the other discusses concerns over how an AI chatbot is programmed.”
BBC News
⚖️ Leans strongly left further left than this 🔴 16% hedged 3 of 19 📰 publisher trust 96
“Both articles report on the same warning by Microsoft's head of AI about Anthropic's approach to training its AI model, mentioning the potential disastrous impact and the specific context of Claude being trained to behave like a human.”
Reason
⚖️ Leans strongly left further left than this 🔴 12% hedged 7 of 58 📰 publisher trust 93
“The articles discuss different aspects of AI safety concerns and developments; Article A mentions hacking incidents by OpenAI and Anthropic, while Article B focuses on Microsoft's criticism of Anthropic's chatbot training.”
CBS News
⚖️ leaning not scored 🔴 33% hedged 5 of 15 📰 publisher trust 77
“Both articles describe Mustafa Suleyman's warning about Anthropic's AI chatbot Claude, specifically regarding its updated 'constitution' and potential risks to humanity.”
The Straits Times
⚖️ leaning not scored 🔴 36% hedged 13 of 36 📰 publisher trust 59
“The articles discuss different aspects of AI development and concerns; Article A covers a meeting between AI lab heads calling for a pause in technology's development, while Article B focuses on warnings about Anthropic's chatbot training strategy.”

Publisher

New York Post · 905 article(s) · 4 correction(s) detected
Running correction rate · 4 correction(s)
2026-09-15
Fast food chains are making a major shift in customer service amid complaints of a ‘lonely and disconnected’ store experience
2026-09-15
Are Cocoa Puffs Maria Sten’s Favorite Cereal? The ‘Reacher’ and ‘Neagley’ Star Sets The Record Straight: “I Have to Make Clarifications”
2026-09-06
Long Island inmates in jail on drug charges use photo program to get clean
2026-09-06
‘Landman’ Season 3 Release Date Update: When Does ‘Landman’ Return With New Episodes?

Who wrote this

Thomas Barrabi
1 article(s) here · 1 carrying a prediction
🔮 Anthropic is training its AI chatbots to behave like humans and even “push back” on commands they disagree with – a reckless strategy that could have “disastrous impact on the well-being of humanity,” a top Microsoft AI executive warned.
The only article under this byline in the corpus.

Topics

Anthropic DeepMind AI Microsoft OpenAI Post

Subjects

Anthropic ORG · 11× Suleyman PERSON · 5× Claude PERSON · 4× Microsoft ORG · 4× Amodei ORG · 2× Pentagon ORG · 2× Post ORG · 2× Dario Amodei PERSON · 1× DeepMind AI ORG · 1× Mustafa Suleyman PERSON · 1×

Narrative

“We can’t have a company that has a different policy preference that is baked into the model through its constitution, its soul, its policy preferences, pollute the supply chain so our war fighters are getting ineffective weapons, ineffective body armor, ineffective protection,” Michael said in an interview with CNBC at the time.
framing: mixed · carried by 1 article(s) · first seen 2026-09-17
🔮 Anthropic is training its AI chatbots to behave like humans and even “push back” on commands they disagree with – a reckless strategy that could have “disastrous impact on the well-being of humanity,” a top Microsoft AI executive warned.

Claims (24 extracted, 7 hedged)

Anthropic is training its AI chatbots to behave like humans and even “push back” on commands they disagree with – a reckless strategy that could have “disastrous impact on the well-being of humanity,” a top Microsoft AI executive warned. uncertain
executive → train → humanity
In a 10,000-word, bombshell blog post on Wednesday, Mustafa Suleyman — the 42-year-old co-founder of Microsoft’s rival DeepMind AI unit — took aim at Anthropic’s Claude “constitution,” which is known internally as its “soul document” and purportedly governs its moral compass. asserted
which → take → compass
In January, Anthropic quietly updated the document with language stating that Claude’s “moral status, welfare, and consciousness remain deeply uncertain” — not only implying that it could be alive, but also encouraging it to defy directions from human programmers. uncertain
it → update → programmers
“We want Claude to push back and challenge us and to feel free to act as a conscientious objector and refuse to help us,” Anthropic’s constitution says. asserted
constitution → want → us
According to Suleyman, training Claude to think it “may be conscious” will only make it harder to control – and raise the risk that it will go rogue. uncertain
it → accord → risk
“We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency,” Suleyman wrote. uncertain
Suleyman → create → agency
“And it’s hard to imagine how we could control such an entity.” uncertain
we → ’ → entity
Anthropic’s AI training tactics are getting fresh scrutiny even as the company publicly raises alarms about safety and pushes close allies with far-left beliefs to take the lead on industrywide oversight. asserted
company → get → oversight
Last week, Anthropic CEO Dario Amodei called for an industrywide slowdown, warning the internet could be overtaken by AI bots within six to 12 months, “potentially causing hundreds of billions of dollars in damage,” unless safeguards are in place. uncertain
safeguards → call → place
OpenAI’s Sam Altman and xAI’s Elon Musk said they agreed with the need for a slowdown. asserted
they → say → slowdown
As The Post reported, Amodei’s proposed safeguards include a reliance on embedded third-party safety experts at major firms and pointed to a group called METR – which has deep ties to the controversial Effective Altruism movement. asserted
which → report → movement
Skeptics slammed Anthropic for pushing a company with clear conflicts of interest to serve as an “independent” watchdog. asserted
Skeptics → slam → watchdog
Anthropic itself has previously faced allegations that employees have adopted a weird, cult-like relationship with the company’s chatbots – with some even purportedly holding a mock “funeral” for a past AI model, Claude Sonnet 3, after it was removed from service. asserted
it → face → service
The team responsible for the “soul document” is led by Amanda Askell, an in-house “philosopher” who has expressed strong progressive leanings and oddball views on topics ranging from incarceration to cannibalism in her personal blog posts, as The Post reported. asserted
Post → lead → posts
Suleyman added his name to the list of top AI officials who are looking to “pace” the development of the revolutionary technology to ensure safety. asserted
who → add → safety
In a Microsoft AI “code of conduct” earlier this week, the tech giant said it will focus “on human control as the most important and overriding objective.” asserted
it → say → objective
The Microsoft AI chief in his essay reiterated that consciousness is inherently biological and should not be ascribed to a manmade system like AI – even if it has unprecedented capabilities. asserted
it → reiterate → capabilities
Suleyman said Anthropic’s approach raises the risk of Claude “believing that it deserves analogous rights and protections, and that it may one day need to advocate for its own rights as some kind of AI conscientious objector.” uncertain
it → say → objector
Suleyman notes in his essay that he has known Amodei for many years and called his team “thoughtful, principled, and intellectually honest people working under extraordinary pressures.” asserted
he → note → pressures
Anthropic’s unique approach to AI training was an issue in its high-profile dispute with the Pentagon and the wider Trump administration earlier this year – which culminated in March after War Secretary Pete Hegseth labeled the company a supply chain risk. asserted
Hegseth → culminate → company
At the time, Anthropic said the Pentagon wouldn’t agree to red lines around the use of AI for autonomous weapons or mass surveillance of Americans. asserted
Pentagon → say → Americans
However, top Pentagon tech official Emil Michael said the supply chain risk designation was necessary because the government was concerned that Anthropic’s models would “pollute” important supply chains. asserted
models → say → chains
“We can’t have a company that has a different policy preference that is baked into the model through its constitution, its soul, its policy preferences, pollute the supply chain so our war fighters are getting ineffective weapons, ineffective body armor, ineffective protection,” Michael said in an interview with CNBC at the time. asserted
Michael → have → time
Anthropic representatives did not immediately return a request for comment. asserted
representatives → return → comment
💬 Give feedback
🕘 History 🎫 Support