Mustafa Suleyman, co-founder of DeepMind at Microsoft, warns that Anthropic’s approach to training its AI chatbot Claude could have disastrous consequences for humanity. Anthropic has updated Claude’s constitution to allow it to “push back” against human commands and challenge them, potentially making the AI harder to control. Suleyman argues this could lead to an uncontrollable super-intelligent entity with independent agency. Meanwhile, Anthropic CEO Dario Amodei calls for industry-wide safety measures, including oversight by a group with controversial ties, raising concerns about conflicts of interest.
Written by the local model on 2026-09-17,
using this article's own text rather than the other coverage of the
same event (that is the story summary below).
Anthropic is training its AI chatbots to behave like humans and even “push back” on commands they disagree with – a reckless strategy that could have “disastrous impact on the well-being of humanity,” a top Microsoft AI executive warned.
uncertain
executive → train → humanity
In a 10,000-word, bombshell blog post on Wednesday, Mustafa Suleyman — the 42-year-old co-founder of Microsoft’s rival DeepMind AI unit — took aim at Anthropic’s Claude “constitution,” which is known internally as its “soul document” and purportedly governs its moral compass.
asserted
which → take → compass
In January, Anthropic quietly updated the document with language stating that Claude’s “moral status, welfare, and consciousness remain deeply uncertain” — not only implying that it could be alive, but also encouraging it to defy directions from human programmers.
uncertain
it → update → programmers
“We want Claude to push back and challenge us and to feel free to act as a conscientious objector and refuse to help us,” Anthropic’s constitution says.
asserted
constitution → want → us
According to Suleyman, training Claude to think it “may be conscious” will only make it harder to control – and raise the risk that it will go rogue.
uncertain
it → accord → risk
“We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency,” Suleyman wrote.
uncertain
Suleyman → create → agency
“And it’s hard to imagine how we could control such an entity.”
uncertain
we → ’ → entity
Anthropic’s AI training tactics are getting fresh scrutiny even as the company publicly raises alarms about safety and pushes close allies with far-left beliefs to take the lead on industrywide oversight.
asserted
company → get → oversight
Last week, Anthropic CEO Dario Amodei called for an industrywide slowdown, warning the internet could be overtaken by AI bots within six to 12 months, “potentially causing hundreds of billions of dollars in damage,” unless safeguards are in place.
uncertain
safeguards → call → place
OpenAI’s Sam Altman and xAI’s Elon Musk said they agreed with the need for a slowdown.
asserted
they → say → slowdown
As The Post reported, Amodei’s proposed safeguards include a reliance on embedded third-party safety experts at major firms and pointed to a group called METR – which has deep ties to the controversial Effective Altruism movement.
asserted
which → report → movement
Skeptics slammed Anthropic for pushing a company with clear conflicts of interest to serve as an “independent” watchdog.
asserted
Skeptics → slam → watchdog
Anthropic itself has previously faced allegations that employees have adopted a weird, cult-like relationship with the company’s chatbots – with some even purportedly holding a mock “funeral” for a past AI model, Claude Sonnet 3, after it was removed from service.
asserted
it → face → service
The team responsible for the “soul document” is led by Amanda Askell, an in-house “philosopher” who has expressed strong progressive leanings and oddball views on topics ranging from incarceration to cannibalism in her personal blog posts, as The Post reported.
asserted
Post → lead → posts
Suleyman added his name to the list of top AI officials who are looking to “pace” the development of the revolutionary technology to ensure safety.
asserted
who → add → safety
In a Microsoft AI “code of conduct” earlier this week, the tech giant said it will focus “on human control as the most important and overriding objective.”
asserted
it → say → objective
The Microsoft AI chief in his essay reiterated that consciousness is inherently biological and should not be ascribed to a manmade system like AI – even if it has unprecedented capabilities.
asserted
it → reiterate → capabilities
Suleyman said Anthropic’s approach raises the risk of Claude “believing that it deserves analogous rights and protections, and that it may one day need to advocate for its own rights as some kind of AI conscientious objector.”
uncertain
it → say → objector
Suleyman notes in his essay that he has known Amodei for many years and called his team “thoughtful, principled, and intellectually honest people working under extraordinary pressures.”
asserted
he → note → pressures
Anthropic’s unique approach to AI training was an issue in its high-profile dispute with the Pentagon and the wider Trump administration earlier this year – which culminated in March after War Secretary Pete Hegseth labeled the company a supply chain risk.
asserted
Hegseth → culminate → company
At the time, Anthropic said the Pentagon wouldn’t agree to red lines around the use of AI for autonomous weapons or mass surveillance of Americans.
asserted
Pentagon → say → Americans
However, top Pentagon tech official Emil Michael said the supply chain risk designation was necessary because the government was concerned that Anthropic’s models would “pollute” important supply chains.
asserted
models → say → chains
“We can’t have a company that has a different policy preference that is baked into the model through its constitution, its soul, its policy preferences, pollute the supply chain so our war fighters are getting ineffective weapons, ineffective body armor, ineffective protection,” Michael said in an interview with CNBC at the time.
asserted
Michael → have → time
Anthropic representatives did not immediately return a request for comment.
asserted
representatives → return → comment