A researcher at a leading AI company has quit over concerns that big tech firms are in an 'out of control' race to build systems that could wipe out humanity.
uncertain
that → lead → humanity
Jacob Coxon has left Anthropic after spending the last three years training new AI models.
asserted
Coxon → leave → models
He claims that neither the firm behind the Claude system – nor ChatGPT creators OpenAI, where he previously worked – were 'acting responsibly' over the safety of AI.
uncertain
he → claim → AI
And he even warned that artificial intelligence could 'kill all humans' by the end of the decade.
uncertain
intelligence → warn → decade
Posting on X, he said: 'I resigned from Anthropic today.
asserted
I → post → Anthropic
I spent the last three years doing pretraining research at both OpenAI and Anthropic.
asserted
I → spend → OpenAI
Neither company is acting responsibly.
asserted
company → act → ?
They are racing straight to self-improving superintelligence and gambling with our lives.'
A researcher at Anthropic has quit over concerns that big tech firms are in an 'out of control' race to build systems that could wipe out humanity
In a series of tweets explaining his decision to leave Anthropic, Mr Coxon urged people to 'not underestimate the power' of AI – particularly if it becomes 'superintelligent'
uncertain
it → race → AI
In response to Mr Coxon, Evan Hubinger, Alignment Science lead at Anthropic, said that the firm believes AI has the potential to kill humans
asserted
AI → say → humans
In a series of tweets explaining his decision to leave Anthropic, Mr Coxon urged people to 'not underestimate the power' of AI – particularly if it becomes 'superintelligent'.
asserted
it → explain → AI
Superintelligence is the point at which an artificial system becomes more powerful than any individual, company or even nation.
asserted
system → become → individual
'These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.
asserted
that → hack → power
We have all witnessed the progress in each of these domains, and progress is not slowing,' Mr Coxon said.
asserted
Coxon → witness → domains
'The people building AI earnestly believe that it could kill us all by the end of the decade,' he said.
uncertain
he → build → decade
If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately.
asserted
people → couch → fear
No other human activity poses this level of danger.
asserted
activity → pose → danger
According to Mr Coxon, this danger is 'well-understood' at Anthropic. However, the company is 'locked in a race to get there first'.
uncertain
company → accord → race
He explained: 'Accepting this race and entering the "endgame" is a hubristic gamble that should not be launched from a private company's [messaging system] Slack.
asserted
that → explain → system
Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.'
asserted
Attempting → attempt → confidence
As an example, the researcher highlights the recent Hugging Face attack, which saw a firm hacked by OpenAI's rogue AI.
asserted
firm → highlight → AI
He said: 'Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable.
asserted
pacing → say → labs
I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
uncertain
which → feel → capabilities
To conclude, he urged fellow AI researchers to 'consider what the next few years will actually look like'.
asserted
years → conclude → what
Hollywood has long warned of the dangers of these technologies, in films such as The Terminator and its sequels.
asserted
Hollywood → warn → Terminator
Hollywood has long warned of technology's threat to humanity
Mr Coxon asked: 'Do you want to kick off a superintelligent RL [reinforcement learning] run without a rigorous understanding of its mind?
asserted
you → warn → mind
Should you put your head down because "it's happening anyway" – or take this moment to call for different conditions?'
asserted
it → put → conditions
In response to the post, Evan Hubinger, Anthropic's AI safety lead, confirmed that the firm believes AI has the potential to kill humans.
asserted
AI → confirm → humans
On X, Mr Hubinger said: 'Jacob is correct here – we really do earnestly believe AI could kill all humans!
uncertain
AI → say → humans
I personally think it is >10% within the next decade.
asserted
it → think → decade
I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.'
asserted
we → believe → track
The news comes as Ed Davey claimed that Anthropic did not submit its latest model to the AI Security Institute for testing due to 'pressure from the Trump administration.'
asserted
Anthropic → come → administration
Speaking at Prime Minister's Questions today, he cited a Financial Times report on the subject, and asked: 'Does the Prime Minister think it is acceptable for President Trump to undermine Britain's efforts to keep the world safe from dangerous AI?'
In response, Andy Burnham said that 'conversations with Anthropic will continue'.
asserted
conversations → speak → Anthropic
He added: 'AI poses risks to our national security, but it also could be the source of solutions.'
Mr Coxon isn't the first expert to raise concerns over superintelligent AI in recent days.
uncertain
Coxon → add → days
His comments come shortly after Geoffrey Hinton – a Canadian researcher often referred to as the 'Godfather of AI' – warned that superintelligent systems could 'lead to human extinction'.
uncertain
systems → come → extinction
'We would be very foolish to develop superintelligence now, when there is no scientific consensus it can be developed safely and controllably,' Dr Hinton said.
'Losing control over AI smarter than ourselves could be catastrophic and could even lead to human extinction.
uncertain
Losing → develop → extinction
Anthropic's Claude is one of the leading large language models (LLMs), which are trained by scraping vast amounts of text so they can understand and generate human-like language and responses to questions.
asserted
they → lead → questions