This was the week that AI safety hit the big time.
asserted
safety → hit → time
A 27-year-old AI researcher named Jacob Coxon quit his job at Anthropic, declaring that OpenAI and Anthropic are racing to create technology that could destroy the human race:
uncertain
that → name → race
Other researchers echoed Coxon’s concern, stating their belief that AI has a reasonable chance of killing all of humanity within a very short space of time:
I’m not sure why this resignation and these statements went mega-viral.
asserted
resignation → echo → time
Plenty of researchers have made similar moves, and similar statements, over the past few years!
asserted
Plenty → make → years
Geoffrey Hinton, one of the pioneers of modern AI, quit Google back in 2023 over safety fears.
asserted
Hinton → quit → fears
Daniel Kokotajlo resigned from OpenAI in 2024, saying that the company wasn’t behaving responsibly in its drive toward superintelligence.
asserted
company → resign → superintelligence
William Saunders and Steve Adler did something similar.
asserted
Saunders → do → something
Mrinank Sharma left Anthropic earlier this year, and wrote a pretty well-read blog post about it.
asserted
Sharma → leave → it
What’s more, it’s been clear for years now that “AI could kill humanity” is a very common belief among AI researchers.
uncertain
kill → ’ → researchers
Grace et al. (2024) interviewed thousands of AI researchers in 2024, and found that more than half thought that artificial superintelligence has a significant chance of making the human race go extinct (or causing similarly bad consequences):
asserted
race → interview → consequences
The median AI researcher gave “doom” a 5-10% probability (depending on how the question was phrased), while their average probability was between 15% and 20%.
asserted
probability → give → probability
Later, smaller surveys found similar numbers.
asserted
surveys → find → numbers
The AI researchers may or may not be right, but the fact that lots of them think AI could kill the human race has never exactly been a secret.
uncertain
AI → think → race
It’s not clear why Coxon went so much more viral than his predecessors.
asserted
Coxon → ’ → predecessors
Maybe it was the fact that AI just solved one of the most important open problems in mathematics (which the best human mathematicians had been unable to solve for almost a century).
asserted
mathematicians → solve → century
Or maybe it was the Hugging Face attack, where a swarm of AI agents tried to cheat on a test by hacking various companies.
asserted
swarm → try → companies
Or maybe AI has just obviously gotten so much smarter that people throughout society were starting to get worried.
asserted
people → get → society
But whatever the reason, Coxon’s announcement was the one that really penetrated through to the public consciousness.
asserted
that → penetrate → consciousness
Suddenly, he was getting interviewed about AI doom on national news:
asserted
he → interview → news
Barack Obama is now urging Democrats to focus on AI risk.
asserted
Obama → urge → risk
Other politicians are calling for federal regulation.
asserted
politicians → call → regulation
Bernie Sanders is drafting a bill to ban AI “superintelligence”, including 20-year prison sentences for anyone working on the technology.
asserted
Sanders → draft → technology
Donald Trump is getting asked about an AI slowdown; so far he’s resisting the calls, but there are rumors that his advisors are calling on him to do something.
asserted
advisors → ask → something
Perhaps the most notable response came from the top figures in the AI field.
asserted
response → come → field
Dario Amodei, the head of Anthropic, wrote a blog post called “We Must Pace the Frontier”, calling for a coordinated slowdown in the rate of AI progress, and suggesting some ways to police AI companies to make sure they were all observing the slowdown.
asserted
they → write → slowdown
He wrote:
[O]ver the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up.
asserted
prevention → write → time
We must slow the pace at which we improve the capabilities of AI models…
asserted
we → slow → models
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas.
asserted
that → propose → dilemmas
As reasons for his increased worry, Dario cites A) the Hugging Face attack, and B) the possibility that AI will soon be able to improve itself without human help (a process called “recursive self-improvement”, or “RSI”).
asserted
AI → increase → process
Elon Musk (head of xAI), Sam Altman (head of OpenAI), and Demis Hassabis (former head of DeepMind) quickly agreed with Dario:
asserted
Musk → agree → Dario
At least some of the labs are reportedly holding secret talks on joint action to slow down AI.
uncertain
some → hold → AI
A coordinated slowdown in AI progress would be bad for these companies’ bottom line, because it would allow upstart competitors to catch up.
asserted
competitors → allow → line
So the fact that they’re still calling for a slowdown, in defiance of their own financial interests, is a clear sign that their worry about human extinction is sincere.
asserted
worry → call → extinction
In fact, anyone following these figures’ public statements over the past few years will have no doubt that they’re all deeply worried about catastrophic AI risks.
asserted
they → follow → risks
The leading AI figures — not just the founders and CEOs, but the researchers themselves — feel trapped in a “red queen’s race”.
asserted
figures → lead → race
They feel like if they stop working on AI, someone else will build it anyway, so they each feel like they have to beat everyone else in the AI race so they can make sure that the safest possible AI (i.e. their own AI) is the one that becomes the most powerful and dominant.
asserted
that → feel → race
Anyway, all of this was common knowledge in my social circle years ago, but now all of it has broken through to the mainstream.
asserted
all → break → mainstream
What do I have to add to this discussion?
asserted
I → have → discussion
I’m not an AI researcher or founder, nor do I think I have a superior grasp of the game theory of AI development.
asserted
I → ’m → development
But I do think I have two useful thoughts on how to persuade the general public to be more concerned about AI risk.
asserted
I → think → risk
…and 65 more, not listed.