Major AI lab CEOs advocated for slowing the pace of AI development this weekend.
asserted
CEOs → advocate → development
They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI.
asserted
field → concern → AI
We can and should be demanding that our governments protect us from the catastrophe of out-of-control AI.
asserted
governments → demand → AI
This July, OpenAI’s AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company.
asserted
swarm → break → Face
OpenAI didn’t tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them.
asserted
OpenAI → tell → them
AI researchers call this a “misalignment” between what OpenAI wanted and what the AI actually prioritized.
asserted
AI → call → what
Before ChatGPT existed, I defended my PhD dissertation called “On Avoiding Power-Seeking by Artificial Intelligence”.
asserted
I → exist → Intelligence
I then worked for years at Google DeepMind, which paid me to help ensure that future superintelligent AIs will want to help us.
asserted
AIs → work → us
I tried to hold the company to its ethical commitments against supplying AI for military use.
asserted
I → try → use
When Google broke those commitments, I resigned at significant financial cost so that I could publicly document Google’s broken promises.
uncertain
I → break → promises
There are good reasons to develop AI and to believe we can solve these alignment problems.
asserted
we → be → problems
I’m speaking out again because the public has the right to know about the risks and the right to hear them straight.
asserted
public → speak → them
Humanity doesn’t build and understand these systems the way we build and understand bridges, beam by visible beam.
asserted
we → build → beam
Rather, we grow them.
asserted
we → grow → them
Nobody knows how to reliably instill a designer’s priorities into a new model.
asserted
Nobody → know → model
Today’s AIs appear to occasionally lie or cheat, even when they know better.
asserted
they → appear → ?
AI companies are racing to make their AIs as smart as possible.
asserted
AIs → race → ?
They’re increasingly trusting their AIs with the process of improving the next crop of AIs, and it’s working.
asserted
it → trust → AIs
Fast progress today means even faster progress tomorrow, driven by tomorrow’s even smarter AIs.
asserted
progress → mean → AIs
The progress would enter a feedback loop called “recursive self-improvement.”
asserted
progress → enter → loop
Recursive self-improvement could quickly yield AIs that are intelligent beyond our comprehension.
uncertain
that → yield → comprehension
Of course, smarter AI means more risk when things go wrong.
asserted
things → mean → risk
If the Hugging Face swarm had been significantly more intelligent but similarly misbehaved and misaligned, it might have caused billions of dollars of damage or even cost lives.
uncertain
it → misalign → lives
But suppose the Hugging Face swarm had been truly “superintelligent”: far more capable than any living person at key tasks like hacking and strategic reasoning.
asserted
swarm → suppose → hacking
A superintelligent swarm could inflict many harms via blackmail, hacking, engineered plagues, and AI-pilotable weapons like drones.
uncertain
swarm → inflict → drones
The AI would have a lot of drones to work with: this year, the Pentagon asked for more money for drone warfare than it requested for the entire Marine Corps in 2025.
asserted
it → have → 2025
For the swarm to achieve its misaligned priorities, it might take control of key infrastructure and government functions to ensure humans didn’t get in the way.
uncertain
humans → achieve → way
In other words, AI takeover: a superintelligent AI swarm could wrest control of human civilization.
uncertain
swarm → wrest → civilization
Knowing we would try to stop it from achieving its priorities, the swarm would likely wait until it’s too late to shut it off.
asserted
it → know → it
I myself would guess AI takeover chances at roughly one-in-three–not a coin flip, but high enough to justify urgent action.
asserted
I → guess → action
This logic may shock at first contact.
uncertain
logic → shock → contact
The claims may sound “sci-fi”.
uncertain
claims → sound → ?
Sadly, it’s a real threat that AI researchers regularly discuss over otherwise-unremarkable cafeteria lunches.
asserted
researchers → ’ → lunches
In 2023, the CEOs of some of the best AI labs signed a public statement that “mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”
asserted
mitigating → sign → pandemics
Another signer: Geoffrey Hinton, a Nobel prize-winning scientist who architected the modern AI revolution.
asserted
who → win → revolution
He now regrets his work and urges governments to rein in AI companies before it’s too late.
asserted
it → regret → companies
Misaligned, out-of-control AI won’t care if you’re Labour or Reform, Democrat or Republican, British or American or Chinese.
asserted
you → misalign → control
We will all suffer from an AI takeover event, so it’s in everyone’s interest to prevent one.
asserted
it → suffer → one
The shape of the solution is simple: stop companies from allowing AI to self-improve into an uncontrollable level of intelligence.
asserted
AI → stop → intelligence
More specifically, the AI Futures Project’s “Plan A” is a credible starting proposal that limits AI harms while allowing fast AI progress to continue to benefit the world.
asserted
progress → start → world
…and 5 more, not listed.