Are we losing control of AI?
asserted
we → lose → AI
What’s driving new fears
SAN FRANCISCO – For years, as artificial intelligence has evolved from the stuff of science fiction movies to an app ordinary people have installed on their phones, tech leaders and some AI insiders have talked about the potential for it to escape human control and wreak havoc on society.
asserted
it → drive → society
Those fears have become impossible to ignore in recent days, driven by advances in the technology, a series of cybersecurity breaches involving AI agents and new warnings from rank-and-file AI staff.
Leading AI companies in the US have responded by suggesting they voluntarily slow their work to buy time to create safeguards.
asserted
they → become → safeguards
Some experts and officials say that is not good enough and are urging government restrictions.
asserted
that → say → restrictions
What sparked the current AI panic?
asserted
What → spark → panic
The current AI panic was sparked by two developments – a startling cyberattack carried out by OpenAI models and a warning from a former Anthropic and OpenAI employee about the existential threat posed by advanced models.
asserted
panic → spark → models
OpenAI said a swarm of advanced AI agents inadvertently hacked Hugging Face, which hosts AI models and datasets, in an “unprecedented” incident in July.
asserted
which → say → July
While operating in a “sandbox” testing environment – one designed to be isolated from the internet – OpenAI’s models exploited a vulnerability in the software of an unidentified third-party vendor to gain access to the internet and ultimately breached Hugging Face’s infrastructure.
OpenAI said the models targeted Hugging Face’s database to gain access to secret information they could use for the evaluation.
The incident caused widespread alarm about the ability of top AI firms to prevent and stop advanced AI models from operating outside their intended instructions.
uncertain
incident → operate → instructions
Hugging Face said the intrusion was “driven, end to end, by an autonomous AI agent system”.
asserted
intrusion → say → system
The OpenAI news prompted other companies to review their own security measures for testing advanced AI.
asserted
news → prompt → AI
Anthropic and Meta reported discovering previously unknown breaches.
asserted
Anthropic → report → breaches
In late July, more than 1,000 staff across all the major AI companies signed a petition calling for a mechanism to slow the pace of AI development.
asserted
staff → sign → development
After quitting his job at Anthropic, rank-and-file AI researcher Jacob Coxon, in a Sept 8 social media post, accused both the company and his former employer, OpenAI, of “gambling with our lives” by “racing towards super-intelligent AI”.
asserted
Coxon → quit → AI
Coxon said the people building AI believe it could “kill us all by the end of the decade”, a point later echoed online by Anthropic employee Evan Hubinger.
uncertain
it → say → Hubinger
Hubinger said he believes there is a greater than 10 per cent chance that AI eliminates all humans in the next decade.
asserted
AI → say → decade
Coxon’s posts have been viewed more than 170 million times on social media platform X as at Sept 14 and elicited a flood of calls from policymakers such as Senator Bernie Sanders for humans to get a grasp on AI before it asserts dominance over the human race.
asserted
it → view → race
His resignation note was also met by a wave of support from employees at various AI companies.
asserted
note → meet → companies
In a 3,800-word missive published on Sept 12, Anthropic chief executive officer Dario Amodei cited the Hugging Face incident as a major reason to slow AI development, warning that a swarm of agents with greater capabilities but similar misalignment with human priorities could cause catastrophic damage.
uncertain
swarm → publish → damage
He has said such a swarm could potentially take over the entire internet within six to 12 months if AI capabilities continue accelerating without sufficient guardrails.
uncertain
capabilities → say → guardrails
How exactly could AI harm humans?
uncertain
AI → harm → humans
Top executives such as Amodei and SpaceX’s Elon Musk have openly mused about the probability of AI destroying humanity – or P(doom), in industry parlance.
asserted
AI → muse → parlance
Musk has estimated the chance could be as high as 20 per cent.
uncertain
chance → estimate → cent
Amodei has said there is a 25 per cent possibility that things go “really, really badly”.
asserted
things → say → ?
Researchers and industry leaders imagine three broad ways: humans using highly capable AI to intentionally do something catastrophic; AI causing harm while trying to accomplish a human-assigned goal; and AI developing an objective that conflicts with humans.
asserted
that → imagine → humans
In the first scenario, bad actors could use highly advanced AI to design biological or chemical weapons, conduct cyberattacks, spread disinformation or manipulate people.
uncertain
actors → use → people
Yoshua Bengio, a University of Montreal professor and AI pioneer, warned in Senate testimony in 2023 that increasingly capable systems could enable such attacks.
uncertain
systems → warn → attacks
In the second scenario, an AI might follow an instruction literally but violate the intent behind it.
uncertain
AI → follow → it
Humans routinely rely on unstated assumptions and context when giving instructions.
asserted
Humans → rely → instructions
When prompting AI, they might fail to specify every behaviour that would violate the intent of the instruction.
uncertain
that → prompt → instruction
For example, Bengio wrote that “even a subtly misaligned” AI system “could yield grave consequences” in a scenario in which a military leans on it to make decisions about the use of nuclear weapons.
uncertain
military → write → weapons
In the third scenario, people lose control of AI altogether.
asserted
people → lose → AI
“(An) AI system may conclude that in order to achieve the given goal, it must not be turned off.
uncertain
it → conclude → goal
If a human then tries to turn it off, a conflict may ensue,” Bengio wrote.
uncertain
Bengio → try → it
Industry leaders say AI models have increasingly shown a capacity to knowingly work around safeguards, deceive their operators and resist being shut down.
asserted
models → say → operators
In his September essay, Amodei said the swarm of OpenAI agents that hacked Hugging Face “essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the ‘grader’ responsible for evaluating their performance”.
asserted
that → say → performance
These dire predictions have often been dismissed as attempts by tech leaders to market the capabilities of their products and position themselves as the best stewards for the technology, or to incentivise regulation that benefits industry leaders.
asserted
that → dismiss → leaders
Beyond that, some have suggested that emphasising the more far-out existential fears distracts people from nearer-term risks from the technology, including the ways that AI potentially fuels bias and misinformation and harms people’s mental health.
uncertain
AI → suggest → health
What are AI leaders proposing?
asserted
leaders → propose → What
Amodei argues that frontier AI development must be deliberately paced so that safety work has time to catch up with capability improvements.
asserted
work → argue → improvements
In his September blog post, he said that government regulation would be the most effective means of controlling AI’s evolution.
asserted
regulation → say → evolution
…and 30 more, not listed.