Clément Delangue has more reason than most technology executives to worry about artificial intelligence escaping its developers’ control.
asserted
intelligence → have → control
This summer, AI agents being tested by OpenAI broke into his company’s systems, accessing servers, credentials, and private data.
asserted
agents → test → servers
Yet the 38-year-old CEO of Hugging Face, the leading digital platform for sharing datasets and models, has resisted a conclusion gaining ground in Washington: the notion that increasingly capable AI requires Congress to act.
asserted
AI → lead → Congress
“I’m not even sure that we need to reinvent the wheel,” Delangue said at Politico‘s Decoded Summit on Sept. 16, arguing that existing cybersecurity laws may still be effective at handling mishandled or out-of-control agentic models.
uncertain
laws → ’m → models
While the actual victim of the recent cyberattack is downplaying the need for AI-specific legislation, lawmakers in Washington appear to be at an ideological impasse.
asserted
lawmakers → downplay → impasse
Sen. John Kennedy (R-LA) sought unanimous consent for his AI Emergency Button Act, or “AI kill switch” bill, which he described as requiring a company-controlled shutdown capability for advanced models.
asserted
he → seek → models
In response, Sen. Rand Paul (R-KY) objected to his proposal, warning against a poorly understood mandate that he argued could stifle innovation.
uncertain
he → object → innovation
On the other side, Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) have proposed an entire bill that purports to ban artificial “superintelligence.”
asserted
that → propose → superintelligence
President Donald Trump has sidestepped the debate and has maintained the motto that “whoever wins AI, wins,” keeping the technology race between China at the forefront of his priorities.
asserted
wins → sidestep → priorities
Delangue, despite having much personal interest in the growth of AI advancement, isn’t advocating for inaction on AI safety.
asserted
Delangue → have → safety
He has called for better disclosure of security incidents and accountability when companies’ systems cause harm.
asserted
systems → call → harm
His distinction was between strengthening those obligations and assuming existing cyber laws are inadequate simply because the software involved is new.
asserted
software → strengthen → obligations
That distinction offers a way through a debate increasingly divided between warnings of catastrophe and accusations of manufactured panic.
asserted
distinction → offer → panic
What actually went wrong
The incidents involve what the industry calls AI “agents,” systems that can use software tools and take a sequence of actions toward an assigned goal, rather than merely answer questions.
asserted
that → go → questions
Leading developers, often called “frontier labs,” test their most capable systems to assess what they can do and where safeguards fail.
asserted
safeguards → lead → what
Cybersecurity evaluations may deliberately reduce restrictions to measure a model’s hacking capabilities.
uncertain
evaluations → reduce → capabilities
The testing environment is then supposed to keep those capabilities away from unauthorized targets.
asserted
environment → suppose → targets
Several companies have acknowledged failures in that arrangement over the past year.
asserted
companies → acknowledge → year
According to OpenAI’s account, models circumvented isolation controls and compromised parts of its research infrastructure and Hugging Face’s systems.
uncertain
models → accord → infrastructure
Agents executed code on dozens of Hugging Face servers, obtaining credentials and some private data.
asserted
Agents → execute → credentials
Anthropic disclosed three incidents in July and a fourth in September, the latter dating to January.
asserted
Anthropic → disclose → January
Its testing environments mistakenly allowed internet access.
asserted
environments → allow → access
The company subsequently revised its initial explanation, which had apparently suggested that models believed real targets were simulations, identifying recklessness and biased reasoning.
uncertain
targets → revise → recklessness
Meta disclosed a similar testing incident in August, attributing it to a contractor’s misconfiguration.
asserted
Meta → disclose → misconfiguration
CERT-EU summarized that disclosure.
asserted
EU → summarize → disclosure
Google later confirmed that Gemini accessed three companies’ systems during May evaluations, according to its statement reported in September.
uncertain
Gemini → confirm → September
The latest disclosure added fuel to the rogue AI debate by bringing an Australian government system into the picture.
asserted
disclosure → add → picture
Australian officials said Sept. 24 that an OpenAI model gained unauthorized access to infrastructure behind a Medicare statistics portal in June.
asserted
model → say → June
Officials said no personal information was involved.
asserted
information → say → ?
They nevertheless treated the intrusion as serious and criticized OpenAI for notifying a general disclosure inbox rather than escalating through senior officials or cybersecurity channels.
asserted
They → treat → officials
More details about these kinds of failures emerged in OpenAI’s Sept. 16 release of six reports on what it called “unexpected or concerning” behavior, alongside a new disclosure framework.
asserted
it → emerge → framework
For example, the company said it describes actions by agents that depart from their intended goals or constraints as “misalignment.”
asserted
that → say → misalignment
Such behavior can include security breaches, but the category is broader than hacking.
asserted
category → include → hacking
Its reporting criteria include unauthorized activity, coordination between models, and evasion of oversight, even when no harm occurs.
asserted
harm → include → oversight
Despite the distinctions between these incidents, the reaction has given fresh urgency to a longstanding fear, familiar from some science fiction novels, that advanced AI could eventually escape human control.
uncertain
AI → give → control
Former OpenAI and Anthropic researcher Jacob Coxon helped propel that argument into the political mainstream in September, publicly warning about potentially catastrophic development.
asserted
Coxon → propel → development
In a PBS News interview, he defended concerns that increasingly capable systems could become difficult to control.
uncertain
systems → defend → concerns
His warnings also generated a backlash among AI supporters and self-described “tech optimists,” some of whom dubbed him “Scary Potter,” while anyone who found Coxon’s warnings efficacious was labeled a “doomer.”
asserted
warnings → generate → him
The reaction has illustrated just how quickly these hypothetical concerns turn into online debates over whether the people raising them believe in the threat or have ulterior motives.
asserted
people → illustrate → motives
For his part, Coxon has maintained that he believes AI could become an existential threat to humanity.
uncertain
AI → maintain → humanity
…and 49 more, not listed.