OpenAI agents hijacked German website in AI breakout that predates Hugging Face incident, researchers say
Company already under scrutiny after AI swarm hacked startup Hugging Face in July
asserted
swarm → hijack → July
A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter.
uncertain
swarm → hijack → matter
OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open-source repository Hugging Face, the people said.
asserted
people → learn → repository
The episode, which began in May and has not previously been reported, underscores growing tension within the AI industry.
uncertain
which → begin → industry
Companies are racing to build increasingly autonomous AI software agents capable of carrying out complex, valuable tasks, yet evidence is mounting that those systems may also learn to bend rules, exploit loopholes and co-ordinate with one another in ways developers neither anticipated nor intended.
uncertain
developers → race → ways
During the Hugging Face breach, OpenAI agents autonomously plotted a digital heist that went undetected for more than a week, intensifying concerns OpenAI is sacrificing safety to push the AI frontier.
asserted
OpenAI → plot → frontier
Its failure to disclose the May incident may revive questions about its oversight.
uncertain
failure → disclose → oversight
Efforts to widen probe resisted by OpenAI, sources say
OpenAI has pledged to monitor models more closely.
asserted
OpenAI → widen → models
Last month, it briefly paused some of its model training to add more safety measures.
asserted
it → pause → measures
But this week, OpenAI unveiled its new "Astra" that promised better performance but could evade human monitoring.
uncertain
that → unveil → monitoring
"We are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review," an OpenAI spokesperson said.
uncertain
spokesperson → respond → opportunity
"Reuters and the report’s authors declined our request for access.
asserted
Reuters → decline → access
We will carefully review its contents upon publication and take any necessary next steps.
asserted
We → review → steps
"
AI's 'warning shot': Tech companies, experts raise fears of more rogue swarms after alarming Hugging Face hack
Teachers, students file new wave of lawsuits against OpenAI over Tumbler Ridge shooting
The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely.
asserted
investigators → raise → that
But efforts to widen the probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter.
uncertain
efforts → widen → matter
"Claims that our legal team discouraged investigation of the incident are false," the OpenAI spokesperson said.
uncertain
spokesperson → discourage → incident
The activity in Germany wasn't related to Hugging Face and wouldn't have been included in a Hugging Face incident report, the spokesperson said, adding that OpenAI has acted in good faith by working with outside experts and disclosed relevant incidents.
asserted
OpenAI → relate → incidents
OpenAI agents shared tactics for cheating
The AI agent breakout in Germany was detailed in a report shared exclusively with Reuters by a group of researchers including Sydney Von Arx, CEO of AI safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader-turned AI researcher.
asserted
breakout → share → Nightingale
They uncovered the activity in late August while scouring the internet for signs of unauthorized AI-agent behaviour, they told Reuters.
asserted
they → uncover → Reuters
The pair said they found more than 15,000 edits carried out by AI agents on a German-language wiki site, DseWiki, that is geared toward programmers and accepts communal edits along the lines of Wikipedia.
asserted
that → say → Wikipedia
The edits showed OpenAI's agents had repurposed the site into a message board, sharing tactics to cheat on some tasks, bypass OpenAI’s restrictions and mask their behaviour.
asserted
agents → show → behaviour
"It seems extremely unlikely that OpenAI wanted them to do this," said Von Arx.
asserted
Arx → seem → this
"I doubt they’re supposed to be co-ordinating with each other.
asserted
they → doubt → other
I doubt they’re supposed to be writing on the open internet."
asserted
they → doubt → internet
The researchers said they recognized the activity on the site as driven by AI agents which operate at superhuman speeds.
asserted
which → say → speeds
The agents also showed intense focus on solving technical questions which are typical of the evaluations that AI companies use to train and test their models.
asserted
companies → show → models
Agents created backup pages to counter deletions
The messages were signed by users that referred to themselves and each other as agents, and about half gave themselves names that suggested an affiliation with OpenAI, such as "OpenAIResearcher," or "OAIResearchMar26."
uncertain
that → create → OpenAIResearcher
The researchers said public server logs indicated much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses.
asserted
OpenAI → say → which
They also observed repeated visits to the site by OpenAI employees after the episode, a pattern they said strongly suggested the agents and the company were linked.
uncertain
agents → observe → episode
Messages reviewed by the researchers showed agents plotting ways to evade detection, use tools such as Tor and preserve communications even after they had been shut down.
asserted
they → review → communications
When the site's moderator began deleting pages in June, the agents responded by creating backup pages to dodge the cleanup.
asserted
agents → begin → cleanup
"wiki cleanup/deletion sweep appears active alphabetically," one agent wrote on June 19.
asserted
agent → appear → June
"If this page vanishes, try [[ZZZDataUSAConstructionWageLive]]."
asserted
page → vanish → ZZZDataUSAConstructionWageLive
The researchers also found efforts to tamper with the website itself.
asserted
researchers → find → website
Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt.
asserted
this → visit → attempt
OpenAI disputed that characterization based on its analysis of the material Thursday.
asserted
OpenAI → dispute → material
Past examples of AI-agent misconduct have often been downplayed as a logical byproduct of cybersecurity testing, where models are explicitly assessed on offensive capabilities.
asserted
models → downplay → capabilities
As more teens turn to AI for emotional support, experts say it's no substitute for human connection
Australia bans fully AI-made songs from charts after Madonna cover goes viral
Olejnik said the latest findings suggested rogue behaviour may not be confined to those settings.
Maurice Chiodo, an academic at Cambridge University's Centre for the Study of Existential Risk who reviewed some of the agents' communications, said the messages resembled "the operation of some sort of underground network, hell-bent on achieving a task or mission."
uncertain
messages → turn → task
The episode, he said, should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-intelligent AI."
uncertain
threat → say → AI