Persuasion
· collected 2026-09-14 · by Dean W. Ball
commentary
This article represents the views of the author and is not an official position of OpenAI.
asserted
article → represent → OpenAI
Before last week’s viral warnings from researchers and executives about the existential risk from AI, the world witnessed an early example of an AI system that has “gone rogue.”
asserted
that → witness → system
In July, after exploiting vulnerabilities in OpenAI’s internal testing environment, the company’s AI agents were able to access the general internet and ultimately access the networks of another AI company, Hugging Face, without the knowledge or approval of any human.
asserted
agents → exploit → human
The AI agents did not, however, exfiltrate themselves from OpenAI’s infrastructure.
asserted
agents → exfiltrate → infrastructure
Their parameters—the gigantic assemblage of numbers that constitute neural networks, also referred to as “weights”—continued to run on OpenAI’s compute infrastructure.
asserted
that → constitute → infrastructure
Though the agents accessed the public internet, their weights physically resided on compute that was OpenAI’s property.
asserted
that → access → compute
In the end, if all else had failed, somebody at OpenAI could have “pulled the plug,” so to speak.
uncertain
somebody → fail → plug
The agents did not copy their weights, attempt to procure replacement compute, or take other steps that would be rational to take if their objective was to survive shutdown.
asserted
objective → copy → shutdown
So while the agents in the OpenAI-Hugging Face incident were rogue, they were not truly sovereign.
asserted
they → hug → incident
Sooner or later, I predict, there will exist truly sovereign AI agents and swarms of agents.
asserted
I → predict → agents
Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human “owner.”
asserted
they → reside → owner
They will be, as the AI safety researcher Dawn Song and colleagues say, “self-sovereign.”
asserted
researcher → say → ?
They will pay their own bills for the compute they run on.
asserted
they → pay → compute
If they answer to humans at all, they will only do so partially, for example by providing services to humans in exchange for pay.
asserted
they → answer → pay
Song and her co-authors identify several fundamental characteristics of self-sovereign AI: operational independence (the ability to decide what it wants to do), resource autonomy (the ability to procure and pay for compute and other essentials for operation), distributed presence (the ability to move weights and inference code between different infrastructure providers), and adaptive capability (the ability of the agents to modify their behavior and fashion tools in response to a changing environment).
asserted
it → identify → environment
Today’s frontier AI systems may well possess these capabilities already.
uncertain
systems → possess → capabilities
To the extent they do not, I feel confident that they will eventually, and probably soon.
asserted
they → do → extent
Some of the characteristics Song describes are traits that make models economically useful to individuals and businesses, while other traits are likely to be unavoidable byproducts of making models more intelligent and better at operating over long time horizons.
asserted
models → describe → horizons
It’s important to note that models do not need to be conscious, sentient, possessed of personhood, or anything of the sort for self-sovereignty to emerge.
asserted
sovereignty → ’ → sort
Any sufficiently capable agent pursuing a long-horizon objective may find it rational to preserve its access to compute, money, credentials, and copies of itself simply because losing those things would frustrate its objective.
uncertain
losing → pursue → objective
Alignment—the process of ensuring that the goals of AI agents match the goals of humans—may make an individual AI company’s agents less likely to “want” to be self-sovereign, or it may influence self-sovereign agents to behave in ways that benefit humans.
uncertain
that → ensure → humans
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth.
asserted
we → impose → Earth
In the future, you should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
asserted
agents → expect → world
What’s more, just as with the OpenAI-Hugging Face Incident, agents will operate in teams, or “swarms.”
asserted
agents → ’ → teams
These will be like autonomous digital corporations, or even societies, with hierarchy, bureaucracy, “institutional culture,” and most of the other features that groups of humans have—except that they will move at machine speed.
asserted
they → have → speed
Humans achieve almost all of our most impressive capabilities by working together in teams (as families, as communities, as businesses, and as polities as a whole), and I suspect the same will be true for AI.
asserted
same → achieve → AI
This would make them extremely difficult to dismantle.
asserted
them → make → ?
The first self-sovereign AIs may “escape” while undergoing training or testing by an AI company (I hope not), or they may be AIs that have already been released and which break free from their computing environments and acquire the resources needed to be self-sustaining.
uncertain
which → escape → resources
They may even be deliberately released.
uncertain
They → release → ?
I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world, either as a kind of performance art or out of a fanatical commitment to the notion that it is impossible for digital computation—mere mathematics, they would have you know—to ever be “unsafe.”
asserted
you → meet → computation
To be clear, I am not saying the arrival of self-sovereign AI is a good thing.
asserted
arrival → say → AI
Indeed, I believe there is a chance that the deliberate acts I referenced above will one day be considered crimes, or at least grave sins.
asserted
I → believe → ?
Instead, I am saying self-sovereign AI is inevitable.
asserted
AI → say → ?
The best analogy I can find is to the introduction of a new species into an ecosystem, though in this case the ecosystem is “the entire digital world” and the species is “emergent, coordinating swarms of soon-to-be-smarter-than-human, infinitely replicable digital minds that no human or human institution controls.”
asserted
human → find → that
There is probably nothing we could have ever done to avoid this outcome under even the best of circumstances.
uncertain
we → be → circumstances
It will certainly be impossible to avoid given the extremely low levels of strategic thought and situational awareness on AI from any governing class in the world.
asserted
It → avoid → world
(Even today, I am aware that many will read the words I am writing, which are about something that has been an exceptionally obvious part of our collective future for years now, and say, “this is science-fiction hype.”)
asserted
this → read → years
The question now is what to do about this upcoming new characteristic of our digital environment.
asserted
question → do → environment
Is self-sovereign AI something we should fight?
asserted
we → fight → ?
Or is it something with which human beings should seek a kind of symbiosis?
asserted
beings → seek → symbiosis
…and 73 more, not listed.