It has been a week of high drama in AI, precipitated by the resignation of the AI safety researcher Jacob Coxon from Anthropic.
asserted
It → precipitate → Anthropic
This followed several weeks of increasingly lurid and disturbing revelations about the OpenAI/Hugging Face incident.
asserted
This → follow → incident
My inbox yesterday included a message from Business Insider with the subject line: “AI doomsday debate reaches boiling point.”
asserted
debate → include → point
Now, the Anthropic CEO, Dario Amodei, has written a 3,800-word, reassuringly phrased letter titled “We Must Pace the Frontier,” describing his proposals for avoiding (or at least postponing) doomsday.
asserted
We → write → doomsday
Sam Altman of OpenAI, Elon Musk of xAI, Demis Hassabis of Google Deepmind and Satya Nadella of Microsoft have all expressed support.
asserted
Altman → express → support
You may be forgiven for not immediately understanding what “pace the frontier” means.
uncertain
pace → forgive → frontier
(My first image was of Amodei walking deep in thought along the Finnish–Russian border.)
asserted
Amodei → walk → border
The phrase also appeared in July’s “Pacing the Frontier” open letter, signed by 1,386 employees of frontier AI labs, including Amodei himself.
asserted
phrase → appear → Amodei
While that letter may have upset the industry’s PR executives with its signatories noting “the complete absence of credible plans for controlling superintelligent AI systems” and asserting that “building things smarter than humans … is, objectively, an insane and suicidal thing to do”, Amodei’s monograph goes out of its way to mollify investors.
uncertain
monograph → upset → investors
The notion of pacing the frontier seems to come from Formula 1: when conditions become too dangerous for racing, a pace car comes onto the track and all the other cars have to follow it at a safe speed.
asserted
cars → pace → speed
Progress continues, without the danger.
asserted
Progress → continue → danger
Amodei writes: “To be clear, pacing does not mean halting model training or technical progress.”
asserted
pacing → write → training
Amodei’s letter is prompted by his concern that “AI has been advancing drastically faster, driven primarily by ... recursive self-improvement”.
asserted
AI → prompt → improvement
It’s as if he and Sam find themselves driving their F1 cars at 200mph neck-and-neck heading into the first corner, only to realize it’s covered in ice and they have no steering wheel.
asserted
they → ’ → wheel
No wonder they want to slow down.
asserted
they → want → ?
In brief, Amodei’s proposal has three parts.
asserted
proposal → have → parts
The first is to have third-party AI system evaluators working inside each company, with full access to the systems; he commits Anthropic to this plan now, without waiting for the government to require it.
asserted
government → have → it
The second part of the plan asks all the frontier AI companies in “democratic countries” to “establish common safety standards as well as limits on the rate of unchecked AI progress”, with government regulation where needed.
asserted
part → ask → regulation
The third part would include “authoritarian countries” in a broader compact.
asserted
part → include → compact
Here, Amodei goes out of his way to reassure those in Washington who see America’s lead in AI as its most important geopolitical asset.
asserted
who → go → asset
On a casual reading, there are many reasons to believe that Amodei is calling for a general slowdown in the rate of progress.
asserted
Amodei → be → progress
He talks about “limits on the rate of unchecked AI progress” and “some kind of ‘speed limit’ on the rate of recursive self- improvement (RSI)”.
asserted
He → talk → improvement
He says: “Progress will still seem fast, and we must make wise use of the time we gain.”
asserted
we → say → time
Slowing down would give companies a bit more time to work on safety; Amodei talks about one to two years of extra time for research on interpretability, alignment, and better testing methods.
asserted
Amodei → slow → interpretability
Let me pause here to respond to Amodei’s critics who say it’s just a bid to cement Anthropic’s lead with the help of government intervention.
asserted
it → let → intervention
In fact, the Wikipedia page on pacing in F1 races says it “eliminates any time and distance advantage that a leading driver may have had over the remaining field of competitors”.
uncertain
driver → say → competitors
Having said that, I think the pacing metaphor is completely misguided.
asserted
metaphor → say → that
We cannot set a slower rate of progress for capabilities and then hope that provides enough time to get the safety right.
asserted
safety → set → time
We must set the safety requirements first, and further progress occurs only when they are met.
asserted
they → set → requirements
Imagine if Boeing said: “We’re going to introduce a new plane every year, and we hope that provides enough time for some flight tests to be completed and for the results to be good.”
asserted
results → imagine → time
We would say: “No, you have that backwards; you can introduce a new plane only when it has passed all the tests and the government has issued an airworthiness certification.
asserted
government → say → certification
If that takes more than a year, so be it.”
asserted
it → take → year
A more careful reading of the document suggests that Amodei agrees with this objection.
uncertain
Amodei → suggest → objection
For example, he says that rules should be of the form: “If models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z.”
asserted
they → say → Y
In other words, we set safety requirements, and developers have to show that they meet those requirements.
asserted
they → set → requirements
This is in fact the “red lines” approach that AI safety researchers have been calling for.
asserted
researchers → call → that
And it means that if developers can’t figure out how to meet the safety requirements, then they will have to halt.
asserted
they → mean → requirements
It would be, in F1 terminology, a red flag and not a pacing car.
asserted
It → pace → terminology
Recursive self-improvement leading to superintelligent AI raises the risk of the irreversible loss of human control.
asserted
improvement → lead → control
The acceptable risk level for loss of control is perhaps one in 100m per year, not the one in 10 or one in five that the AI CEOs currently estimate.
asserted
CEOs → estimate → that
…and 6 more, not listed.