On one hand, I love AI technology.
asserted
I → love → technology
On the other hand, I do think there’s a substantial chance that AI will kill most people on Earth within the next decade or two, by designing superviruses.
asserted
AI → think → superviruses
AI is already capable of designing viruses not found in nature, so this isn’t a sci-fi scenario.
asserted
this → design → nature
Whether these superviruses would be designed and unleashed by nihilistic human individuals, doomsday cults, or rogue AI agents themselves might end up being a secondary question.
uncertain
designed → design → individuals
We know we have nihilistic human individuals who might decide to destroy civilization in a fit of depression or pique.
uncertain
who → know → depression
We know we have doomsday cults.
asserted
we → know → cults
The will to destroy humanity exists, and sufficiently capable AI will probably provide a way, if sufficient precautions are not taken.
asserted
precautions → destroy → way
But right now, nobody really knows what precautions will be sufficient.
asserted
precautions → know → ?
One idea — promoted by the big AI labs themselves! — is to intentionally slow down the development of AI capabilities.
asserted
idea → promote → capabilities
This could conceivably buy us time to take other precautions, such as improved security around bio-labs, better AI alignment, and so on.
uncertain
This → buy → labs
Intentionally slowing AI development is called “pacing”.
asserted
slowing → slow → development
The biggest question facing the “pacing” debate right now is whether to curb the use of AI to design better AI — often called “recusive self-improvement”, or “RSI” for short.
asserted
question → face → AI
I haven’t waded into the pacing debate myself, but as a start, I thought it would be interesting to publish the thoughts of the good folks at the Institute for Progress, whose judgement I generally trust.
asserted
I → wad → judgement
Part 1 (today’s post) covers how seriously we should take this possibility of RSI, and whether it justifies slowing down frontier AI development.
asserted
it → cover → development
Part 2 will cover policy recommendations.
asserted
Part → cover → recommendations
If you work in US policy and would like to connect with the authors, you can reach Tim Fist at tim.fist@ifp.org and Saif Khan at saif@ifp.org.
asserted
you → work → saif@ifp.org
Frontier AI companies are racing to automate the development of AI, but they seem increasingly worried about what will happen if they succeed.
asserted
they → race → AI
More than 1,300 employees across every US frontier AI company recently signed an open letter calling for the government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
asserted
government → sign → development
The official OpenAI and Anthropic accounts tweeted messages in support of the letter, and the same day Sam Altman told an interviewer “we may have to pace the rate of AI development.”
uncertain
we → tweet → development
But before considering whether the letter is relevant for government policy, we have to answer two questions: What does “pacing” actually mean?
asserted
pacing → consider → What
And does the argument for it stand up to scrutiny?
asserted
argument → stand → scrutiny
In our view, the letter is implicitly arguing three things:
Frontier AI companies are close to fully automating AI R&D.
Automating AI R&D would pose serious risks.
asserted
Automating → argue → risks
Building the option to “pace” — i.e., somehow slow down progress toward fully automated AI R&D — is a good way to address those risks.
asserted
Building → build → risks
Slowing down AI progress should not be taken lightly: Advances in AI could unlock massive societal benefits, from new cures for diseases to abundant robotic labor.
uncertain
Advances → slow → labor
Yet if the AI researchers and CEOs are right about automating AI R&D — both that they could do it and that it would be extremely risky — the right tradeoffs for policymakers might look very different when we get there.
uncertain
we → automate → policymakers
With the right preparation, we might be able to manage the risks of automated AI R&D while having AI’s capabilities progress faster and diffuse more broadly than they do today.
uncertain
they → manage → R&D
So, despite substantial uncertainty, we believe the US should take low-regret policy actions now to prepare for a possible future in which serious risks from automated AI R&D require some form of “pacing.”
asserted
risks → believe → pacing
Here we’ll explain why, including what makes us take the open letter’s claims seriously and our principles for choosing policies with minimal downside if the risks prove overblown.
uncertain
risks → explain → downside
In the next post, we’ll provide a detailed list of specific policy recommendations.
asserted
we → provide → recommendations
Are frontier AI companies close to fully automating AI R&D?
asserted
companies → automate → R&D
Frontier AI companies are racing to automate AI R&D. Sam Altman, for example, stated last year that OpenAI aimed to have a “true automated researcher” by March 2028, and Anthropic’s leaders have made similar predictions.1
asserted
leaders → race → March
But how do those goals stack up to reality?
asserted
goals → stack → reality
One way of answering is to look at how models are getting better over time at AI R&D.
asserted
models → answer → R&D.
You can break down the skills required for an AI model to do AI R&D into two broad categories: software engineering, where the model writes code for research experiments and training runs, and research taste, where AI models decide which experiments are worth trying.
asserted
experiments → break → experiments
For software engineering, AI capabilities appear to be increasing exponentially.
asserted
capabilities → appear → engineering
When AI models are evaluated against how long it would take humans to complete the same engineering tasks — so-called “time horizons” — their capabilities seem to be doubling every 7 months.2
These capability improvements apply to the software engineering tasks required for AI R&D.
asserted
improvements → evaluate → R&D.
In a long-running experiment, researchers at Anthropic have found that their models now significantly outperform humans under a fixed time budget on an AI R&D task focused on speeding up AI model training.
asserted
models → run → training
For research taste, frontier AI models also show signs of fast improvement, though the evidence is less clear.
asserted
evidence → show → improvement
According to Anthropic’s research, its models have rapidly become proficient at solving “open-ended” tasks that Anthropic’s technical staff work on, where the model must solve problems with no clear specification by generating multiple approaches and then deciding among them.
uncertain
model → accord → them
In a separate open-ended research project where the AI was tasked with proposing and testing hypotheses about an open problem in AI safety, Anthropic’s models significantly outperformed two human researchers (97% performance improvement vs. 23%) when given a similar time budget (5 to 7 days).
asserted
models → task → budget
…and 61 more, not listed.