A few days ago, I published a guest post by Tim Fist and Saif Khan of the Institute for Progress, discussing the question of whether we should deliberately try to slow down the rate of AI progress:
asserted
we → publish → progress
The authors promised a raft of specific policy recommendations, and they didn’t disappoint.
asserted
they → promise → recommendations
Did you know that 23 is my lucky number?
asserted
23 → know → ?
In our last post, we evaluated the claims of a recent open letter by AI company employees calling for governments to “pace” frontier AI development.
uncertain
governments → evaluate → development
Rapid progress towards fully automated AI R&D has empirical support, but it’s less clear how much it will accelerate AI capabilities or pose severe risks.
asserted
it → automate → risks
Despite substantial uncertainty, we believe some preparatory policy action is warranted.
asserted
action → believe → uncertainty
This follows both from how serious the possible direct risks are and the risk that political backlash to AI-driven disruptions results in poorly-reasoned policy measures, such as broad bans on new data centers.
asserted
backlash → follow → centers
If “pacing” becomes necessary, we think it should consist of two parts: first, specifying thresholds for when automated AI R&D is likely to pose severe risks; and second, if a threshold is exceeded, incentivizing AI companies to reallocate resources away from the most risky research, and towards activities that make further automation safer, or diffuse the benefits of existing AI faster.
asserted
automation → become → AI
Without preparation now, however, our preferred pacing strategy will be impossible to implement.
asserted
strategy → prefer → preparation
In this post, we’ll describe how the US can concretely prepare for the further automation of AI R&D and the risks it entails.
asserted
it → describe → R&D
Still, we aren’t certain whether the benefits of pacing outweigh the downsides, especially given the risk that government regulation is implemented counterproductively.
asserted
regulation → outweigh → risk
So to make policy preparation as targeted and low-regret as possible, we think any intervention should meet the following five criteria:
Target only AI development activities that could lead to serious and irreversible harms.
Minimize any slowdown in the diffusion of existing AI capabilities, and ideally accelerate it.
Impose low costs, or deliver clear benefits, even if automated AI R&D and its attendant risks prove unlikely.
uncertain
R&D → make → benefits
Avoid establishing a new regulatory apparatus that is likely to be misused (e.g., by concentrating power in a small set of companies).
asserted
that → establish → companies
A surprisingly wide range of policy moves meet these criteria.
asserted
range → meet → criteria
We’ve identified 23 of them, and they span 7 areas:
asserted
they → identify → areas
Transparency: giving the government and public more visibility into automated AI R&D
State capacity: improving the government’s ability to understand and respond to automated AI R&D
Risk management: developing a risk management strategy for automated AI R&D that accelerates defensive and commercial AI use
Verification: accelerating the development of AI verification technologies to enable agreements between mutually distrustful parties
Resilience: accelerating the development of technologies that improve society’s ability to withstand and recover from AI-driven disruptions
Competition with China: extending the US AI lead over China to buy more time to manage risks and increase US leverage in international negotiations
Diplomacy: creating option value for international cooperation to manage the risks of automated AI R&D
asserted
cooperation → give → R&D
In the rest of this post, we’ll explain why we think policy action is justified across each area and give specific recommendations for each.
asserted
action → explain → each
For more details on this and the material from yesterday’s post, you can read our full report here.
asserted
you → read → report
Transparency
AI now regularly makes impressive breakthroughs in math and is superhuman at many aspects of software development and cybersecurity, with capabilities doubling every 7 and 5 months, respectively.
asserted
capabilities → make → development
But outside of frontier AI companies, how exactly the drivers of AI progress — e.g., curating more and better data, improving training algorithms, applying more reinforcement learning, simply making the model bigger — are unlocking AI capabilities
asserted
model → curate → capabilities
The same is true for many of the crucial questions surrounding AI R&D automation.
asserted
same → surround → automation
Much of the best information remains inside company walls.1 Given that automated R&D could rapidly accelerate AI progress with little warning, this dynamic could leave the government and public unprepared.
uncertain
dynamic → remain → government
It doesn’t have to be this way.
asserted
It → have → ?
If we want society to respond well, we’ll need much more information about what’s going on at the frontier of AI.
asserted
what → want → AI
More transparency could help us understand the science behind automated AI R&D as well as what it looks like within specific companies (e.g., how much they’re automating, what policies they use to manage the risks, and any R&D-related incidents).
uncertain
they → help → risks
A recent NVIDIA-led letter supported open-weight models from an open science perspective — a valuable goal.
asserted
letter → lead → goal
Building on the findings of CSET and the Elasticity Institute, we suggest transparency measures for information in five categories:
The science of general AI progress.3
AI’s ability to automate specific AI R&D tasks.4
Company progress toward automating AI R&D.5
Company AI R&D automation risk management practices and incidents.6
Company “model behavior specifications,” documents that describe the values and principles an AI model is trained to follow (e.g., OpenAI’s Model Spec for its GPT models and Anthropic’s Constitution for its Claude models).7
The vast majority of this information is not subject to disclosure requirements, and so it is either disclosed voluntarily in a limited way or not at all.8
asserted
it → build → way
Although some particularly sensitive information may only be suitable for disclosure to the US government, we generally recommend transparency measures that involve public disclosure.
uncertain
that → recommend → disclosure
As AI companies automate more of their R&D, this information would enhance public and policymaker understanding, improve policy responses, and enable the broader scientific community to do better work on alignment and security.
1.
asserted
information → automate → alignment
companies and relevant industry bodies should publicly share information relevant to trends and risks in AI R&D automation
asserted
companies → share → automation
The categories of information outlined above would improve policymakers’ and the public’s understanding of the extent and nature of AI R&D automation at frontier AI companies, enabling outside experts to model, project, and publish on associated trends and impacts.
asserted
categories → outline → trends
In turn, this would improve policy responses and bolster the broader scientific community’s work on alignment and security.
asserted
this → improve → alignment
Disclosing information about the science of AI and general AI capabilities would be a return to the historical norms of open scientific publication in the US AI industry.
asserted
information → disclose → industry
As recently as 2020, OpenAI published detailed information on the architecture, training recipe, and data of GPT-3, while in 2022, Google published detailed scaling laws showing how training compute, data, and model architecture correlate with AI model capabilities.
asserted
compute → publish → capabilities
To reestablish this norm, employees at frontier AI companies should encourage their leadership to publicly share information in the categories outlined above, and advocate for reestablishing broader industry norms, including through industry bodies such as the Frontier Model Forum.
asserted
employees → reestablish → Forum
We believe commercial and geopolitical concerns over the sensitivity of this information are manageable.
asserted
concerns → believe → information
Industry-level technical metrics related to the science of AI and general AI capabilities are likely well understood by most or all frontier AI companies in the US and China.
asserted
metrics → relate → US
They are therefore unlikely to alter the balance of AI capabilities between US AI companies and between the US and China.
asserted
They → alter → US
Additionally, specific company-level AI R&D automation activities are of extraordinary public interest, such that improving the US government’s and public’s ability to mount a policy response outweighs concerns over commercial sensitivity.
asserted
improving → improve → sensitivity
Of course, some specific information on cutting-edge breakthroughs — particularly where non-US companies lack comparable knowledge — will be crucial to US strategic interests and AI leadership, and may thus be less suited to public disclosure.
uncertain
companies → lack → disclosure
…and 342 more, not listed.