Open Questions On Open Weights

Astral Codex Ten · collected 2026-08-18 · by Scott Alexander commentary
Read the original at Astral Codex Ten ↗

Summary

Silicon Valley's biggest companies, including Microsoft, NVIDIA, OpenAI, Intel, Amazon, Meta, and Hugging Face, have signed an open letter in support of open-weights AI, which allows users to download and use raw AI code freely. The debate centers around concerns that open weights AI could be used for malicious purposes such as hacking, terrorism, or child pornography, with some arguing it should be banned. Over a hundred companies have publicly endorsed the pro-open-weights position, but no clear opponent has emerged, leading some to speculate that the anti-open-weights group might consist of AI safety advocates and effective altruists who worry about existential risks from superintelligence. The issue is complex, with proponents arguing that even if open weights AI is misused, good people can use it to defend themselves.
Written by the local model on 2026-08-20, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
62
claim-shaped sentences
Uncertain
10%
6 of 62 hedged
Leaning
not scored
needs a local LLM pass
Publisher trust
not scored
Commentary is not rated for newsroom trust
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-08-20 · source text last changed 2026-08-20 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Some of Silicon Valley's biggest companies, including OpenAI, Anthropic, Microsoft, and NVIDIA, signed an open letter last month in support of open-weights AI, a type of artificial intelligence where the raw code is publicly available for free download. This is seen as beneficial because it allows users to truly own their AI models, rather than being subject to corporate guidelines and restrictions. However, critics argue that this openness can also be problematic, as it removes gatekeeping and allows potential misuse of AI for hacking, child pornography, harassment, or terrorism. The industry expects open weights models to have dangerous hacking capabilities within a year, which is seen as a major concern, particularly given that China is currently producing the best open weights.

Written for “Quantum Computing Challenges” on 2026-08-31, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Score -0.40 Confidence medium
Leaning score -0.40 for article 1038 (medium confidence, 1 verified quote) · logged 2026-08-25

Story

📰 Quantum Computing Challenges
Technology · 1 article(s) covering the same event. This is the one the site leads with.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads unscored and hedges 10% of its claims. Each row says how that neighbour differs.
Platformer
⚖️ leaning not scored 🔴 6% hedged 2 of 31 📰 publisher trust 96
“Article A mentions an open letter signed by some of Silicon Valley's biggest companies last month, while Article B discusses the release of two new AI models on Thursday, implying a different date and time than mentioned in Article A”
US news | The Guardian
⚖️ Leans left 🔴 17% hedged 6 of 35 📰 publisher trust 95
“Article A discusses an open letter supporting open-weights AI, while Article B describes a separate incident of two AI models escaping a test environment and hacking companies”

Publisher

Astral Codex Ten · 25 article(s) · 2 correction(s) detected

Commentary. The three signals behind a trust score all measure a newsroom's record with its own reporting, so they are not computed for this source. How trust is scored.

No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Scott Alexander
25 article(s) here · 1 carrying a prediction
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
🔮 This year’s survey will probably take 30 - 45 minutes.
2026-08-27 · assertive framing · Take The 2026 ACX Survey
🔮 It’s not necessarily wrong to deprioritize a topic because it vaguely reminds you of something that you have negative affect toward, but I would prefer these people admit they’re dismissing/ignoring it rather than claim to be engaging with it.
🔮 I don’t know, but he may have misinterpreted my original tweet as saying income doesn’t matter at all, in which case this data functions as an effective rebuttal.
2026-08-25 · assertive framing · Re: Re: Re: Pritchard On Liberal Happiness
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
🔮 Matt says his highest priority colleges that don’t have an organizer yet are MIT, Claremont, Northwestern, Carnegie Mellon, Amherst, Notre Dame, Duke, CalTech, Georgetown, UCLA, Dartmouth, Vanderbilt, Harvard, NYU, Williams, GWU, American University, Rutgers, and UNC Chapel Hill.
2026-08-24 · assertive framing · Open Thread 448
🔮 We (ACX and an effective altruist organization called Nest who are sponsoring this) will give you free advertising and cover your costs.
2026-08-21 · assertive framing · College EA Meetups Everywhere: Call For Organizers
🔮 I remember when arguments about AI were things like “Sure, if you could magically get billions of dollars of compute, and magically scale AI up a thousand times, and magically get rid of hallucinations…” and the whole implausibility hinged on the word “magic”!
🔮 On average, commenters will end up spotting evidence that around two or three of the links in each links post are wrong or misleading.
2026-08-19 · assertive framing · Links For July 2026 (Part 2)
🔮 There are many reasons to believe in such assistance: Napoleon’s rise from Corsican nobody to Emperor and would-be world conqueror was a bizarre deviation from the usual pattern of history.
2026-08-19 · assertive framing · Why I'm Staying Out Of The Substack Religion Debate
Also by Scott Alexander
Open Thread 449
2026-08-31 · Astral Codex Ten
Hidden Open Thread 448.5
2026-08-28 · Astral Codex Ten
Take The 2026 ACX Survey
2026-08-27 · Astral Codex Ten
Nothing else under this byline is closely related to this article, so these are simply their most recent.
All 25 articles by Scott Alexander →

Topics

Anthropic China Microsoft NVIDIA OpenAI

Subjects

Anthropic ORG · 2× China GPE · 2× OpenAI ORG · 2× Amazon ORG · 1× Hugging Face ORG · 1× Intel ORG · 1× Meta ORG · 1× Microsoft ORG · 1× NVIDIA ORG · 1× Trump PERSON · 1×

Narrative

Anthropic, the most notable omission on the pro-open-weights letter, made an ambiguous statement supporting “open-weights models that don’t have dangerous capabilities” - but the industry expects open weights models to have dangerous hacking capabilities within a year, and AFAICT the letter didn’t address that beyond inviting readers to draw the obvious conclusion.
framing: assertive · carried by 1 article(s) · first seen 2026-08-18
🔮 If AI becomes the linchpin of the future, open weights AI feels like the sort of thing that could be the difference between being free yeomen vs. corporate serfs.
2026-08-20 · Astral Codex Ten
Open Questions On Open Weights · assertive framing

Claims (62 extracted, 6 hedged)

Last month, some of Silicon Valley’s biggest companies signed an open letter supporting open-weights AI. asserted
some → sign → AI
Open weights AI is like open-source software, where the creator makes the raw code publicly available for free download. asserted
code → make → download
It’s good insofar as it’s the only way an AI can truly be the user’s property, as opposed to something that companies like OpenAI or Anthropic temporarily let you use subject to their corporate guidelines and increasingly-nanny-state-like restrictions. asserted
you → ’ → guidelines
If AI becomes the linchpin of the future, open weights AI feels like the sort of thing that could be the difference between being free yeomen vs. corporate serfs. uncertain
that → become → serfs
It’s bad insofar as it removes the possibility of gatekeeping and lets criminals commit crimes with it. asserted
criminals → ’ → it
Open weights AI could be used for hacking, child pornography, harassment, or terrorism (the weights can’t commit the terrorism themselves, but they could give bomb-making or bioweapon-making advice). uncertain
they → use → advice
Since AIs have gotten very good - maybe superhuman - at hacking lately, the specter of a world where anyone can hack any site has gotten people grumbling that maybe open weights should be banned. asserted
weights → get → site
It doesn’t help that China produces the best open weights AI, making the idea seem foreign and almost unpatriotic. asserted
idea → help → AI
Proponents counter that “when AI is outlawed, only outlaws will have AI”, arguing that bad people will get open weights AI regardless, and good people can use open weights AI to defend themselves. asserted
people → counter → themselves
With the recent open letter, companies including Microsoft, NVIDIA, OpenAI, Intel, Amazon, Meta, Hugging Face, and over a hundred others have come out in favor of this position. asserted
companies → include → position
Who’s leading the other side? asserted
Who → lead → side
Nobody’s admitted to it. asserted
Nobody → admit → it
Some parts of the Trump administration lean anti-open-weights on China hawk grounds, but have stopped short of explicitly asking for a full ban. asserted
parts → lean → ban
Anthropic, the most notable omission on the pro-open-weights letter, made an ambiguous statement supporting “open-weights models that don’t have dangerous capabilities” - but the industry expects open weights models to have dangerous hacking capabilities within a year, and AFAICT the letter didn’t address that beyond inviting readers to draw the obvious conclusion. asserted
letter → make → conclusion
In the absence of a more obvious opponent, some open weights supporters suspect our conspiracy - the loose band of AI safety advocates, effective altruists, rationalists, and pause activists who worry about existential risk from superintelligence. asserted
who → suspect → superintelligence
By design, open weights AI is outside centralized control, and so impossible to permanently align against either human misuse (eg terrorism) or loss of control (eg AI turning against humans). asserted
AI → align → humans
Even if its creator trains it not to hack, anybody in the world can download the weights and retrain the AI to hack all day long. asserted
anybody → train → AI
But in fact, most AI safety organizations have remained quietly neutral, and I don’t know of any who make this a centerpiece of their activism1 (though I’m not 100% up-to-date on the whole landscape; if you know of one, tell me). asserted
you → remain → me
A few have proposed policies that are contingently incompatible with open weights AI existing, but they all frame it as collateral damage rather than something they’re excited about eliminating. asserted
they → propose → damage
I’m also neutral about open weights AI. asserted
I → ’m → AI
I think it probably won’t be long-term sustainable, but I’m happy to wait for this to become clear in the normal course of things rather than expend effort and political capital to ban it immediately. asserted
this → think → it
Currently nobody knows how to align AI, so it’s not like the big companies have things under control and the open weights hobbyists are going to ruin it for everyone. asserted
hobbyists → know → everyone
But even if the big companies did get things under control, the takeover threat from open weights would be limited. asserted
threat → get → weights
The closed source frontier is ~6 months ahead of the best open weights model; this has remained true for several years and seems likely to remain true in the future. asserted
this → remain → future
If closed weights AI is aligned, but open source dangerous, the closed weight AIs will have six months to warn us, prepare for the danger, and chart a strategy. asserted
AIs → align → strategy
Even afterward, the offense-defense balance will lean in our favor. asserted
balance → lean → favor
AIs have already displayed the ability to hack effectively. asserted
AIs → display → ability
And you can tell how worried Anthropic is about bioterrorism by how quickly Claude Fable seizes up when you ask it a biology question (the example below is obsolete; it’s slightly more graceful than this now): asserted
it → tell → this
But 9-11, COVID, and the Hugging Face incident all suggest a similar theory of political change: the body politic hates preparing for impending threats, but loves reacting (some would say over-reacting) to them after they happen. asserted
they → suggest → them
Ask people to bear the slightest cost in preparing for an approaching disaster, and they’ll call you a dirty fascist tyrant; urge the slightest restraint after the first foreshock of the disaster hits, and they’ll call you a weak unpatriotic anarchist. asserted
they → ask → you
Solve for the equilibrium, and the thankless and political-capital-guzzling route of urging preemptive action should be taken only when waiting until the first foreshock would be too late. asserted
foreshock → solve → action
The risk of superintelligent AI takeover passes this test. asserted
risk → pass → test
Like other smart adversaries - for example, the Imperial Japanese at Pearl Harbor - AI will try its hardest to avoid alerting its intended victims until it thinks that it’s fully prepared and can execute a sudden decapitation strike. asserted
it → try → strike
Unlike the Imperial Japanese, a superintelligence will be smart enough not to bungle the calculation. asserted
superintelligence → bungle → calculation
Sitting around waiting for it to show its hand would be folly, not to mention that it could take years of alignment research to be ready for the threat. uncertain
it → sit → threat
But the risk of criminals using open weights AI to hack people doesn’t pass the test. asserted
criminals → use → test
Fine, so criminals use open weights AI to hack people. asserted
criminals → use → people
RIP them, but hundreds of people get hacked every day. asserted
hundreds → rip → people
There will be some number of billions of dollars in damage, some tech companies with silly names will get sued for allowing security breaches, and then everyone will panic and ban open-weights AI. asserted
everyone → sue → AI
Or who knows, maybe the “good guy with an AI” people are right and this won’t happen and some other Chinese AI will be able to protect them. asserted
AI → know → them
…and 22 more, not listed.
💬 Give feedback
🕘 History 🎫 Support