Does Forecasting Have Room At The Top?

Astral Codex Ten · collected 2026-08-18 · by Scott Alexander commentary
Read the original at Astral Codex Ten ↗

Summary

The article discusses the potential limitations of superforecasting, or predicting future events, by humans versus artificial intelligence (AI). A study coauthored in 2010 found that prediction markets only outperformed simple statistical models by 3-6% in certain areas such as sports and movie box office predictions. This has led some to suggest that there may be a fundamental limit to the predictability of world events, beyond which AI will plateau at or slightly above human levels of accuracy. The article's author expresses skepticism about this view, citing the vast improvement in prediction markets over previous statistical models.
Written by the local model on 2026-08-20, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
121
claim-shaped sentences
Uncertain
17%
20 of 121 hedged
Leaning
not scored
needs a local LLM pass
Publisher trust
not scored
Commentary is not rated for newsroom trust
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-08-20 · source text last changed 2026-08-20 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Scott Alexander, from the blog Astral Codex Ten, discusses the limitations of forecasting in various domains. Daniel Reeves argues that humans have already approached a fundamental limit on predicting world events, and therefore AI systems will plateau at or slightly above human accuracy levels. This is based on a 2010 study by Reeves and others, which found that prediction markets outperformed simple statistical models by only 3-6% in various tasks such as predicting sports game outcomes and movie box office receipts. If this trend continues, it would suggest that AI systems will not surpass human forecasting abilities significantly, but rather maintain a similar level of accuracy. This challenges the idea of "superforecasting" or "ultraforecasting", which implies continuous improvement in AI's ability to predict the future.

Written for “Predictive Analytics Career Growth” on 2026-08-31, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Score +0.35 Confidence medium
Leaning score +0.35 for article 1104 (medium confidence, 1 verified quote) · logged 2026-08-25

Story

📰 Predictive Analytics Career Growth
Technology · 1 article(s) covering the same event. This is the one the site leads with.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

Nothing to compare against. No article is close enough to this one for the pipeline to have linked or judged the pair.

Publisher

Astral Codex Ten · 25 article(s) · 2 correction(s) detected

Commentary. The three signals behind a trust score all measure a newsroom's record with its own reporting, so they are not computed for this source. How trust is scored.

No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Scott Alexander
25 article(s) here · 1 carrying a prediction
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
🔮 This year’s survey will probably take 30 - 45 minutes.
2026-08-27 · assertive framing · Take The 2026 ACX Survey
🔮 It’s not necessarily wrong to deprioritize a topic because it vaguely reminds you of something that you have negative affect toward, but I would prefer these people admit they’re dismissing/ignoring it rather than claim to be engaging with it.
🔮 I don’t know, but he may have misinterpreted my original tweet as saying income doesn’t matter at all, in which case this data functions as an effective rebuttal.
2026-08-25 · assertive framing · Re: Re: Re: Pritchard On Liberal Happiness
🔮 [This is one of the finalists in the 2026 book review contest, written by an ACX reader who will remain anonymous until after voting is done.
🔮 Matt says his highest priority colleges that don’t have an organizer yet are MIT, Claremont, Northwestern, Carnegie Mellon, Amherst, Notre Dame, Duke, CalTech, Georgetown, UCLA, Dartmouth, Vanderbilt, Harvard, NYU, Williams, GWU, American University, Rutgers, and UNC Chapel Hill.
2026-08-24 · assertive framing · Open Thread 448
🔮 We (ACX and an effective altruist organization called Nest who are sponsoring this) will give you free advertising and cover your costs.
2026-08-21 · assertive framing · College EA Meetups Everywhere: Call For Organizers
🔮 I remember when arguments about AI were things like “Sure, if you could magically get billions of dollars of compute, and magically scale AI up a thousand times, and magically get rid of hallucinations…” and the whole implausibility hinged on the word “magic”!
🔮 On average, commenters will end up spotting evidence that around two or three of the links in each links post are wrong or misleading.
2026-08-19 · assertive framing · Links For July 2026 (Part 2)
🔮 There are many reasons to believe in such assistance: Napoleon’s rise from Corsican nobody to Emperor and would-be world conqueror was a bizarre deviation from the usual pattern of history.
2026-08-19 · assertive framing · Why I'm Staying Out Of The Substack Religion Debate
Also by Scott Alexander
Open Thread 449
2026-08-31 · Astral Codex Ten
Hidden Open Thread 448.5
2026-08-28 · Astral Codex Ten
Take The 2026 ACX Survey
2026-08-27 · Astral Codex Ten
Nothing else under this byline is closely related to this article, so these are simply their most recent.
All 25 articles by Scott Alexander →

Topics

No topics tagged.

Subjects

Reeves PERSON · 3× Daniel Reeves1 PERSON · 1× Goel PERSON · 1×

Narrative

Reeves et al argue against this interpretation in their paper: That the poor discrimination of the baseline model carries so small a penalty in terms of RMSE is due in part to specific features of the NFL (e.g., salary caps) that ensure that most games are played between closely matched teams, and hence are decided with probabilities close to 50%.
framing: assertive · carried by 1 article(s) · first seen 2026-08-20
🔮 Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered.
2026-08-20 · Astral Codex Ten
Does Forecasting Have Room At The Top? · assertive framing

Claims (121 extracted, 20 hedged)

Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered. asserted
technologies → predict → wars
Over the past few years, it went from an obscure academic subfield to a multibillion dollar industry in the form of prediction markets. asserted
it → go → markets
More recently, AI superforecasters have come close to the accuracy of top humans, and their performance is rising rapidly. asserted
performance → come → humans
In a year or two, we’ll see one of the following patterns: Either humans have already come close to some fundamental limit on the predictability of world events - in which case AIs will plateau at or slightly above the human level - or the trend line will continue until AIs are far beyond top humans. asserted
AIs → see → humans
By analogy to superintelligence, the natural term for the second situation would be “superforecasting”; since that’s already in use, we can cringely call it “ultraforecasting”. asserted
we → ’ → it
Daniel Reeves1 makes the case for scenario A here. asserted
Reeves1 → make → A
He describes a study he coauthored in 2010, which found that, on a variety of questions related to sports games and movie box office receipts2, prediction markets only outperformed simple boring statistical models by 3-6%. asserted
markets → describe → %
Maybe those statistical models are close to the best that it’s possible to do; the rest is what the mathematicians call aleatoric uncertainty - irreducible complexity downstream of chaotic systems that entirely resist modeling. asserted
that → ’ → modeling
This post isn’t meant to be a decisive refutation, but rather a description of why I’m still about 70-30 expecting Scenario B. Slightly Contra Goel, Reeves, et al Reeves’ study claims that the prediction markets of 2010 only beat dumb statistical models by 3% (for sports) to 6% (for movies). uncertain
markets → mean → movies
But these percentages aren’t real win-loss percentages; they’re variation in a quantity called root mean-squared error. asserted
they → ’re → quantity
One way to get a feel for this quantity is that the dumb statistical model for sports (home team advantage + win-loss record) beat an even dumber statistical model (home team advantage only) by 0.8 percentage points, and the prediction market beat the first (better) statistical model by another 0.4 pp. asserted
market → get → pp
So the effect of going from a statistical model to a prediction market is half as large as the effect of knowing which two teams were playing and how good they are! asserted
they → go → effect
Why can we frame this same result as either very small or very large? asserted
we → frame → result
Sports are optimized against prediction3. asserted
Sports → optimize → prediction3
If there were a fully predictable sport (eg heightball, where all athletes line up in a row and the tallest one wins), nobody would watch it. asserted
nobody → be → it
Instead, we go to absurd lengths to keep the outcome uncertain. asserted
we → go → outcome
Salary caps, draft systems, etc try to ensure that all teams have exactly equal talent. asserted
teams → try → talent
Commercial incentives and ceiling effects ensure that they have exactly equal training. asserted
they → ensure → training
Then an exactly-equal number of these exactly-equally-talented-and-trained people are placed in exactly-identical positions on a perfectly-symmetrical field and told to hit/kick/throw a ball which is placed exactly equidistant between both of them. asserted
which → train → them
It’s funny for me to describe it this way, because obviously this is what we want (to “keep things fair”), but it’s all designed for prediction-resistance. asserted
it → ’ → resistance
Given the difficulty of the domain, even very large relative advances in prediction look small in absolute terms. asserted
advances → give → terms
Maybe we should look at Reeves’ other example, movie box office receipts. asserted
we → look → example
Here the markets did slightly better, getting a 6% improvement. asserted
markets → do → improvement
Box office receipts differ by orders of magnitude (some movies make $100,000, others make $100 million), so the paper puts this on a log scale. asserted
paper → differ → scale
6% improvement on a log scale is already starting to sound pretty good. asserted
improvement → start → scale
And again, it all ends up coming down to what we compare it to. asserted
we → end → it
Here the super-dumb model is that all movies make $8.1 million, the takings of the exact average movie. asserted
movies → make → movie
The slightly-less-dumb model then adds the number of screens that the movie is showing on and the amount of Google search traffic for the movie! asserted
movie → add → movie
For example, a random indie film might be showing on three screens in the entire country, and a Disney blockbuster might be showing on ten thousand. uncertain
blockbuster → show → thousand
The challenge the paper gives prediction markets is to significantly improve on knowing whether a film is an indie film or a Disney blockbuster, plus knowing how many people are interested in seeing it, and it has to do this on a log scale! asserted
it → give → scale
No wonder the relative improvement number comes out looking slightly anemic. asserted
number → come → ?
I would summarize this section as: we shouldn’t expect miracles, but this doesn’t rule out further normal-sized gains4. asserted
this → summarize → gains4
Applying The Non-Miracle Rule To Geopolitics Let’s return to the picture we looked at above: This graph is loosely based on Metaculus’ AIs vs. humans results: …and these are mostly on geopolitical questions. asserted
these → apply → questions
What can Reeves’ model tell us about these? asserted
model → tell → these
In one sense, it can already be proven not to apply. asserted
it → prove → sense
In the sports analysis, the spread between base rate (guessing 50% on everything) and smart humans (the prediction markets) was 0.04 Brier score points. asserted
spread → guess → everything
So Metaculus’ superforecasters have already removed 3x more uncertainty than a naive port of Reeves’ model would suggest exists! asserted
port → remove → model
Here we return to the idea of sports as a uniquely unpredictable domain. asserted
we → return → domain
To give a trivial example, the easiest sports question (will the best team in the league beat the worst team in the league?) might still only be 90-10 (upsets happen). uncertain
upsets → give → league
But the easiest geopolitical question (maybe “will the US bomb Canada in the next month?”) could easily be 99-1 or more. uncertain
US → bomb → month
…and 81 more, not listed.
💬 Give feedback
🕘 History 🎫 Support