Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered.
asserted
technologies → predict → wars
Over the past few years, it went from an obscure academic subfield to a multibillion dollar industry in the form of prediction markets.
asserted
it → go → markets
More recently, AI superforecasters have come close to the accuracy of top humans, and their performance is rising rapidly.
asserted
performance → come → humans
In a year or two, we’ll see one of the following patterns:
Either humans have already come close to some fundamental limit on the predictability of world events - in which case AIs will plateau at or slightly above the human level - or the trend line will continue until AIs are far beyond top humans.
asserted
AIs → see → humans
By analogy to superintelligence, the natural term for the second situation would be “superforecasting”; since that’s already in use, we can cringely call it “ultraforecasting”.
asserted
we → ’ → it
Daniel Reeves1 makes the case for scenario A here.
asserted
Reeves1 → make → A
He describes a study he coauthored in 2010, which found that, on a variety of questions related to sports games and movie box office receipts2, prediction markets only outperformed simple boring statistical models by 3-6%.
asserted
markets → describe → %
Maybe those statistical models are close to the best that it’s possible to do; the rest is what the mathematicians call aleatoric uncertainty - irreducible complexity downstream of chaotic systems that entirely resist modeling.
asserted
that → ’ → modeling
This post isn’t meant to be a decisive refutation, but rather a description of why I’m still about 70-30 expecting Scenario B.
Slightly Contra Goel, Reeves, et al
Reeves’ study claims that the prediction markets of 2010 only beat dumb statistical models by 3% (for sports) to 6% (for movies).
uncertain
markets → mean → movies
But these percentages aren’t real win-loss percentages; they’re variation in a quantity called root mean-squared error.
asserted
they → ’re → quantity
One way to get a feel for this quantity is that the dumb statistical model for sports (home team advantage + win-loss record) beat an even dumber statistical model (home team advantage only) by 0.8 percentage points, and the prediction market beat the first (better) statistical model by another 0.4 pp.
asserted
market → get → pp
So the effect of going from a statistical model to a prediction market is half as large as the effect of knowing which two teams were playing and how good they are!
asserted
they → go → effect
Why can we frame this same result as either very small or very large?
asserted
we → frame → result
Sports are optimized against prediction3.
asserted
Sports → optimize → prediction3
If there were a fully predictable sport (eg heightball, where all athletes line up in a row and the tallest one wins), nobody would watch it.
asserted
nobody → be → it
Instead, we go to absurd lengths to keep the outcome uncertain.
asserted
we → go → outcome
Salary caps, draft systems, etc try to ensure that all teams have exactly equal talent.
asserted
teams → try → talent
Commercial incentives and ceiling effects ensure that they have exactly equal training.
asserted
they → ensure → training
Then an exactly-equal number of these exactly-equally-talented-and-trained people are placed in exactly-identical positions on a perfectly-symmetrical field and told to hit/kick/throw a ball which is placed exactly equidistant between both of them.
asserted
which → train → them
It’s funny for me to describe it this way, because obviously this is what we want (to “keep things fair”), but it’s all designed for prediction-resistance.
asserted
it → ’ → resistance
Given the difficulty of the domain, even very large relative advances in prediction look small in absolute terms.
asserted
advances → give → terms
Maybe we should look at Reeves’ other example, movie box office receipts.
asserted
we → look → example
Here the markets did slightly better, getting a 6% improvement.
asserted
markets → do → improvement
Box office receipts differ by orders of magnitude (some movies make $100,000, others make $100 million), so the paper puts this on a log scale.
asserted
paper → differ → scale
6% improvement on a log scale is already starting to sound pretty good.
asserted
improvement → start → scale
And again, it all ends up coming down to what we compare it to.
asserted
we → end → it
Here the super-dumb model is that all movies make $8.1 million, the takings of the exact average movie.
asserted
movies → make → movie
The slightly-less-dumb model then adds the number of screens that the movie is showing on and the amount of Google search traffic for the movie!
asserted
movie → add → movie
For example, a random indie film might be showing on three screens in the entire country, and a Disney blockbuster might be showing on ten thousand.
uncertain
blockbuster → show → thousand
The challenge the paper gives prediction markets is to significantly improve on knowing whether a film is an indie film or a Disney blockbuster, plus knowing how many people are interested in seeing it, and it has to do this on a log scale!
asserted
it → give → scale
No wonder the relative improvement number comes out looking slightly anemic.
asserted
number → come → ?
I would summarize this section as: we shouldn’t expect miracles, but this doesn’t rule out further normal-sized gains4.
asserted
this → summarize → gains4
Applying The Non-Miracle Rule To Geopolitics
Let’s return to the picture we looked at above:
This graph is loosely based on Metaculus’ AIs vs. humans results:
…and these are mostly on geopolitical questions.
asserted
these → apply → questions
What can Reeves’ model tell us about these?
asserted
model → tell → these
In one sense, it can already be proven not to apply.
asserted
it → prove → sense
In the sports analysis, the spread between base rate (guessing 50% on everything) and smart humans (the prediction markets) was 0.04 Brier score points.
asserted
spread → guess → everything
So Metaculus’ superforecasters have already removed 3x more uncertainty than a naive port of Reeves’ model would suggest exists!
asserted
port → remove → model
Here we return to the idea of sports as a uniquely unpredictable domain.
asserted
we → return → domain
To give a trivial example, the easiest sports question (will the best team in the league beat the worst team in the league?) might still only be 90-10 (upsets happen).
uncertain
upsets → give → league
But the easiest geopolitical question (maybe “will the US bomb Canada in the next month?”) could easily be 99-1 or more.
uncertain
US → bomb → month
…and 81 more, not listed.