How can estimating three unrelated things together beat estimating each one alone?
Estimating wheat yield, Wimbledon crowds and a candy bar's weight all at once beats estimating them one by one, on average: a result Charles Stein found in 1955.
▶ Start the storySuppose you must estimate three unrelated numbers: the US wheat yield for 1993, the number of spectators at Wimbledon in 2001, and the weight of a randomly chosen candy bar. You have one measurement of each, blurred by independent random noise. The obvious move is to take each measurement as the estimate of the thing it measures. Numerous methods for building estimators, including maximum likelihood and least squares, all lead to exactly that rule.
Stein's example says the obvious move is not the best. When three or more parameters are estimated simultaneously, there exist combined estimators more accurate on average, with lower expected mean squared error, than any method that handles the parameters separately. Charles Stein of Stanford discovered it in 1955. So a better estimate of the vector of three can be had, on average, by using the three unrelated measurements together.
Ordinary: one at a time
- Use each measurement as its own estimate
- Maximum likelihood and least squares all give this rule
Stein: all together
- Shrink the estimates towards a common point
- Lower total error on average, with three or more parameters
The catch is in the words "on average" and "the vector". The new estimator is no better for the wheat yield by itself. What it gives is an estimator for the vector of all three means with a reduced total risk, and any particular set of estimates will not necessarily beat the measured values. The intuitive explanation is that optimizing the mean-squared error of a combined estimator is not the same as optimizing the errors of separate estimators of each parameter.
The improved estimators work by shrinkage: a naive or raw estimate is improved by combining it with other information, pulled towards a common point. Many standard estimators can be improved in mean squared error by shrinking them, because the improvement from a narrower confidence interval can outweigh the worsening from biasing the estimate.
Quiz me
0/3
Recap
Pulling noisy estimates toward a common point lowers the total error on average, even though any one estimate can get worse.
💡 A trick to remember it · Three noisy guesses, one shared anchor: tug them all toward it and the team's total error shrinks.
Surprising fact · Three unrelated quantities, such as wheat yield, Wimbledon attendance and a candy bar's weight, are better estimated together than separately, on average.
Connects to
- 🗳️ How can a poll of a thousand people speak for millions, when one of 2.4 million got it wrong?
- 🔔 Why does the bell curve keep showing up, even when nothing is bell-shaped?
- 🔴 How did a checkers program teach itself to beat a checkers master?
- 🧀 Why should you never be 100% sure of anything?
- ↩️ Why is a brilliant performance usually followed by a worse one?
Sources (4)
No source, no claim. Every fact in this lesson (16 claims) cites at least one of these.