Unit content
Bootstrap resampling
The bootstrap approximates sampling variability by repeatedly resampling the observed dataset itself.
Given observations
$$x_1,\ldots,x_n,$$
a bootstrap sample contains $n$ observations drawn from these data with replacement. Some original observations may appear several times and others not at all.
To estimate uncertainty in a statistic $T$:
- draw a bootstrap sample;
- compute $T$ on that sample;
- repeat many times;
- examine the distribution of the resulting bootstrap statistics.
For example, from measured values $2,4,7,9$, one bootstrap sample might be $2,2,7,9$ and another $4,7,7,9$. Their means vary, approximating how the sample mean could vary under repeated sampling from the underlying population.
Bootstrap methods are useful when analytic sampling distributions are inconvenient, but they rely on the observed sample being informative about the population. Strong dependence or a tiny unrepresentative sample can make naive resampling misleading.
The same resampling idea also powers bagging: instead of using bootstrap replicates to estimate uncertainty, bagging fits a predictor to each replicate and aggregates the predictions.