Learning path

Full curriculum

Full curriculum

Unit content

Populations, samples, parameters and statistics

Statistical questions usually concern a larger collection than the observations we can conveniently measure.

The population is the full collection of units or outcomes we want to understand. A sample is the subset actually observed.

Parameters and statistics

A numerical property of a population is a parameter. Examples include a population mean $\mu$ or population proportion $p$.

A numerical summary calculated from a sample is a statistic, such as the sample mean $\bar x$ or sample proportion $\hat p$.

Parameters are fixed properties of the chosen population, although they may be unknown. Statistics vary from sample to sample.

Example

Suppose we want the average commuting time of all employees in a company. The commuting times of all employees form the population, and the true average is a population parameter.

If $100$ employees are sampled and their average commuting time is calculated, that average is a statistic used to learn about the unknown parameter.

Census and sampling

A census attempts to observe every member of the population. A sample observes only part of it.

Sampling is often cheaper or more practical, but the connection between sample and population depends on how the sample was obtained.

Target population and sampling units

The population must be defined carefully. A study about “university students”, for example, needs to specify which universities, which time period and which students are included.

The individual objects selected for observation are the sampling units. Statistical inference is meaningful only relative to the population the sampling process can reasonably represent.