Unit content
Populations, samples, parameters and statistics
Statistical questions usually concern a larger collection than the observations we can conveniently measure.
The population is the full collection of units or outcomes we want to understand. A sample is the subset actually observed.
Parameters and statistics
A numerical property of a population is a parameter. Examples include a population mean $\mu$ or population proportion $p$.
A numerical summary calculated from a sample is a statistic, such as the sample mean $\bar x$ or sample proportion $\hat p$.
Parameters are fixed properties of the chosen population, although they may be unknown. Statistics vary from sample to sample.
Example
Suppose we want the average commuting time of all employees in a company. The commuting times of all employees form the population, and the true average is a population parameter.
If $100$ employees are sampled and their average commuting time is calculated, that average is a statistic used to learn about the unknown parameter.
Census and sampling
A census attempts to observe every member of the population. A sample observes only part of it.
Sampling is often cheaper or more practical, but the connection between sample and population depends on how the sample was obtained.
Target population and sampling units
The population must be defined carefully. A study about “university students”, for example, needs to specify which universities, which time period and which students are included.
The individual objects selected for observation are the sampling units. Statistical inference is meaningful only relative to the population the sampling process can reasonably represent.