Unit content
Descriptive statistics
A dataset is a collection of observed values. Descriptive statistics summarize those values without yet making claims about a larger population.
Suppose the observations are
$$x_1,x_2,\ldots,x_n.$$
Measures of center
The sample mean is
$$\bar x=\frac1n\sum_{i=1}^n x_i.$$
The median is the middle value after ordering the data, or the average of the two middle values when $n$ is even.
The mean uses every observation and is sensitive to extreme values. The median depends on order and is often more resistant to outliers.
Measures of spread
The range is the difference between the largest and smallest values.
A common measure of spread around the mean is the sample variance
$$s^2=\frac1{n-1}\sum_{i=1}^n(x_i-\bar x)^2,$$
with sample standard deviation
$$s=\sqrt{s^2}.$$
Quantiles and quartiles
A quantile divides ordered data according to position. The first and third quartiles, $Q_1$ and $Q_3$, delimit the middle half of the observations. Their difference
$$\operatorname{IQR}=Q_3-Q_1$$
is the interquartile range.
Looking at the distribution
Tables and plots such as histograms, box plots and dot plots reveal features a single summary number cannot: skewness, clusters, gaps, outliers and multiple modes.
Describing data well therefore means combining numerical summaries with the shape and context of the observations.