Unit content
Hypothesis testing
A hypothesis test asks how surprising the observed data would be if a specified statistical claim were true.
The null hypothesis, written $H_0$, is the claim used to generate the reference sampling distribution. The alternative hypothesis, $H_1$ or $H_A$, describes the competing possibility of interest.
Test statistic
The sample is reduced to a test statistic whose behavior under $H_0$ is known or approximated.
A statistic far into the tail of its null distribution represents data that are difficult to reconcile with the null model.
The p-value
The p-value is the probability, assuming $H_0$ is true, of obtaining a test statistic at least as incompatible with $H_0$ as the one observed.
A small p-value means the observed data would be unusual under the null model. It is not the probability that $H_0$ is true.
Significance level
A significance level $\alpha$ is chosen before interpreting the result. If
$$p\le\alpha,$$
the null hypothesis is rejected according to that decision rule.
Failure to reject $H_0$ does not prove it true; it means the data did not provide sufficiently strong evidence against it under the chosen procedure.
Type I and Type II errors
A Type I error occurs when a true null hypothesis is rejected. Its probability is controlled by the significance level under the test assumptions.
A Type II error occurs when the test fails to reject a false null hypothesis. Its probability depends on the true alternative, sample size and test design.
The probability of correctly rejecting a particular false null is the power of the test.
Statistical and practical significance
A tiny effect can produce a small p-value with enough data, while an important effect may remain uncertain in a small sample. Hypothesis testing measures evidence relative to a model; practical importance must be judged from effect sizes, uncertainty and context.