Unit content
Likelihood and maximum likelihood estimation
A statistical model assigns probabilities or probability densities to possible observations according to unknown parameters. Once data has been observed, the same model can be viewed as a function of those parameters. This function is the likelihood.
For observations $x_1,\ldots,x_n$ modeled as independent with density $p(x\mid\theta)$,
$$L(\theta)=\prod_{i=1}^n p(x_i\mid\theta).$$
Likelihood is not a probability distribution over the parameter
The data are held fixed while $\theta$ varies. The likelihood compares how well different parameter values explain the observed data; it does not by itself say that the parameter is random.
Maximum likelihood
A maximum likelihood estimator (MLE) chooses
$$\hat\theta_{ML}=\arg\max_\theta L(\theta).$$
Because products of many probabilities can be inconvenient, the log-likelihood is usually maximized instead:
$$\ell(\theta)=\log L(\theta)=\sum_i\log p(x_i\mid\theta).$$
The logarithm preserves the maximizing parameter while turning products into sums.
The model determines the estimator
Different assumptions about the distribution of observations produce different likelihood functions and therefore different estimation rules. Maximum likelihood is a general principle; the familiar formulas that result depend on the chosen probabilistic model.