Unit content
Maximum a posteriori estimation as regularized inference
Bayes' theorem combines a likelihood with a prior distribution:
$$p(\theta\mid D)\propto p(D\mid\theta)p(\theta).$$
A maximum a posteriori estimate chooses the parameter value that maximizes the posterior density:
$$\theta_{MAP}=\arg\max_\theta p(\theta\mid D).$$
Taking negative logarithms turns this into minimization:
$$\theta_{MAP}=\arg\min_\theta\left[-\log p(D\mid\theta)-\log p(\theta)\right].$$
The negative log-likelihood acts as a data-fit objective, while the negative log-prior acts as a regularization term.
A Gaussian prior on parameters produces an $L_2$-type quadratic penalty. A Laplace prior produces an $L_1$-type penalty.
The same penalized optimization formula can be used without a probabilistic interpretation. MAP adds a specific meaning: the penalty represents the logarithm of a prior distribution over parameters.