Learning path

Full curriculum

Full curriculum

Unit content

Maximum a posteriori estimation as regularized inference

Bayes' theorem combines a likelihood with a prior distribution:

$$p(\theta\mid D)\propto p(D\mid\theta)p(\theta).$$

A maximum a posteriori estimate chooses the parameter value that maximizes the posterior density:

$$\theta_{MAP}=\arg\max_\theta p(\theta\mid D).$$

Taking negative logarithms turns this into minimization:

$$\theta_{MAP}=\arg\min_\theta\left[-\log p(D\mid\theta)-\log p(\theta)\right].$$

The negative log-likelihood acts as a data-fit objective, while the negative log-prior acts as a regularization term.

A Gaussian prior on parameters produces an $L_2$-type quadratic penalty. A Laplace prior produces an $L_1$-type penalty.

The same penalized optimization formula can be used without a probabilistic interpretation. MAP adds a specific meaning: the penalty represents the logarithm of a prior distribution over parameters.