Learning path

Full curriculum

Full curriculum

Unit content

Logistic regression for binary classification

Logistic regression models the probability of a binary outcome using a linear score.

For feature vector $x$, form

$$z=\beta_0+\beta^Tx.$$

The logistic, or sigmoid, function converts this unrestricted score into a probability:

$$p(y=1\mid x)=\sigma(z)=\frac{1}{1+e^{-z}}.$$

Equivalently, the log-odds are linear:

$$\log\frac{p}{1-p}=\beta_0+\beta^Tx.$$

If $z=\ln 3$, then the odds are $3:1$ and the predicted probability is $3/4$.

The coefficients are commonly fitted by maximizing the Bernoulli likelihood, which is equivalent to minimizing binary cross-entropy. A positive coefficient means that increasing that feature, while holding the others fixed, increases the predicted log-odds of the positive class.

A threshold such as $p\ge 0.5$ can convert probabilities into labels, but the threshold is a decision choice rather than part of the probabilistic model itself. Different applications can use the same fitted probabilities with different thresholds.