Unit content
Logistic regression for binary classification
Logistic regression models the probability of a binary outcome using a linear score.
For feature vector $x$, form
$$z=\beta_0+\beta^Tx.$$
The logistic, or sigmoid, function converts this unrestricted score into a probability:
$$p(y=1\mid x)=\sigma(z)=\frac{1}{1+e^{-z}}.$$
Equivalently, the log-odds are linear:
$$\log\frac{p}{1-p}=\beta_0+\beta^Tx.$$
If $z=\ln 3$, then the odds are $3:1$ and the predicted probability is $3/4$.
The coefficients are commonly fitted by maximizing the Bernoulli likelihood, which is equivalent to minimizing binary cross-entropy. A positive coefficient means that increasing that feature, while holding the others fixed, increases the predicted log-odds of the positive class.
A threshold such as $p\ge 0.5$ can convert probabilities into labels, but the threshold is a decision choice rather than part of the probabilistic model itself. Different applications can use the same fitted probabilities with different thresholds.