Unit content
Neural networks and activation functions
A neural network builds a complex function by composing simple parameterized transformations.
A basic layer computes
$$\mathbf z=W\mathbf x+\mathbf b,$$
then applies an activation function componentwise:
$$\mathbf h=\phi(\mathbf z).$$
Stacking layers gives a nested function whose parameters are the weights and biases throughout the network.
Why nonlinear activations matter
If every layer were only linear, composing many layers would still produce one linear transformation. Nonlinear activations let the network represent relationships that cannot be reduced to one matrix multiplication.
Common activations include ReLU,
$$\operatorname{ReLU}(x)=\max(0,x),$$
and smooth functions such as sigmoid or tanh.
Hidden representations
Intermediate layer outputs are often called activations or hidden representations. Training determines which intermediate features are useful for reducing the final loss.
Output layers
The final transformation depends on the task. Regression may output real values directly; multiclass classification commonly converts logits into probabilities using softmax.
A neural network is therefore not a collection of biological neurons simulated literally. It is a differentiable function family built from repeated linear transformations and nonlinearities.