← Articles

AI Recap, Sidebar 1: The Soft Threshold's Bargain

Why sigmoid units made gradient-based training practical, how saturation complicates learning in deep networks, and how later methods address gradient propagation.

Yesterday's post — Part 1 of this recap — introduced the vanishing gradient through a small calculation: sigmoid derivatives max out at $0.25$, and multiplying twenty of those factors gives about $10^{-12}$. That left me with two questions. I'd coded sigmoid neurons without asking why we used this particular S-shaped function. And when I read about Hochreiter's 1991 analysis of vanishing gradients, I wondered how it had affected the field. Had a mathematical explanation of the difficulty discouraged people from pursuing neural networks?

From a threshold to a sigmoid

The sigmoid maps any real number into the open interval $(0, 1)$:

$$\sigma(z) = \frac{1}{1 + e^{-z}}$$

Here $z$ is the neuron's pre-activation — the weighted sum of its inputs, $z = w \cdot x + b$. As that sum increases, the output follows an S-shaped curve:

  • $z$ very negative → output $\approx 0$
  • $z = 0$ → output exactly $0.5$
  • $z$ very positive → output $\approx 1$

Early artificial-neuron models used a hard threshold: sum the inputs, output 1 above the threshold and 0 otherwise. This gave a simplified description of a neuron firing or remaining silent. The sigmoid provides a smooth transition between those two output levels. Changing the bias $b$ shifts the input sum at which the output crosses $0.5$.

Differentiability and training

Training a multilayer network by gradient descent means asking, for every weight: if I nudge this weight a little, how much does the error change? Backpropagation computes those derivatives using the chain rule, working backward from the loss through each layer. The activation derivatives enter into that calculation alongside the weights.

A hard threshold presents a problem. Away from the threshold, its output doesn't change under a small adjustment: the derivative is zero. At the threshold, the function jumps and has no derivative. Applying the ordinary chain rule through it therefore gives gradient descent no useful signal for adjusting the preceding weights.

The sigmoid is differentiable everywhere, which made it a practical replacement for the step function when training multilayer networks by backpropagation. The trade is explicit: you give up the crisp on/off decision and accept a graded "how strongly on" — precisely so that calculus works. That's the bargain in the title, and by the standards of 1986 it was a superb deal. It made gradient-based training possible through these units.

Saturation and depth

The sigmoid's derivative is

$$\sigma'(z) = \sigma(z)(1 - \sigma(z))$$

It reaches its maximum, $0.25$, at $z = 0$. In the tails, where the output approaches $0$ or $1$, the derivative approaches zero. We call a sigmoid in that regime saturated: a change in its input produces very little change in its output. Its contribution to the backpropagated gradient is correspondingly small.

In a deep stack, those small factors accumulate. Considering the activation derivatives alone, even twenty factors at their maximum give

$$\left(\frac{1}{4}\right)^{20} \approx 10^{-12}$$

Saturation makes the factors smaller still. This calculation isolates one part of the gradient; the weight matrices also matter, and can amplify as well as attenuate it. When the combined effect is a very small gradient, the early layers receive too little signal to learn effectively. When the products grow excessively, we encounter the related problem of exploding gradients.

The sigmoid's smoothness lets us compute a gradient, but it doesn't ensure that the gradient remains useful across many layers. Its bounded output comes with increasingly flat tails. That distinction between having a derivative and having a useful learning signal becomes important as we add depth.

What the 1991 analysis led to

Hochreiter's diploma thesis examined this difficulty in recurrent networks, where learning may need to connect events separated by many time steps. An error measured at step 100 can require adjusting how the network processed an input at step 1. Unroll a recurrent network over $T$ time steps and the computation resembles a feedforward network $T$ layers deep, with shared weights. Backpropagating through time involves the same kind of repeated multiplication.

The subsequent research gives a fairly direct answer to my historical question. Hochreiter and Schmidhuber's 1997 LSTM paper reviews the 1991 analysis and uses it to motivate a different architecture. Its "constant error carousel" provides a path along which the error signal can persist, with gates controlling access to the stored information. The analysis helped them identify what the architecture needed to preserve.

That doesn't explain the whole course of neural-network research in the 1990s. As the parent post discusses, researchers also had to choose among competing methods under the data and computing constraints of the time. A difficulty with training long chains of computation was one consideration among several. It wasn't a proof that neural networks had no useful applications, or that a different architecture could not address the difficulty.

My original question had compressed those separate issues into a single event: someone proves a limitation, and research stops. The connection between the 1991 analysis and LSTM shows continued work on the limitation itself. It also gives us a concrete way to examine the later methods for training deep networks: look at what each one changes in the gradient calculation.

Four approaches to gradient propagation

ReLU, careful initialization, batch normalization and residual connections operate on different parts of a network. They arose at different times and have uses beyond vanishing gradients. We can nevertheless relate them through the derivatives of successive layers.

For a fully connected sigmoid layer, the derivative of its output with respect to the previous layer's output is the Jacobian

$$J_l = \operatorname{diag}\big(\sigma'(z_l)\big) \cdot W_l$$

The diagonal matrix contains the activation derivatives, and $W_l$ contains the weights. Across $L$ layers, the chain rule composes these Jacobians in order:

$$J_L \cdot J_{L-1} \cdots J_1$$

Backpropagation uses their transposes in reverse through the network to compute the loss gradient. The size and direction of that gradient depend on the combined action of the matrices. Repeated attenuation can leave earlier layers with very small updates; repeated amplification can make the updates unstable.

ReLU changes the activation derivative. $\mathrm{ReLU}(z) = \max(0, z)$ has derivative $1$ for positive inputs and $0$ for negative inputs, with a corner at zero. An active unit therefore contributes a factor of one, rather than a sigmoid derivative of at most $0.25$. That removes this source of shrinkage along active paths. Inactive units still block their paths, and the weights can still change the gradient's scale.

Careful initialization sets the starting scale of the weights. Glorot and He initialization account for layer width and the activation's behavior when choosing that scale. The aim is to keep signals from shrinking or growing rapidly as they cross layers at the start of training. Initialization gives training a workable starting point; it cannot guarantee how the weights will evolve afterward.

Batch normalization changes the values entering the activation. In its original arrangement, it normalizes pre-activations using batch statistics, then applies a learned scale and shift. This can make saturating nonlinearities easier to train with. It doesn't guarantee that every unit stays out of saturation, and avoiding saturation is not a complete account of why the method helps. The original paper discusses its use with saturating nonlinearities as well as its broader effects on training.

Residual connections add an identity path. For a block computing $y = f(x) + x$, the Jacobian is $J_f + I$. The direct path contributes an identity term, allowing the gradient to propagate without depending entirely on the derivative of $f$. LSTM's persistent cell-state path addresses a related need in recurrent networks, although its gated architecture differs from a residual block. I'll return to that architecture in the LSTM post later in the series.

These connections help me reason about the methods beyond remembering which ones tend to work. An activation changes one derivative; initialization changes the scale of the weights; normalization changes the inputs to the nonlinearity; a skip connection changes the paths available through the computation. The gradient calculation provides a way to examine each choice and its limits.

The next sidebar takes up a different question from the parent post: how to understand the representations these networks learn. A network may learn features that help it recognise birds, but establishing what those features represent takes further investigation.