AI Recap, Part 1: How Neural Networks Survived Their Second Winter
Returning to neural networks after two decades: the training difficulties, competing methods and continuing research that connect backpropagation to AlexNet.
In 1996, a friend gave me his copy of Stuart Russell and Peter Norvig's Artificial Intelligence: A Modern Approach. For 30 years it has been within arm's reach. There was a day when I could work problems from the book by heart, but by the late 2000s my attention had turned away from AI and towards distributed systems.
Since the winter of 2024/2025, I've been using LLMs daily. Working around their limitations drew me into a cycle of fixes, deeper analysis, research and revised approaches. I began reading more current papers, trying to understand the systems I was relying on.
A few weeks ago, I reached the point where I needed to brush up on my differential geometry and stopped: “How the hell did we end up here?” When I last studied AI seriously, I was writing code to perform backpropagation in neural networks. Now I was reading about the geometry of representations inside models whose development I had largely missed.
I could use these systems, but I couldn't explain enough of what had happened between those two points. This series is my attempt to close that gap: five chronological posts, eight companion pieces on questions that came up while reading, and an overview of the mathematics. I'll start with the period I had mostly filed away as “AI winter.”
Research during the winter
The story I carried around for twenty years compressed a great deal: neural networks had disappointed, little happened, and then suddenly there was ChatGPT. Looking back through the research, I can see how much that account leaves out.
Neural networks did lose standing relative to other machine-learning methods in the late 1990s and early 2000s. The historical account in Goodfellow, Bengio and Courville's Deep Learning describes that decline alongside continuing applications and research. “Winter” is useful for describing the loss of enthusiasm; it gives a poor account of the work itself.
There were practical reasons to choose another method. Adding layers to a network could make training much harder. A model might have enough capacity to express a useful solution while the training procedure struggled to reach it. More computation made a larger experiment possible, but did not by itself ensure that the experiment would learn anything useful.
Researchers also had alternatives that worked well. To understand the choices of the period, I needed to look at both the training problem and what those alternatives offered.
How the gradient becomes too small
A common activation function was the sigmoid:
$$\sigma(z) = \frac{1}{1 + e^{-z}}$$
Here $z$ is the neuron's weighted input, including its bias. The sigmoid maps that input into the interval $(0, 1)$, giving a smooth transition between low and high output. This makes it possible to calculate how a small change in an earlier weight affects the loss. Backpropagation carries out that calculation by applying the chain rule through the network.
A hard threshold has a zero derivative away from its jump and no derivative at the jump. The sigmoid has a derivative everywhere:
$$\sigma'(z) = \sigma(z)(1 - \sigma(z))$$
Its maximum is $0.25$, reached at $z = 0$. In the tails, where the output approaches $0$ or $1$, the derivative approaches zero. A unit in that regime is saturated: changing its input has very little effect on its output.
Backpropagation through successive layers involves products of activation derivatives and weight matrices. Isolating just the activation factors, twenty sigmoid derivatives contribute at most
$$\left(\frac{1}{4}\right)^{20} \approx 10^{-12}$$
That is a small multiplier even before saturation makes it smaller. It is not, however, the size of the whole gradient. The weight matrices can amplify as well as attenuate signals, and their combined action determines what reaches an earlier layer. If the resulting gradient is very small, learning there can become impractically slow. If it grows excessively, updates can become unstable.
The same issue appears in recurrent networks, where a signal may need to travel backward through many time steps. Hochreiter's 1991 analysis and the subsequent work by Bengio, Simard and Frasconi in 1994 investigated the difficulty of learning long-term dependencies. Hochreiter and Schmidhuber's 1997 LSTM paper reviews that problem and introduces an architecture with a path designed to preserve the error signal, with gates controlling access to stored information.
So the analysis was useful for designing another network. It identified a problem with how learning signals travelled through the computation, and gave researchers something specific to change. The sigmoid sidebar will work through that distinction in more detail: having a derivative and having a useful learning signal across many layers are different requirements.
What the alternatives offered
Support vector machines, as developed by Cortes and Vapnik in 1995, offered a classification objective with a different optimisation structure. The model balances a wide separation between classes against penalties for margin violations. Kernels let it work with inner products in a transformed feature space without explicitly constructing every transformed vector.
For a fixed valid kernel and chosen hyperparameters, the standard SVM training problem is convex. A local minimum is therefore global. That is a useful property, but it doesn't choose the features, kernel or regularisation strength, and it doesn't guarantee good predictions on new data. Those decisions still require evidence from the task.
Graphical models provided a way to express probabilistic relationships among variables. In a Bayesian network, for example, a directed graph helps specify how the joint distribution factors into conditional distributions. The structure can incorporate knowledge about the problem and support questions about unobserved variables. Exact inference can be expensive, so the modelling choice includes deciding which calculations are tractable and where approximation is acceptable. Kevin Murphy's introduction lays out that connection between representation and inference.
Boosting built a predictor from a sequence of simpler ones. In AdaBoost, successive learners receive reweighted training examples, placing more emphasis on examples the current combination gets wrong. Freund and Schapire's analysis connects this procedure to guarantees under stated assumptions about the learners. It offered another way to improve predictive performance without training a deep neural network.
Each method made some aspects of the problem easier to manage while leaving others to the practitioner. A convex objective could simplify optimisation; a graph could express useful structure; an ensemble could combine limited predictors. Neural networks had to justify their additional training difficulty against these alternatives on the actual task.
The division was never as simple as theory on one side and guesswork on the other. Statistical theory also applied to neural networks, and probabilistic models could be components of neural systems. The pretraining work below is one example of that overlap.
Work that continued
Three lines of research helped me connect the networks I remembered with the later developments. They are examples of continuing work, with many collaborators and predecessors behind them.
Convolutional networks used the structure of images. Local connections let a unit respond to a small region, while shared weights let the same filter be applied at different locations. This reduces the number of parameters relative to connecting every input pixel to every unit and gives the network an architectural reason to reuse a detected pattern across an image.
LeCun, Bottou, Bengio and Haffner's 1998 paper on document recognition describes convolutional networks, including LeNet-5, and systems that combine recognition with other stages of document processing. Their later review describes commercial cheque-reading systems already in use in the 1990s. A network recognising handwritten digits had a narrower task than an ImageNet classifier, but it was learning useful features and contributing to a working application.
Neural language models learned word representations along with prediction. Bengio, Ducharme, Vincent and Jauvin's 2003 paper jointly trained word vectors and a model for predicting word sequences. Similar representations let learning from one sequence help with related sequences, addressing some of the difficulty of encountering combinations absent from the training data.
The paper did not invent distributed representations. It demonstrated how learned representations and a language-model objective could work together. That gives me a route from predicting word sequences to the geometric relationships explored in later word-embedding research.
Layerwise pretraining addressed how a deep model was initialised. Hinton, Osindero and Teh's 2006 deep-belief-net paper trained successive layers using restricted Boltzmann machines, then refined the generative model with a contrastive wake-sleep procedure. It reported results on handwritten digits with three hidden layers.
In a separate 2006 paper, Hinton and Salakhutdinov used layerwise pretraining to initialise deep autoencoders, then fine-tuned them with backpropagation. An autoencoder learns to reconstruct its input through an intermediate representation; in this work, a narrow central layer provided a compressed code. Pretraining gave the subsequent optimisation a more useful starting point.
These papers offered methods for particular training problems. Their lasting value to this account is concrete: they show researchers changing architectures and initialisation procedures, measuring the results, and examining the representations they learned. The work was already underway before the large image-classification results that drew my attention back to it.
What came together in AlexNet
In the 2012 ImageNet competition, Krizhevsky, Sutskever and Hinton's entry achieved a top-5 test error of 15.3%, compared with 26.2% for the runner-up. Top-5 error counts cases where the correct class is absent from the model's five highest-ranked predictions. The winning result combined predictions from seven networks, two of which also used pretraining on a larger ImageNet release. The paper distinguishes that ensemble from the individual models.
The result brought several parts of the earlier story together: learned convolutional features, enough labelled data to train a large classifier, hardware that made the computation practical, and choices that helped optimisation and generalisation. Looking at what each contributed gives a more useful account than treating 2012 as the arrival of a single new idea.
The dataset defined a larger task. ImageNet grew through the work of Fei-Fei Li and collaborators, who organised labelled images using WordNet's category structure. The 2009 paper described 3.2 million images across 5,247 categories at that stage of the project. The collection continued to grow, while the competition used a smaller selection for a specific classification task.
The challenge used roughly 1.2 million training images across 1,000 classes. It supplied both training material and a common evaluation on which approaches could be compared. Dataset construction and labelling were part of what made the experiment possible, alongside the model itself.
GPUs made the training run feasible. The AlexNet paper reports five to six days to train a network on two NVIDIA GTX 580 GPUs. Convolution and other neural-network operations contain substantial parallel arithmetic, which can be organised to use GPU hardware effectively. Memory capacity, data movement and implementation choices still matter; the hardware's usefulness is more specific than a universal speedup over a CPU.
That time scale also affects research. A change that can be evaluated in days gives a group more opportunities to compare architectures and training choices than one requiring months. Computation affects both which model can be trained and which alternatives can be investigated.
ReLU changed the activation derivative. The rectified linear unit is
$$\mathrm{ReLU}(z) = \max(0, z)$$
For positive inputs its derivative is $1$, so an active unit doesn't introduce the sigmoid's small activation factor. For negative inputs the derivative is $0$, and at zero the function has a corner. This changes one part of gradient propagation. The weights and the rest of the architecture still affect it, and inactive units can block a learning signal.
The appeal is easy to see after the sigmoid calculation: an active ReLU avoids that particular source of repeated attenuation. It doesn't follow that every network using ReLU will train successfully, or that activation choice accounts for every improvement in deep learning.
Regularisation helped the model generalise. AlexNet used dropout in its first two fully connected layers, with a drop probability of one half. During training, randomly selected units were omitted from the computation. It also used data augmentation, presenting altered versions of training images.
The later dropout paper describes the method more generally. The probability of dropping a unit is a choice, rather than a requirement that every layer always lose half its units. Training with these omissions limits reliance on particular combinations of units. Evaluation uses the corresponding scaled prediction procedure, rather than simply retaining the random training-time omissions.
These choices address different difficulties. ReLU changes derivatives; dropout changes the training procedure to reduce overfitting; augmentation changes the examples presented to the model. More data and faster hardware support the experiment, but neither replaces those decisions. That is the relationship I want to carry into the later history: progress depends on how the parts work together.
Learning features and investigating them
The connection to my earlier AI studies becomes clearer when I look at the representation between input and prediction. In many conventional vision systems, people designed a feature extractor and trained a classifier on its output. A neural network can learn intermediate features jointly with the final prediction, allowing the task's loss to influence what gets represented along the way.
This still involves design. Choosing convolution introduces assumptions about locality and weight sharing. Choosing the labels and objective determines which distinctions training rewards. The individual filters can be learned without being specified one by one, while the architecture and task strongly shape what is possible.
Those features can sometimes be reused. DeCAF, for example, studied features extracted from an ImageNet-trained convolutional network on other visual-recognition tasks. Training a new classifier on those features let researchers investigate what transferred beyond the original labels. Transfer has limits: usefulness on one collection of photographs doesn't establish usefulness on every image domain.
Studying these representations was already part of neural-network research. The title of Rumelhart, Hinton and Williams's 1986 paper was Learning representations by back-propagating errors. The word vectors and autoencoders above continue that concern. AlexNet's success provided a larger and more effective system to examine, and further evidence that learned features could serve difficult tasks.
For me, the historical correction changes the question. I had been trying to account for a sudden break between the AI I remembered and the systems I was using. Following these results gives me a sequence of problems and responses: gradients that were hard to preserve, architectures that used known structure, representations learned alongside prediction, and experiments made practical by data and computation.
It also leaves a distinction worth pursuing. Predicting accurately is evidence that a trained system works on an evaluation; explaining which internal features it uses requires further investigation. That question will be the subject of the second sidebar and will return when the series reaches mechanistic interpretability.
A reading map for this period
These sources cover different parts of the gap I am trying to close:
- Russell and Norvig, Artificial Intelligence: A Modern Approach. The first edition at my desk gives me a reference point for the broader AI I studied, beyond neural networks alone.
- Bengio, Learning Deep Architectures for AI (2009). An account of deep architectures, representation learning and training methods before AlexNet.
- Krizhevsky, Sutskever and Hinton, ImageNet Classification with Deep Convolutional Neural Networks (2012). The architecture, training procedure and comparisons behind the competition result.
- Schmidhuber, Deep Learning in Neural Networks: An Overview (2015). A broad historical survey, with extensive references to work that a short account built around AlexNet can miss.
- LeCun, Bengio and Hinton, Deep Learning (2015). A shorter review connecting representation learning, convolutional networks and sequence models.
- Goodfellow, Bengio and Courville, Deep Learning (2016). Part I reviews the mathematical background; Parts II and III develop the methods and research questions of the period.
The next chronological post follows the years from 2012 to 2017, including convolutional networks, sequence models and early attention mechanisms. First, the companion posts take up the two questions this history has left me with: how learning signals survive through a network, and how we establish what its learned representations mean.