← Articles

AI Recap, Sidebar 2: How Neural Networks Represent the World

From authored ontologies to learned features: how we investigate neural representations, what their geometry can tell us, and where observation gives way to interpretation.

In a symbolic knowledge base, I can declare that a penguin is a bird and inspect the statement later. In a trained image classifier, I can look at the weights, but there may be no identifiable statement to read. The classifier might recognise penguins reliably while leaving me with considerable work to establish which features it uses to do so.

That difference interests me as someone returning to neural networks from the symbolic AI tradition. Knowledge Representation was a central part of the AI I grew up on: frames, semantic networks, description logics, explicit accounts of what a system knows and how it can reason with that knowledge. Studying representations was part of the work from the beginning.

It also has a long history within neural networks. Rumelhart, Hinton and Williams called their 1986 paper Learning representations by back-propagating errors. They described hidden units learning features of the task domain through training. The research predates the 2012 turning point in Part 1 of this recap by decades.

The distinction I want to explore is between specifying a representation and investigating one that training has produced. Both involve human design. They give us different kinds of evidence about what a system is doing.

From declarations to experiments

An authored ontology makes some of its commitments explicit. It might define Penguin as a subclass of Bird, give birds a relationship to habitats, and constrain the values that relationship can take. Languages such as OWL let us express those commitments formally and derive consequences from them.

Knowing the declarations doesn't make every consequence obvious. A large knowledge base can be difficult to understand, and a perfectly consistent ontology can still describe its subject badly. But the declarations provide something concrete to inspect: we can ask whether a relationship was stated correctly, whether a conclusion follows, or whether the categories serve the task.

A neural network's designers make consequential choices too. They choose an architecture, training data, an objective and an optimisation procedure. A convolutional architecture, for example, already builds in assumptions about local structure and the reuse of filters across an image. The learned features develop within those constraints. Saying that nobody designed them means that nobody specified the individual detectors in advance; it doesn't mean that training proceeds without human choices shaping the result.

Zeiler and Fergus's work on convolutional networks gives an example of how researchers investigate what training produces. They visualised intermediate features and removed parts of the network to measure their contribution to classification. Their results helped them revise the architecture. Investigation of the learned representation fed back into engineering the next model.

For a simpler example, suppose a unit responds strongly to photographs of birds. It might respond to feathers, to a particular texture, or to the branches and sky that often accompany birds in the training images. A collection of high-activation images gives us candidates. Images that separate those properties help us distinguish them. Changing the relevant activation and observing the effect on the classifier addresses a further question: how does the network use that response?

This is the kind of work I have in mind when I think of empirical ontology: investigating which distinctions a trained system makes and how it uses them. The phrase is an analogy to my old field. A feature we label “bird” need not behave like a symbolic Bird category, and finding a useful label doesn't establish that we have understood the computation. Mechanistic interpretability, which I'll return to later in the series, asks for an account of how the internal components produce the behaviour.

Geometry and shared dimensions

A word embedding assigns a vector of numbers to a word. Training can arrange these vectors so that geometric relationships capture regularities in language. In their 2013 study of word representations, Mikolov, Yih and Zweig examined offsets between word vectors. Some relationships appeared as approximately consistent offsets, including the familiar example king − man + woman ≈ queen.

This gives us a way to investigate a learned relationship without finding an explicit rule that names it. We can measure the offset, look for other word pairs with a similar relationship, and test how consistently it holds. The approximation matters: vector arithmetic succeeding on some analogies doesn't make addition a general account of how meaning composes.

An individual coordinate need not have an intelligible meaning. A useful direction can run across many coordinates, and the same coordinates can contribute to several features. Looking for one neuron per concept would then miss part of the representation's structure.

Elhage and colleagues' Toy Models of Superposition examines a particular form of this sharing. In small networks trained on synthetic data, they demonstrate how sparse features — features that are usually inactive — can be represented in greater numbers than the available dimensions. Their directions overlap, introducing interference; nonlinear filtering helps recover the features. Sparsity makes that tradeoff workable because fewer features need to be active together.

The controlled setting matters. The researchers know which features generated the inputs, so they can compare the learned representation with those features directly. Applying the account to a large trained model requires further evidence about its features and computations. The toy model demonstrates a mechanism and conditions under which it occurs; it doesn't supply a complete description of every network.

These examples change what I would look for when inspecting a system. An explicit ontology directs attention to named categories and relations. A learned representation may require measurements across many activations before a useful distinction becomes apparent. Even then, the name I give that distinction remains a hypothesis about what the system does.

What the observations let us infer

Suppose several classifiers learn features that respond to edges. There is structure in the images that makes edges useful, but there are also shared choices in how those models were built and trained. Finding similar features invites an investigation of both. It doesn't, by itself, tell us how much of the agreement comes from the subject matter and how much comes from the training setup.

This is where the question of representation meets a philosophical interest of mine. What relationship does a useful representation have to the world it represents? How much can we infer from its success?

The practical questions give us ways to make progress. We can ask whether a feature continues to predict behaviour on different data, whether a corresponding feature appears after changing the architecture, or whether an intervention has the effect our interpretation predicts. Each experiment establishes something within its conditions. Transfer across models would broaden the evidence for an interpretation; failure to transfer would give us a reason to examine its limits.

My background in Madhyamika Buddhism shapes how I approach the further question of whether these features are “real.” I am comfortable discussing conventional reality: what we observe, what we can infer from observations, and how those inferences hold up when we act on them. I don't take a successful account of a network's behaviour as grounds for a claim about ultimate reality.

That leaves substantial room for investigation. An interpretation that predicts a model's errors can help us understand where it will fail. One that supports a reliable intervention can help us change its behaviour. If it stops working when we alter the training data, that dependence becomes part of the account. I want to understand those relationships as precisely as the evidence allows.

It also keeps the investigator's contribution visible. We choose what to measure and which distinctions to name. The model may use a feature consistently without that feature corresponding neatly to a category we brought to the investigation. Calling it “bird” is useful only to the extent that the name helps us explain and predict what happens.

Mathematics, observation and inference

The mathematics is where another of my interests enters. Faced with a collection of activations, there are many possible descriptions: individual responses, distances between vectors, directions associated with a property, or relations among groups of units. Choosing a description takes imagination as well as calculation.

A proposed direction, for example, can do more than summarise the examples that suggested it. It can give us a measurement to try on new inputs and an intervention whose effects we can compare with a prediction. A mathematical framing earns our confidence as those consequences are tested. When it fails, the failure can point to structure the framing left out.

I see a connection here with painting and music. A portraitist makes choices about which relationships of colour and shape will convey a face; a composer can develop a phrase by changing its rhythm or setting it against another voice. When I think of the melodies in a Philip Glass violin concerto, part of what interests me is how the surrounding relationships change what I hear. Mathematical descriptions can change what I notice in observations in a comparable way, while carrying the additional obligation to support claims we can test.

That is the spirit in which I want to approach the mathematics at the end of this series: what each method lets us describe, what it lets us predict, and what remains unresolved. The penguin classifier is still a useful place to start. Naming a response is one step; finding out what sustains it, where it fails, and how it affects a decision is the work that follows.