← Articles

Designing Supervision Trees for Recovery

How process dependencies, restart policies and state reconstruction determine whether an Erlang service can recover from a crash.

A worker crashes. Its supervisor starts a replacement, and the process count returns to normal. A second worker, which survived the crash, still holds a reference to a resource that belonged to the first. The next request fails.

The supervisor has carried out its policy. Whether that policy restores a usable service depends on choices we made elsewhere: which processes depend on one another, what their state means, and what startup establishes.

Those are the choices I want to make visible when designing a supervision tree. The tree gives OTP a recovery procedure it can execute even when the code doing the work has failed. To make that procedure useful, we need to connect it to the application's own definition of a working service.

Francesco Cesarini and Steve Vinoski develop this connection particularly well in Designing for Scalability with Erlang/OTP, my favourite Erlang book. Their treatment of supervisors keeps returning to dependencies, startup order and reconstructing state from known-good sources. A small example shows why those concerns belong together.

A dependency that survives a crash

Imagine a local search service with two long-lived processes. index_owner owns an unnamed ETS table containing an index derived from source documents. query_worker obtains that table's identifier during startup and retains it for lookups. It relies on the owner having built the table before its own initialization completes.

For this example, there is no ETS heir and no ownership transfer. If index_owner terminates, its table is destroyed. A replacement owner can create another table, but the identifier saved in the surviving query worker still refers to the old one. Trying to read through that identifier raises badarg. These lifetime rules are part of the ETS contract.

Putting both processes under a one_for_one supervisor would restart the owner alone. The query worker could survive until its next lookup, then crash in turn. That may eventually cause it to reacquire the table, but it leaves a failed request to discover a dependency the design already knew about.

We have at least two reasonable ways to handle this. The query worker could explicitly detect replacement of the index and reacquire access, with defined behaviour while the index is unavailable. Or we could give the pair a restart policy that rebuilds the dependent process when its resource disappears.

The second choice fits rest_for_one: start the owner first and the query worker second. When the owner needs restarting, terminate the query worker and restart the pair in that order. When only the query worker fails, leave the owner running.

That choice depends on the example's contract. Processes that communicate do not automatically need to restart together. A client designed to reconnect may remain useful while its server recovers. We should group lifetimes where the application requires it.

Expressing the recovery order

Here is the supervisor module for that arrangement. The two worker modules are application-specific; each exports start_link/0. The owner must establish the table before reporting successful initialization, and the query worker must acquire the current table during its own initialization.

-module(search_sup).
-behaviour(supervisor).

-export([start_link/0, init/1]).

start_link() ->
    supervisor:start_link(?MODULE, []).

init([]) ->
    Flags = #{strategy => rest_for_one,
              intensity => 3,
              period => 10},
    Children = [
        #{id => index_owner,
          start => {index_owner, start_link, []},
          restart => permanent,
          shutdown => 5000,
          type => worker},
        #{id => query_worker,
          start => {query_worker, start_link, []},
          restart => permanent,
          shutdown => 5000,
          type => worker}
    ],
    {ok, {Flags, Children}}.

The list order is significant. OTP starts children in that order and shuts them down in reverse order. It invokes each start function synchronously, so the function's success has to mean what the next child needs it to mean. Starting a background index rebuild and immediately returning success would require a separate readiness protocol before the query worker could use the index.

The supervisor does not inspect the workers and discover their dependency. We encode it through the strategy and the order. For these two permanent children, assuming restarts succeed and the restart limit has not been exceeded, the strategies give us different behaviour:

StrategyOwner terminatesQuery worker terminates
one_for_oneRestart the owner.Restart the query worker.
rest_for_oneStop the query worker; restart the owner, then the query worker.Restart the query worker.
one_for_allStop the query worker; restart both in start order.Stop the owner; restart both in start order.

With more children, rest_for_one includes every child after the failed child in start order. It cannot infer which of those children is independent. If the suffix contains unrelated work, separating that work into another subtree may give it a more appropriate lifetime. The OTP supervision guide describes the strategies and their ordering.

I compiled this module and exercised the example with small ETS-backed workers on OTP 28. Restarting the owner under one_for_one left the surviving query worker with the old table identifier. Under rest_for_one, both processes were replaced and a query succeeded against the rebuilt table. Killing only the query worker left the owner and its table intact. This checks the particular dependency in the example; it does not establish a production search service's recovery guarantees.

What a restart has to rebuild

The index is useful as an example because it is derived state. If its source documents remain available and valid, rebuilding it can restore the information the service needs. That may take time, during which callers need an explicit unavailable or retry response.

Other state has different recovery requirements. A process holding the only copy of an accepted update cannot recover that update merely by running init/1 again. A ledger needs a durable record and a defined acknowledgement boundary. A worker retrying an external action needs to know what happened before it failed.

Suppose a worker completes a write, then crashes before replying. A replacement process does not by itself tell the caller whether to retry. Request identities, duplicate handling, transactions and reconciliation may be needed, depending on the operation. Cesarini and Vinoski discuss these questions in their treatment of reliability and message-delivery semantics. The relevant guarantee belongs to the whole request path, including its durable effects.

Even derived state can reproduce a fault. Reloading the same invalid input into the same buggy parser may cause another crash. The book's emphasis on known-good sources is useful here: startup must have a reason to establish a usable state. A stored snapshot is not trustworthy simply because it survived the process that produced it.

This also shapes the role of shutdown. In the example, shutdown => 5000 gives a child up to five seconds to terminate after the supervisor sends a shutdown signal; OTP kills it if necessary after the timeout. A gen_server needs to trap exits for supervisor-initiated shutdown to reach its terminate/2 callback. Crashes, forced termination and node failure make cleanup unsuitable as the only place to preserve essential state. See the child shutdown specification and gen_server termination documentation.

Recovery design therefore includes both what can be discarded and what must already be safe when a process disappears.

Deciding when to stop restarting

There are two separate restart choices in the module. rest_for_one determines which children are affected. The restart field on each child determines whether its termination calls for a restart in the first place.

Both example children are permanent, so even a normal exit calls for replacement. A transient child is restarted after an abnormal exit, excluding normal, shutdown and {shutdown, Term}. A temporary child is not restarted. A one-shot job that completes successfully often needs a different choice from a long-lived service. In particular, restarting a successfully completed permanent child still contributes to the supervisor's restart count. The supervisor reference defines these distinctions.

The intensity and period flags bound automatic restart attempts across the supervisor's children. The example allows three within ten seconds; a further attempt inside that window exceeds the limit. The supervisor then terminates its remaining children and exits with reason shutdown. A parent supervisor responds according to its own policy.

Three and ten are demonstration values, not a recommendation for a search service. A real threshold should account for acceptable bursts, the cost of reconstruction and how long we are willing to repeat an unsuccessful recovery. A busy supervisor's children share that budget; it is not a separate allowance for every worker.

The period is also not a delay between attempts. A dependency that is unavailable for several minutes may need a worker that remains alive in a disconnected state, reconnects with backoff and reports its availability honestly. Rapidly restarting that worker can exhaust a budget without changing the condition preventing it from working.

A larger restart can help when it recreates another process whose state contributed to the failure. It can also fail for exactly the same reason. Escalation expands the recovery action available to the parent; someone still has to decide what that action can repair. The restart-intensity guidance is worth reading alongside the actual tree, including the policies of its ancestors.

Leaving room for ordinary errors

None of this requires turning every unsuccessful request into a process crash. A missing document can be an ordinary result. A malformed query can be rejected at the input boundary. A temporarily unavailable dependency can have a documented error response.

The harder case is a worker that encounters a condition under which it can no longer safely continue. Catching the exception and inventing a default may keep the process alive while making its answers unreliable. Terminating the worker lets recovery happen outside the compromised execution path, using a policy and initialization procedure designed for that purpose.

The separation keeps workers focused. They handle the normal work and its expected failures; supervisors manage process lifetimes. Supervisors should remain simple enough that recovery is not itself entangled with the business logic that failed.

That still leaves diagnostic work. An exit reason and crash report help us investigate; the supervisor does not establish the cause or prove that the process boundary was wrong. A recovered service can contain a bug that will happen again. We need enough evidence to recognise that recurrence and fix it.

Checking the service after the process returns

For the search example, a useful recovery check goes beyond observing a new PID. Terminate the owner, wait for the replacement pair, and issue a query whose answer must come from the rebuilt table. Verify that the old table is gone and the query worker has obtained the replacement. Then terminate just the query worker and check that the owner was preserved.

The failure cases matter too. If index reconstruction cannot finish, what does a caller observe? If several children consume the restart budget, does the parent respond as intended? If a request was in flight, can the client distinguish a failed operation from an uncertain outcome? Those questions lead to different tests because they concern different promises.

I would draw the first version of the tree early, alongside the state ownership and dependency model, then revise it as those tests expose bad assumptions. Process isolation gives us somewhere to contain failure. The restart policy, reconstruction logic and request contract determine what we can safely do next.

When the owner in our example returns, the query worker must be able to use its table again. That observable relationship is what the supervision tree is there to restore.