Introduction to Supervisors
Discover how supervisors automatically restart failed processes, ensuring fault tolerance and high availability in your Erlang applications.
Introduction to Supervisors is a free Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson on CoddyKit — lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Meet Erlang Supervisors
In Erlang, processes are designed to crash! But who handles the mess? That's where Supervisors come in.
A supervisor is a special Erlang process whose job is to start, stop, and monitor other processes, called its children.
If a child process crashes, the supervisor automatically restarts it. This makes your applications incredibly resilient and fault-tolerant!
Why Fault Tolerance Matters
Imagine a web server process handling user requests. What happens if it crashes due to an error?
- Without a supervisor, the server stops, and users lose service.
- With a supervisor, the crashed process is detected and restarted instantly, often without users even noticing!
This "let it crash" philosophy, combined with supervisors, is key to Erlang's legendary reliability.
How Supervisors Work
Supervisors are part of Erlang's Open Telecom Platform (OTP) framework. They follow a simple hierarchy:
- A supervisor has a list of child processes it's responsible for.
- Each child is defined by a child specification.
- If a child terminates unexpectedly, the supervisor steps in to restart it according to a defined strategy.
They form "supervision trees" where supervisors can supervise other supervisors.
Defining Child Processes
Before a supervisor can manage a process, it needs to know how to start it. This is done via a child specification.
A child spec is a record (or map) containing details like:
id: A unique name for the child.start: The module, function, and arguments to call to start the process.restart: When and how to restart (e.g.,permanent,temporary).type: Whether it's aworkeror anothersupervisor.
Restart Strategy: One For One
Supervisors use restart strategies to decide what to do when a child crashes. The most common is one_for_one.
With one_for_one:
- If a child process terminates, only that specific child process is restarted.
- Other sibling processes managed by the same supervisor are unaffected.
This strategy is ideal when children are independent and a failure in one doesn't impact the others.
Our First Supervised Worker
Let's create a simple Erlang module that will act as a worker process. It will just start, print a message, and then we'll make it crash.
-module(my_worker).
-behaviour(gen_server).
-export([start_link/0, init/1, handle_call/3, handle_cast/2, handle_info/2, terminate/2, code_change/3]).
-export([crash_me/0]).
start_link() ->
gen_server:start_link({local, ?MODULE}, ?MODULE, [], []).
init([]) ->
io:format("Worker started!~n", []),
{ok, #{}}.
handle_call(crash, _From, State) ->
io:format("Worker told to crash!~n", []),
exit(reason_for_crash),
{reply, ok, State};
handle_call(_Request, _From, State) ->
{noreply, State}.
handle_cast(_Msg, State) ->
{noreply, State}.
handle_info(_Info, State) ->
{noreply, State}.
terminate(_Reason, _State) ->
io:format("Worker terminating!~n", []).
code_change(_OldVsn, State, _Extra) ->
{ok, State}.
crash_me() ->
gen_server:call(?MODULE, crash).Setting up Our Supervisor
Now, let's create a supervisor module that will manage our my_worker. We'll specify the one_for_one restart strategy.
-module(my_supervisor).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
init([]) ->
WorkerSpec = #{
id => my_worker,
start => {my_worker, start_link, []},
restart => permanent,
type => worker,
shutdown => 5000,
via => [{local, my_worker}]
},
Children = [WorkerSpec],
Strategy = #{
strategy => one_for_one,
intensity => 10,
period => 1
},
{ok, {Strategy, Children}}.Launching the Application
To see our supervisor in action, we need to start it. We can do this directly from the Erlang shell or a main application module.
Here's how to start it and check its children:
-module(app_starter).
-export([start/0, stop/0]).
start() ->
my_supervisor:start_link(),
io:format("Supervisor started. Worker should be running.~n", []).
stop() ->
supervisor:stop(my_supervisor),
io:format("Supervisor stopped.~n", []).
% To run this in the shell:
% 1. Compile: c(my_worker), c(my_supervisor), c(app_starter).
% 2. Start: app_starter:start().
% 3. Crash: my_worker:crash_me().
% 4. Observe restarts!Witnessing Fault Tolerance
After compiling and running app_starter:start()., you should see "Worker started!". Now, call my_worker:crash_me(). in the shell.
What happens?
- The worker process will terminate ("Worker terminating!").
- The supervisor detects the crash and restarts the worker.
- You'll see "Worker started!" again, demonstrating automatic recovery!
This shows the power of supervisors in keeping your system running even when individual components fail.
Supervisor Check-up
Which of the following statements correctly describe the purpose or behavior of an Erlang supervisor with a one_for_one restart strategy?
Supervisors: Your Reliability Hero
Great job! You've learned the fundamentals of Erlang supervisors:
- They are special processes that monitor and restart child processes.
- They ensure fault tolerance and high availability by automatically recovering from crashes.
- Child specifications define how supervisors manage their children.
- The
one_for_onestrategy restarts only the failed child.
Supervisors are a cornerstone of robust Erlang/OTP applications. Next, we'll explore more advanced restart strategies!
Frequently asked questions
Is the “Introduction to Supervisors” lesson free?
Yes — the full text of “Introduction to Supervisors” is free to read here on the web, and the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course, upgrade to CoddyKit PRO.
What will I learn in “Introduction to Supervisors”?
Discover how supervisors automatically restart failed processes, ensuring fault tolerance and high availability in your Erlang applications. You practise Erlang OTP: Distributed & Fault-Tolerant Systems Programming with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Erlang OTP: Distributed & Fault-Tolerant Systems Programming?
No prior experience is required. Erlang OTP: Distributed & Fault-Tolerant Systems Programming on CoddyKit is structured for beginners through advanced learners; this is — lesson 3 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Introduction to Supervisors” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson?
Yes. Every Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Understanding OTP & Behaviors
- Implementing GenServer Behavior
- Introduction to Supervisors
- Building OTP Applications & Releases