0Pricing
Erlang OTP: Distributed & Fault-Tolerant Systems Programming · 课时

监管者入门

了解监管者如何自动重启失败的进程,确保 Erlang 应用具备容错能力和高可用性。

监管者入门 是 CoddyKit 上的免费 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Meet Erlang Supervisors

In Erlang, processes are designed to crash! But who handles the mess? That's where Supervisors come in.

A supervisor is a special Erlang process whose job is to start, stop, and monitor other processes, called its children.

If a child process crashes, the supervisor automatically restarts it. This makes your applications incredibly resilient and fault-tolerant!

Why Fault Tolerance Matters

Imagine a web server process handling user requests. What happens if it crashes due to an error?

  • Without a supervisor, the server stops, and users lose service.
  • With a supervisor, the crashed process is detected and restarted instantly, often without users even noticing!

This "let it crash" philosophy, combined with supervisors, is key to Erlang's legendary reliability.

How Supervisors Work

Supervisors are part of Erlang's Open Telecom Platform (OTP) framework. They follow a simple hierarchy:

  • A supervisor has a list of child processes it's responsible for.
  • Each child is defined by a child specification.
  • If a child terminates unexpectedly, the supervisor steps in to restart it according to a defined strategy.

They form "supervision trees" where supervisors can supervise other supervisors.

Defining Child Processes

Before a supervisor can manage a process, it needs to know how to start it. This is done via a child specification.

A child spec is a record (or map) containing details like:

  • id: A unique name for the child.
  • start: The module, function, and arguments to call to start the process.
  • restart: When and how to restart (e.g., permanent, temporary).
  • type: Whether it's a worker or another supervisor.

Restart Strategy: One For One

Supervisors use restart strategies to decide what to do when a child crashes. The most common is one_for_one.

With one_for_one:

  • If a child process terminates, only that specific child process is restarted.
  • Other sibling processes managed by the same supervisor are unaffected.

This strategy is ideal when children are independent and a failure in one doesn't impact the others.

Our First Supervised Worker

Let's create a simple Erlang module that will act as a worker process. It will just start, print a message, and then we'll make it crash.

-module(my_worker).
-behaviour(gen_server).

-export([start_link/0, init/1, handle_call/3, handle_cast/2, handle_info/2, terminate/2, code_change/3]).
-export([crash_me/0]).

start_link() ->
    gen_server:start_link({local, ?MODULE}, ?MODULE, [], []).

init([]) ->
    io:format("Worker started!~n", []),
    {ok, #{}}.

handle_call(crash, _From, State) ->
    io:format("Worker told to crash!~n", []),
    exit(reason_for_crash),
    {reply, ok, State};
handle_call(_Request, _From, State) ->
    {noreply, State}.

handle_cast(_Msg, State) ->
    {noreply, State}.

handle_info(_Info, State) ->
    {noreply, State}.

terminate(_Reason, _State) ->
    io:format("Worker terminating!~n", []).

code_change(_OldVsn, State, _Extra) ->
    {ok, State}.

crash_me() ->
    gen_server:call(?MODULE, crash).

Setting up Our Supervisor

Now, let's create a supervisor module that will manage our my_worker. We'll specify the one_for_one restart strategy.

-module(my_supervisor).
-behaviour(supervisor).

-export([start_link/0, init/1]).

start_link() ->
    supervisor:start_link({local, ?MODULE}, ?MODULE, []).

init([]) ->
    WorkerSpec = #{
        id => my_worker,
        start => {my_worker, start_link, []},
        restart => permanent,
        type => worker,
        shutdown => 5000,
        via => [{local, my_worker}]
    },
    Children = [WorkerSpec],
    Strategy = #{
        strategy => one_for_one,
        intensity => 10,
        period => 1
    },
    {ok, {Strategy, Children}}.

Launching the Application

To see our supervisor in action, we need to start it. We can do this directly from the Erlang shell or a main application module.

Here's how to start it and check its children:

-module(app_starter).
-export([start/0, stop/0]).

start() ->
    my_supervisor:start_link(),
    io:format("Supervisor started. Worker should be running.~n", []).

stop() ->
    supervisor:stop(my_supervisor),
    io:format("Supervisor stopped.~n", []).

% To run this in the shell:
% 1. Compile: c(my_worker), c(my_supervisor), c(app_starter).
% 2. Start: app_starter:start().
% 3. Crash: my_worker:crash_me().
% 4. Observe restarts!

Witnessing Fault Tolerance

After compiling and running app_starter:start()., you should see "Worker started!". Now, call my_worker:crash_me(). in the shell.

What happens?

  • The worker process will terminate ("Worker terminating!").
  • The supervisor detects the crash and restarts the worker.
  • You'll see "Worker started!" again, demonstrating automatic recovery!

This shows the power of supervisors in keeping your system running even when individual components fail.

Supervisor Check-up

Which of the following statements correctly describe the purpose or behavior of an Erlang supervisor with a one_for_one restart strategy?

Supervisors: Your Reliability Hero

Great job! You've learned the fundamentals of Erlang supervisors:

  • They are special processes that monitor and restart child processes.
  • They ensure fault tolerance and high availability by automatically recovering from crashes.
  • Child specifications define how supervisors manage their children.
  • The one_for_one strategy restarts only the failed child.

Supervisors are a cornerstone of robust Erlang/OTP applications. Next, we'll explore more advanced restart strategies!

常见问题解答

「监管者入门」课时是免费的吗?

是的 — 「监管者入门」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程的其余内容,请升级到 CoddyKit PRO。 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程共包含 4 节课。

「监管者入门」这节课中我会学到什么?

了解监管者如何自动重启失败的进程,确保 Erlang 应用具备容错能力和高可用性。 你通过在浏览器中直接运行的动手代码来练习 Erlang OTP: Distributed & Fault-Tolerant Systems Programming,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「监管者入门」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课中编写并运行代码吗?

能。每节 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 理解 OTP 与行为
  2. 实现 GenServer 行为
  3. 监管者入门
  4. 构建 OTP 应用与发布包
← 返回 Erlang OTP: Distributed & Fault-Tolerant Systems Programming