高级重启策略
深入了解 one_for_one、one_for_all 和 rest_for_one 重启策略,以及它们对容错能力的影响。
高级重启策略 是 CoddyKit 上的免费 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Intro to Restart Strategies
Welcome to a crucial topic in Erlang: restart strategies! These define how a supervisor reacts when one of its child processes crashes.
Understanding these strategies is key to building fault-tolerant and self-healing systems, which is a hallmark of Erlang/OTP.
Supervisors: The Fault Managers
Before diving in, let's quickly recap: a supervisor is a special process that monitors other processes (its children).
- If a child process dies, the supervisor detects it.
- Based on its configured restart strategy, the supervisor decides how to bring the system back to a healthy state.
- This ensures your application can recover from individual process failures automatically.
`one_for_one`: Isolated Restarts
The one_for_one strategy is the simplest and most common. When a child process terminates:
- Only the failing child process is restarted.
- All other sibling processes remain unaffected and continue running.
This strategy is ideal for systems where child processes are largely independent of each other, such as individual client connections.
`one_for_one` in Action
Let's see one_for_one. We'll have a supervisor managing a single worker. When the worker crashes, only it restarts.
To run:
1. Compile: c(worker_gen). c(supervisor_one_for_one).
2. Start supervisor: supervisor_one_for_one:start_link().
3. Crash worker: worker_gen:crash(worker_1).
Observe the output in your Erlang shell.
-module(worker_gen).
-behaviour(gen_server).
-export([start_link/1, init/1, handle_call/3, terminate/2, crash/1]).
start_link(Id) ->
gen_server:start_link({local, Id}, ?MODULE, Id, []).
init(Id) ->
io:format("Worker ~p (~p) started.~n", [Id, self()]),
{ok, Id}.
handle_call(crash, _From, State) ->
exit(i_crashed),
{reply, ok, State};
handle_call(_Req, _From, State) ->
{reply, ok, State}.
terminate(_Reason, Id) ->
io:format("Worker ~p (~p) terminated.~n", [Id, self()]).
crash(Id) ->
gen_server:call(Id, crash).
-module(supervisor_one_for_one).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
init([]) ->
ChildSpec = #{
id => worker_1,
start => {worker_gen, start_link, [worker_1]},
restart => permanent,
shutdown => 5000,
type => worker,
modules => [worker_gen]
},
{ok, #{
strategy => {one_for_one, 3, 5},
children => [ChildSpec]
}}.
`one_for_all`: All for One Failure
The one_for_all strategy takes a more drastic approach:
- If any child process terminates, all other child processes are first terminated.
- Then, all child processes (including the one that crashed) are restarted.
This is useful when your child processes are tightly coupled and require a consistent, synchronized state. A failure in one implies a need to reset the entire group.
`one_for_all` in Action
Here, a supervisor manages two workers. Crash one, and both will restart.
To run:
1. Compile: c(worker_gen). c(supervisor_one_for_all).
2. Start supervisor: supervisor_one_for_all:start_link().
3. Crash worker: worker_gen:crash(worker_A).
Notice how both worker_A and worker_B restart.
-module(worker_gen).
-behaviour(gen_server).
-export([start_link/1, init/1, handle_call/3, terminate/2, crash/1]).
start_link(Id) ->
gen_server:start_link({local, Id}, ?MODULE, Id, []).
init(Id) ->
io:format("Worker ~p (~p) started.~n", [Id, self()]),
{ok, Id}.
handle_call(crash, _From, State) ->
exit(i_crashed),
{reply, ok, State};
handle_call(_Req, _From, State) ->
{reply, ok, State}.
terminate(_Reason, Id) ->
io:format("Worker ~p (~p) terminated.~n", [Id, self()]).
crash(Id) ->
gen_server:call(Id, crash).
-module(supervisor_one_for_all).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
init([]) ->
Child1 = #{
id => worker_A,
start => {worker_gen, start_link, [worker_A]},
restart => permanent, shutdown => 5000, type => worker, modules => [worker_gen]
},
Child2 = #{
id => worker_B,
start => {worker_gen, start_link, [worker_B]},
restart => permanent, shutdown => 5000, type => worker, modules => [worker_gen]
},
{ok, #{
strategy => {one_for_all, 3, 5},
children => [Child1, Child2]
}}.
`rest_for_one`: Cascade Restarts
The rest_for_one strategy offers a middle ground:
- If a child process terminates, it and all subsequent children (those defined after it in the supervisor's child list) are terminated.
- Then, the failed child and all subsequent children are restarted.
- Children defined before the failed child are left untouched.
This is useful when processes have sequential dependencies, where a failure in an earlier stage might invalidate the state of later stages.
`rest_for_one` in Action
We have three workers: X, Y, Z. If Y crashes, Y and Z restart, but X remains active.
To run:
1. Compile: c(worker_gen). c(supervisor_rest_for_one).
2. Start supervisor: supervisor_rest_for_one:start_link().
3. Crash worker: worker_gen:crash(worker_Y).
See worker_X stay running, while worker_Y and worker_Z restart.
-module(worker_gen).
-behaviour(gen_server).
-export([start_link/1, init/1, handle_call/3, terminate/2, crash/1]).
start_link(Id) ->
gen_server:start_link({local, Id}, ?MODULE, Id, []).
init(Id) ->
io:format("Worker ~p (~p) started.~n", [Id, self()]),
{ok, Id}.
handle_call(crash, _From, State) ->
exit(i_crashed),
{reply, ok, State};
handle_call(_Req, _From, State) ->
{reply, ok, State}.
terminate(_Reason, Id) ->
io:format("Worker ~p (~p) terminated.~n", [Id, self()]).
crash(Id) ->
gen_server:call(Id, crash).
-module(supervisor_rest_for_one).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
init([]) ->
ChildA = #{
id => worker_X,
start => {worker_gen, start_link, [worker_X]},
restart => permanent, shutdown => 5000, type => worker, modules => [worker_gen]
},
ChildB = #{
id => worker_Y,
start => {worker_gen, start_link, [worker_Y]},
restart => permanent, shutdown => 5000, type => worker, modules => [worker_gen]
},
ChildC = #{
id => worker_Z,
start => {worker_gen, start_link, [worker_Z]},
restart => permanent, shutdown => 5000, type => worker, modules => [worker_gen]
},
{ok, #{
strategy => {rest_for_one, 3, 5},
children => [ChildA, ChildB, ChildC]
}}.
Selecting the Best Strategy
Choosing the right strategy depends on your application's architecture and process dependencies:
one_for_one: Use for independent processes, like individual client connections or request handlers.one_for_all: Best for tightly coupled processes that must always be in a consistent state together (e.g., a group of processes managing a single resource).rest_for_one: Suitable for sequential pipelines or layered systems where a failure in an earlier stage affects subsequent stages.
Test Your Knowledge
You are building a system where a primary worker fetches data, and two secondary workers process different aspects of that data. If the primary worker fails, the secondary workers cannot continue with stale data and must also restart. If a secondary worker fails, the others are unaffected. Which strategy is most appropriate for the primary worker and its dependent secondary workers?
Recap: Mastering Restarts
You've explored Erlang's powerful restart strategies:
one_for_one: Restarts only the failed child, keeping others running.one_for_all: Restarts all children if any one fails, ensuring full consistency.rest_for_one: Restarts the failed child and all subsequent children in the list.
These strategies are fundamental to building robust, self-healing applications in Erlang, allowing your system to recover from failures gracefully.
常见问题解答
「高级重启策略」课时是免费的吗?
是的 — 「高级重启策略」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程的其余内容,请升级到 CoddyKit PRO。 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程共包含 4 节课。
「高级重启策略」这节课中我会学到什么?
深入了解 one_for_one、one_for_all 和 rest_for_one 重启策略,以及它们对容错能力的影响。 你通过在浏览器中直接运行的动手代码来练习 Erlang OTP: Distributed & Fault-Tolerant Systems Programming,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。
「高级重启策略」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课中编写并运行代码吗?
能。每节 Erlang OTP: Distributed & Fault-Tolerant Systems Programming 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。