Robust Error Handling
Implement strategic error handling using try/catch, exit signals, and trapping exits to gracefully manage failures.
Robust Error Handling is a free Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Erlang's Error Philosophy
Erlang is renowned for its fault tolerance. This isn't achieved by preventing all errors, but by expecting them and designing systems that can recover gracefully. We embrace the idea of 'let it crash' where appropriate, allowing supervisors to handle failures.
Catching Internal Errors
For errors that occur within a single process, Erlang provides the try...catch construct. This is useful for handling expected, localized issues like invalid function arguments, file not found errors, or custom application-specific exceptions.
It works similarly to exception handling in other languages, but it's less common for handling failures between different processes.
`try...catch` in Action
Let's see try...catch in a simple calculation. If an error happens, we can catch it and provide a fallback or log it. Notice how we match on error:badarith for a division-by-zero.
-module(calculator).
-export([safe_divide/2]).
safe_divide(A, B) ->
try A / B of
Result -> {ok, Result}
catch
error:badarith ->
{error, division_by_zero}
end.
% To run in shell:
% calculator:safe_divide(10, 2).
% calculator:safe_divide(10, 0).Matching Different Exceptions
Erlang's try...catch allows matching on different types of exceptions:
throw: For expected conditions, often used to jump out of deep function calls.exit: When a process terminates (e.g.,exit(Reason)).error: For unexpected runtime issues (e.g., division by zero, undefined function calls).
Each type can be caught and handled differently.
-module(exception_matcher).
-export([test_catch/1]).
test_catch(Val) ->
try
case Val of
throw_it -> throw(something_thrown);
exit_it -> exit(something_exited);
error_it -> 1 / 0;
_ -> "no error"
end
catch
throw:something_thrown -> {caught, thrown};
exit:something_exited -> {caught, exited};
error:badarith -> {caught, error_badarith};
_ -> {caught, unknown}
end.
% To run in shell:
% exception_matcher:test_catch(throw_it).
% exception_matcher:test_catch(exit_it).
% exception_matcher:test_catch(error_it).Exit Signals: Erlang's Core
Beyond `try...catch` for internal errors, Erlang processes communicate their termination using exit signals. When a process dies (either gracefully or due to an error), it sends an exit signal to all processes it's linked to.
This mechanism is fundamental for building fault-tolerant systems in Erlang.
Links Propagate Exits
By default, if two processes are linked and one terminates with an exit signal (other than normal), the other linked process will also terminate with the same reason. This is Erlang's 'let it crash' philosophy in action.
This propagation allows supervisors to detect and restart entire sub-systems, ensuring failures don't leave lingering, inconsistent state.
Trapping Exits with `process_flag`
Sometimes, a process needs to handle the exit of a linked process instead of crashing itself. This is achieved by "trapping exits". A process can set its trap_exit flag to true.
When trap_exit is true, exit signals from linked processes are converted into messages that are sent to the trapping process's mailbox.
Handling a Linked Process Exit
This example shows a 'parent' process linking to a 'worker'. The parent sets trap_exit to true. When the worker crashes, the parent doesn't crash but receives an {'EXIT', Pid, Reason} message, which it can then process.
-module(exit_trap_demo).
-export([start/0, worker/0]).
start() ->
ParentPid = self(),
WorkerPid = spawn_link(fun() -> worker() end),
process_flag(trap_exit, true), % Parent traps exits
io:format("Parent (~p) linked to Worker (~p)~n", [ParentPid, WorkerPid]),
receive
{'EXIT', WorkerPid, Reason} ->
io:format("Parent caught worker exit: ~p~n", [Reason]),
{worker_died, Reason}
after 5000 ->
io:format("Parent timed out waiting for worker exit.~n"),
timeout
end.
worker() ->
io:format("Worker (~p) starting...~n", [self()]),
timer:sleep(1000), % Do some work
exit(bad_calculation). % Worker crashes
% To run in shell:
% exit_trap_demo:start().Choosing Your Strategy
When should you trap exits versus letting them crash?
- Let it Crash (default): Use when a failure in one process means the whole component is compromised. Supervisors will handle the restart logic.
- Trap Exits: Use when a process needs to clean up resources, log the event, or attempt recovery from a linked process's failure without itself dying. This is often used by supervisors themselves.
Quick Check
Consider a scenario where process_A is linked to process_B. process_B crashes with reason error_condition.
Robust Error Handling Summary
We've explored key Erlang error handling strategies:
try...catchfor localized, internal exceptions within a single process.- Exit signals as the primary mechanism for inter-process failure notification via links.
- Trapping exits using
process_flag(trap_exit, true)to convert exit signals from linked processes into messages, allowing a process to react to a linked process's termination without crashing itself.
Understanding these mechanisms is crucial for building resilient, fault-tolerant Erlang systems.
Frequently asked questions
Is the “Robust Error Handling” lesson free?
Yes — the full text of “Robust Error Handling” is free to read here on the web, and the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course, upgrade to CoddyKit PRO.
What will I learn in “Robust Error Handling”?
Implement strategic error handling using try/catch, exit signals, and trapping exits to gracefully manage failures. You practise Erlang OTP: Distributed & Fault-Tolerant Systems Programming with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start Erlang OTP: Distributed & Fault-Tolerant Systems Programming?
No prior experience is required. Erlang OTP: Distributed & Fault-Tolerant Systems Programming on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Robust Error Handling” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson?
Yes. Every Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Links and Monitors Explained
- Robust Error Handling
- Designing for Crash-First
- The Let-It-Crash Philosophy