0Pricing
Erlang OTP: Distributed & Fault-Tolerant Systems Programming · Lesson

Tracing & Debugging Distributed Systems

Master advanced tracing and debugging techniques for diagnosing issues across multiple Erlang nodes in a distributed environment.

Tracing & Debugging Distributed Systems is a free Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson on CoddyKit — lesson 2 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.

Debugging Distributed Erlang

Debugging a single Erlang process is already fun, but diagnosing issues across multiple interconnected Erlang nodes can be a real challenge! Why is it so hard?

  • Concurrency: Many processes running in parallel.
  • Distribution: Processes spread across different machines.
  • Asynchrony: Messages sent, not always immediately received.

Erlang provides powerful built-in tools to help us peer into these complex systems.

Starting the Erlang Debugger

The primary tool for in-depth debugging is the Erlang Debugger. It's a graphical interface that lets you inspect processes, set breakpoints, and trace code execution.

You start the debugger from the Erlang shell. Once open, you can connect to local or remote Erlang nodes to begin your investigation.

1> debugger:start().
{ok,<0.88.0>}
2> % The debugger GUI will now appear.
3> % To connect to another node:
4> % debugger:start([{node, 'other_node@hostname'}]).

Inspecting Live Processes

Once connected, the debugger allows you to select any running process on the chosen node. You can then:

  • View its current state (process dictionary, stack trace).
  • Examine its message queue (mailbox).
  • Set breakpoints on functions it's executing.

This is crucial for understanding what a process is doing, or waiting for, at any given moment.

Tracing Local Function Calls

The debugger's tracing capabilities, often accessed via the dbg module, let you monitor function calls. You can trace specific functions and see their arguments and return values.

Let's trace a simple module. Compile it, then use dbg:tp/2 to trace its add/2 function.

-module(my_math).
-export([add/2, multiply/2]).

add(A, B) ->
    A + B.

multiply(A, B) ->
    A * B.

% To run:
% 1> c(my_math).
% 2> dbg:tracer(), dbg:tp(my_math, add, 2, []).
% 3> my_math:add(5, 3).
% {trace, <0.78.0>, call, {my_math,add, [5,3]}}
% {trace, <0.78.0>, return_from, {my_math,add,2}, 8}
% 8
% 4> dbg:stop_clear().

Tracing Across Nodes

One of Erlang's powerful features is the ability to trace across a distributed system. The dbg module can be instructed to trace events on other connected nodes, allowing you to follow the flow of execution and messages between them.

This is essential when a problem involves interaction between processes residing on different machines.

Example: Tracing Remote Calls

Imagine you have a 'worker' process on a remote node. You can instruct your local debugger to trace functions on that remote node. Here's a simple worker module.

If this module was running on worker@remotehost, you could trace its process_task/1 function from your local node using dbg:tp({'worker@remotehost', worker}, process_task, 1, []) after connecting the nodes.

-module(remote_worker).
-export([start_link/0, process_task/1]).

start_link() ->
    gen_server:start_link({local, ?MODULE}, ?MODULE, [], []).

process_task(Task) ->
    io:format("~p processing task: ~p~n", [self(), Task]),
    timer:sleep(100), % Simulate work
    {ok, Task}.

% gen_server callbacks omitted for brevity.
% This module would be running on a remote node.

Analyzing Traces with TTB

For complex scenarios, viewing trace output directly in the shell or debugger GUI can be overwhelming. The Trace Tool Builder (TTB) helps by recording traces to a file for later, detailed analysis.

With TTB, you can visually replay and filter events, providing a clearer picture of system behavior over time. It's excellent for post-mortem debugging.

1> ttb:tracer().
% Start tracing and save to a file.
2> dbg:tp(my_module, my_function, 1, []).
3> my_module:my_function(data).
% ... run your system ...
4> ttb:stop().
% Later, to analyze:
5> ttb:start().
6> ttb:p(ttb_file_name, []).

Quick Debugging with Redbug

While dbg is powerful, it can have overhead. For quick, lightweight, on-the-fly tracing in a running system (even production), Redbug is often preferred.

Redbug lets you trace calls to specific functions and see their arguments and return values with minimal impact. It's perfect for quickly verifying assumptions or pinpointing recent activity.

1> redbug:start("my_module:my_function/1").
% Trace my_module:my_function/1
2> my_module:my_function(hello).
% Redbug will print trace output to the shell.
3> redbug:stop().

Redbug for Remote Tracing

Like dbg, Redbug can also trace functions on remote nodes. This makes it invaluable for quickly checking what's happening on a specific process or module on another server without bringing down the system or attaching a heavy debugger.

Here's a module we could trace on a remote node using Redbug.

-module(sensor_data).
-export([collect/1]).

collect(SensorId) ->
    Value = erlang:phash2(SensorId, 100), % Simulate reading
    io:format("Sensor ~p collected value: ~p~n", [SensorId, Value]),
    Value.

% To trace remotely (assuming nodes are connected):
% From node 'local@host':
% redbug:start("sensor_data:collect/1", [{node, 'remote@host'}]).
% Then, on 'remote@host':
% sensor_data:collect(temperature_sensor).

Distributed Debugging Check

Which of the following statements about Erlang's distributed debugging tools are true?

Recap: Tracing & Debugging

You've now explored key tools for tracing and debugging distributed Erlang systems:

  • The Erlang Debugger (dbg) for in-depth inspection and tracing.
  • TTB for recording and visualizing complex traces.
  • Redbug for lightweight, on-the-fly tracing, even in production.

Mastering these tools is essential for building and maintaining robust, fault-tolerant distributed applications with Erlang.

Frequently asked questions

Is the “Tracing & Debugging Distributed Systems” lesson free?

Yes — the full text of “Tracing & Debugging Distributed Systems” is free to read here on the web, and the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the Erlang OTP: Distributed & Fault-Tolerant Systems Programming course, upgrade to CoddyKit PRO.

What will I learn in “Tracing & Debugging Distributed Systems”?

Master advanced tracing and debugging techniques for diagnosing issues across multiple Erlang nodes in a distributed environment. You practise Erlang OTP: Distributed & Fault-Tolerant Systems Programming with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start Erlang OTP: Distributed & Fault-Tolerant Systems Programming?

No prior experience is required. Erlang OTP: Distributed & Fault-Tolerant Systems Programming on CoddyKit is structured for beginners through advanced learners; this is — lesson 2 of 4, so you can start here or from the beginning and move at your own pace.

How long does the “Tracing & Debugging Distributed Systems” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson?

Yes. Every Erlang OTP: Distributed & Fault-Tolerant Systems Programming lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. Erlang Profiling Techniques
  2. Tracing & Debugging Distributed Systems
  3. Metrics & Monitoring Integration
  4. Memory Analysis & Garbage Collection Tuning
← Back to Erlang OTP: Distributed & Fault-Tolerant Systems Programming