การจัดการการแบ่งแยกเครือข่าย
สำรวจแนวทางจัดการการแยกและการรวมเครือข่ายในคลัสเตอร์ Erlang แบบกระจายอย่างเหมาะสม เพื่อรักษาความถูกต้องสมบูรณ์ของระบบ
การจัดการการแบ่งแยกเครือข่าย เป็นบทเรียน Erlang OTP: Distributed & Fault-Tolerant Systems Programming ฟรีบน CoddyKit นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน คุณสามารถอ่านบทเรียนทั้งหมดด้านล่างฟรี — จากนั้นลองปฏิบัติด้วยตัวคุณเองในเบราว์เซอร์พร้อมตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7 บทเรียนนี้เป็นส่วนหนึ่งของเส้นทางการเรียน Erlang OTP: Distributed & Fault-Tolerant Systems Programming และความก้าวหน้าของคุณจะซิงค์ข้ามเว็บและแอป CoddyKit คอร์ส Erlang OTP: Distributed & Fault-Tolerant Systems Programming มีบทเรียนทั้งหมด 4 บทเรียน
บางส่วนของบทเรียนนี้ยังไม่ได้รับการแปล และแสดงเป็นภาษาอังกฤษ
Understanding Network Partitions
In distributed systems, a network partition happens when parts of the system can no longer communicate with each other due to network failures. Think of it like a bridge collapsing, splitting a city into disconnected districts.
This can lead to a "split-brain" scenario, where different parts of your Erlang cluster believe they are the only active ones. This often results in data inconsistency and service disruption.
Erlang Node Connectivity
Erlang nodes communicate by forming a distributed system. They connect to each other using a process called net_kernel. When a node starts, it tries to find and connect to other known nodes.
- Use
-snamefor short names (local network). - Use
-namefor full names (across networks). - All nodes must share the same magic cookie for security.
Here's a simple module. Compile it and run MyNode.get_name(). in the Erlang shell after starting with erl -sname mynode:
-module(my_node).
-export([get_name/0]).
get_name() ->
node().Monitoring Node Status
Erlang provides built-in mechanisms to detect when a node disconnects. The monitor_node/2 function allows a process to receive messages when the status of another node changes (e.g., up or down).
This is crucial for reacting to unexpected node failures or network issues. Let's see how a process can monitor another node:
-module(node_monitor).
-export([start/1]).
start(OtherNode) ->
Pid = spawn(fun() -> init(OtherNode) end),
{ok, Pid}.
init(OtherNode) ->
io:format("~p monitoring ~p~n", [self(), OtherNode]),
erlang:monitor_node(OtherNode, true),
receive
{nodeup, Node} ->
io:format("Node ~p is UP~n", [Node]);
{nodedown, Node} ->
io:format("Node ~p is DOWN!~n", [Node])
end,
io:format("Monitor process ~p exiting.~n", [self()]).Beyond Simple Disconnection
While monitor_node is powerful, it primarily tells you if a TCP connection to a node has dropped. This might not always mean a full "partition".
Short network blips or a slow network can cause temporary disconnections, leading to false positives. A true partition implies a sustained inability to communicate between groups of nodes.
- Network lag can delay detection.
- Brief outages might not warrant full system reaction.
- Application-level health checks are often needed.
Quorum and Majority Wins
To avoid "split-brain" in a network partition, distributed systems often use quorum. A quorum is the minimum number of nodes that must agree on an operation (or simply be reachable) for it to be considered valid.
The "majority wins" strategy is a common quorum approach:
- Only the partition containing more than half of the total nodes is allowed to continue operations.
- Other partitions (minority) should halt or become read-only.
This prevents conflicting updates and ensures data consistency.
Tracking Active Membership
To implement "majority wins," each node needs to know the total cluster size and which nodes are currently reachable. This creates a "membership oracle".
While a full implementation is complex, we can simulate a basic reachability check by having each node periodically "ping" its known peers. If a node can reach a majority of its peers, it considers itself "active".
Here's a conceptual module for a node to ping others:
-module(ping_checker).
-export([start/2, ping_peers/1]).
start(KnownPeers, Interval) ->
Pid = spawn(fun() -> init(KnownPeers, Interval) end),
{ok, Pid}.
init(KnownPeers, Interval) ->
ping_peers(KnownPeers),
timer:sleep(Interval),
init(KnownPeers, Interval).
ping_peers(Peers) ->
io:format("~p: Pinging peers: ~p~n", [node(), Peers]),
ActivePeers = lists:filter(fun(Peer) ->
case net_adm:ping(Peer) of
pong -> true;
pang -> false
end
end, Peers),
io:format("~p: Reachable peers: ~p~n", [node(), ActivePeers]),
TotalNodes = length(Peers) + 1, % Include self
ReachableCount = length(ActivePeers) + 1,
if
ReachableCount > TotalNodes / 2 ->
io:format("~p: I am in the MAJORITY partition!~n", [node()]);
true ->
io:format("~p: I am in the MINORITY partition or isolated.~n", [node()])
end.Fencing for Safety
When a network partition occurs and a minority partition is identified, it's crucial to prevent it from causing harm (e.g., writing conflicting data). This process is called fencing.
Fencing ensures that only the "winning" (majority) partition can continue to operate and modify shared state. Common fencing actions include:
- Shutting down services in the minority partition.
- Disabling write operations.
- Isolating resources (e.g., database access).
The goal is to prevent "split-brain" from corrupting data.
Reconciling Divergent States
After a network partition heals and nodes reconnect, their states might have diverged. This is because the active partition continued operations while the isolated ones were inactive or performing different actions.
Data reconciliation is the process of resolving these conflicts and bringing all nodes back to a consistent state. Common strategies include:
- Last Write Wins (LWW): The most recent update (based on timestamp) is chosen.
- Conflict Resolution Functions: Application-specific logic to merge data.
Designing for eventual consistency is key.
Partition Strategy Check
Consider a 5-node Erlang cluster. A network partition occurs, splitting it into two groups: Node A, B (Group 1) and Node C, D, E (Group 2). Which of the following statements about handling this partition are generally TRUE to maintain data integrity and availability?
Recap: Resilient Partitions
We've explored how to handle network partitions, a critical aspect of building resilient distributed Erlang applications. Key takeaways include:
- Detection: Beyond simple disconnections, using application-level health checks.
- Quorum: Employing strategies like "majority wins" to ensure only one active partition.
- Fencing: Preventing minority partitions from causing data inconsistencies.
- Reconciliation: Strategies for merging divergent states when partitions heal.
These principles help your Erlang systems remain available and consistent even in the face of network instability.
เรียนรู้ Erlang ด้วย AI tutor — ฟรี
เขียนและเรียกใช้โค้ดจริงในเบราว์เซอร์ของคุณ รับความช่วยเหลือทันทีจาก AI tutor 24/7 และเรียนรู้ต่อจากที่คุณหยุดบนเว็บหรือในแอป
- คอร์ส
- 12
- บทเรียน
- 48
คำถามที่พบบ่อย
บทเรียน “การจัดการการแบ่งแยกเครือข่าย” ฟรีหรือไม่
ใช่ — ข้อความเต็มของ “การจัดการการแบ่งแยกเครือข่าย” ฟรีให้อ่านที่นี่บนเว็บ เพื่อปฏิบัติแบบโต้ตอบ (ตัวแก้ไขโค้ดในตัวและติวเตอร์ AI ตลอด 24/7) และปลดล็อคส่วนที่เหลือของคอร์ส Erlang OTP: Distributed & Fault-Tolerant Systems Programming ให้อัปเกรดเป็น CoddyKit PRO คอร์ส Erlang OTP: Distributed & Fault-Tolerant Systems Programming มีบทเรียนทั้งหมด 4 บทเรียน
คุณจะเรียนรู้อะไรในบทเรียน “การจัดการการแบ่งแยกเครือข่าย”
สำรวจแนวทางจัดการการแยกและการรวมเครือข่ายในคลัสเตอร์ Erlang แบบกระจายอย่างเหมาะสม เพื่อรักษาความถูกต้องสมบูรณ์ของระบบ คุณปฏิบัติ Erlang OTP: Distributed & Fault-Tolerant Systems Programming ด้วยโค้ดที่ใช้งานได้จริงที่คุณเรียกใช้โดยตรงในเบราว์เซอร์ และติวเตอร์ AI ตลอด 24/7 ตอบคำถามของคุณขณะที่คุณไปผ่านบทเรียน
คุณต้องมีประสบการณ์ก่อนที่จะเริ่มเรียน Erlang OTP: Distributed & Fault-Tolerant Systems Programming หรือไม่
ไม่จำเป็นต้องมีประสบการณ์มาก่อน Erlang OTP: Distributed & Fault-Tolerant Systems Programming บน CoddyKit ออกแบบมาสำหรับผู้เริ่มต้นไปจนถึงผู้เรียนขั้นสูง คุณสามารถเริ่มต้นที่นี่หรือเริ่มจากตัวแรกและเรียนด้วยความเร็วของคุณเอง นี่คือบทเรียนที่ 1 จากทั้งหมด 4 บทเรียน
บทเรียน “การจัดการการแบ่งแยกเครือข่าย” ใช้เวลานานแค่ไหน
บทเรียน CoddyKit ส่วนใหญ่ใช้เวลาประมาณ 5–10 นาที แต่ละบทเรียนจึงสั้นและเป็นแบบโต้ตอบ คุณสามารถก้าวหน้าอย่างต่อเนื่องและกลับมาเรียนต่อจากตรงที่เพิ่งหยุดบนเว็บและแอปได้เลย
ฉันเขียนและรันโค้ดในบทเรียน Erlang OTP: Distributed & Fault-Tolerant Systems Programming นี้ได้ไหม
ได้ บทเรียน Erlang OTP: Distributed & Fault-Tolerant Systems Programming ทุกบทมีตัวแก้ไขโค้ดในตัว คุณจึงเขียนและรันโค้ดจริงได้เลยในเบราว์เซอร์ และได้รับข้อเสนอแนะจาก AI ในทันที — ไม่ต้องติดตั้งในเครื่องของคุณ
บทเรียนทั้งหมดในหลักสูตรนี้
- การจัดการการแบ่งแยกเครือข่าย
- ข้อมูลแบบกระจายด้วย ETS และ Mnesia
- การออกแบบเพื่อการขยายขนาดและความยืดหยุ่น
- การกระจายภาระงานและการสลับไปใช้ระบบสำรองระหว่างโหนด