熔断器模式与舱壁隔离
通过触发熔断并按依赖隔离资源池,阻止故障级联。
熔断器模式与舱壁隔离 是 CoddyKit 上的免费 Node.js Backend Development Bootcamp 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Node.js Backend Development Bootcamp 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Node.js Backend Development Bootcamp 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Cascading Failures Happen
In a microservice backend, your Node.js API often depends on downstream services: a payments provider, an inventory service, a database. When one dependency becomes slow (not even down — just slow), every request that touches it piles up.
- Each pending request holds an event-loop slot, a socket, and memory.
- Callers retry, multiplying load on the already-struggling dependency.
- Your healthy endpoints starve because the process is saturated waiting on the sick one.
This is a cascading failure: one slow dependency drags down the entire service. Two patterns fight this — the circuit breaker (stop calling a failing dependency) and the bulkhead (cap how many resources any one dependency can consume).
The Naive Retry Makes It Worse
A common first instinct is to wrap a flaky call in a retry loop. But blind retries amplify a partial outage into a full one — you triple the traffic at the exact moment the dependency can least handle it.
Below, a fragile downstream is hammered. Notice how retries turn 1 logical request into many physical calls.
async function callDependency(attempt) {
// Simulate a dependency that fails 70% of the time
if (Math.random() < 0.7) {
throw new Error('dependency timeout');
}
return 'ok';
}
async function withBlindRetry(maxRetries) {
let physicalCalls = 0;
for (let i = 0; i <= maxRetries; i++) {
physicalCalls++;
try {
const res = await callDependency(i);
console.log(`Success after ${physicalCalls} physical call(s)`);
return res;
} catch (e) {
console.log(`Attempt ${i + 1} failed: ${e.message}`);
}
}
console.log(`Gave up after ${physicalCalls} physical calls`);
}
withBlindRetry(3);The Circuit Breaker States
A circuit breaker wraps a call and tracks its health. It is a small state machine with three states:
- CLOSED — calls flow through normally. Failures are counted.
- OPEN — too many recent failures. Calls are rejected immediately without touching the dependency (fail fast).
- HALF_OPEN — after a cooldown, a few trial calls are allowed. If they succeed, go back to CLOSED; if they fail, snap back to OPEN.
The key insight: when OPEN, you stop wasting resources on a dependency that is already failing. This gives it room to recover and keeps your event loop free.
A Minimal Circuit Breaker
Here is a self-contained breaker. It opens after a failure threshold, rejects fast while OPEN, then probes once the reset timeout passes. This is the core logic every production library implements.
class CircuitBreaker {
constructor(fn, { threshold = 3, resetMs = 5000 } = {}) {
this.fn = fn;
this.threshold = threshold;
this.resetMs = resetMs;
this.failures = 0;
this.state = 'CLOSED';
this.nextTry = 0;
}
async exec(...args) {
if (this.state === 'OPEN') {
if (Date.now() < this.nextTry) {
throw new Error('Circuit OPEN - failing fast');
}
this.state = 'HALF_OPEN';
}
try {
const result = await this.fn(...args);
this.failures = 0;
this.state = 'CLOSED';
return result;
} catch (err) {
this.failures++;
if (this.failures >= this.threshold) {
this.state = 'OPEN';
this.nextTry = Date.now() + this.resetMs;
}
throw err;
}
}
}
let n = 0;
const flaky = async () => { n++; if (n <= 5) throw new Error('boom'); return 'ok'; };
const breaker = new CircuitBreaker(flaky, { threshold: 3, resetMs: 1000 });
(async () => {
for (let i = 0; i < 4; i++) {
try { console.log(await breaker.exec()); }
catch (e) { console.log(`call ${i}: ${e.message} [state=${breaker.state}]`); }
}
})();Fail Fast Beats Slow Failure
The real win of OPEN state is latency. A timeout might be 10 seconds; a tripped breaker rejects in microseconds. Under load this is the difference between survival and collapse.
Compare the cost of 100 requests hitting a 2-second-timeout dependency versus a breaker that rejects instantly once tripped.
function estimate(reqs, timeoutMs, openAfter) {
let totalMs = 0;
for (let i = 0; i < reqs; i++) {
if (i < openAfter) {
totalMs += timeoutMs; // these waited for the full timeout
} else {
totalMs += 0.01; // breaker rejected instantly
}
}
return totalMs;
}
const reqs = 100, timeout = 2000, openAfter = 5;
const withBreaker = estimate(reqs, timeout, openAfter);
const noBreaker = reqs * timeout;
console.log(`No breaker: ${noBreaker} ms of blocked time`);
console.log(`With breaker: ${withBreaker.toFixed(2)} ms of blocked time`);
console.log(`Saved: ${(noBreaker - withBreaker).toFixed(0)} ms`);Using Opossum in Production
Don't ship a hand-rolled breaker. The de-facto Node.js library is opossum. It adds rolling-window error rates, a fallback, per-call timeouts, and rich events/metrics.
timeout— abort a call that hangs too long.errorThresholdPercentage— trip when this % of calls in the window fail.resetTimeout— how long to stay OPEN before probing.
This snippet needs the opossum package and a network call, so it is illustrative, not runnable here.
const CircuitBreaker = require('opossum');
async function getInventory(sku) {
const res = await fetch(`http://inventory.internal/items/${sku}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
return res.json();
}
const breaker = new CircuitBreaker(getInventory, {
timeout: 3000, // fail a call after 3s
errorThresholdPercentage: 50, // open at 50% failures in the window
resetTimeout: 10000, // try again after 10s
rollingCountTimeout: 10000, // 10s stats window
});
// Serve stale/cached data instead of erroring out
breaker.fallback((sku) => ({ sku, stock: 'unknown', cached: true }));
breaker.on('open', () => console.warn('inventory breaker OPEN'));
breaker.on('halfOpen', () => console.info('inventory breaker probing'));
module.exports = (sku) => breaker.fire(sku);Fallbacks: Degrade, Don't Die
A tripped breaker should usually return something useful rather than a 500. This is graceful degradation:
- Serve a cached or stale value.
- Return a safe default (empty recommendations, last-known price).
- Queue the work for later (write to a buffer, process when the dependency recovers).
Pick the fallback per dependency based on business rules. A pricing service should fail closed (block the sale); a recommendations service should fail open (show nothing) so the page still loads.
The Bulkhead Pattern
A breaker stops calls after things go wrong. A bulkhead prevents one dependency from ever monopolizing resources in the first place. The name comes from ships: watertight compartments so one flooded section doesn't sink the whole vessel.
In a backend, you give each dependency its own bounded pool:
- A capped number of concurrent in-flight calls.
- A separate connection pool per database or HTTP target.
- Optionally a bounded queue; overflow is rejected fast.
So if the payments API hangs, at most N slots are stuck there — your inventory and auth calls keep their own slots and stay healthy.
Implementing a Concurrency Bulkhead
The simplest bulkhead is a concurrency limiter (a semaphore). It caps how many calls to a given dependency run at once and rejects (or queues) the rest, so a slow dependency can only ever tie up maxConcurrent slots.
class Bulkhead {
constructor(maxConcurrent, maxQueue = 0) {
this.max = maxConcurrent;
this.maxQueue = maxQueue;
this.active = 0;
this.queue = [];
}
async run(task) {
if (this.active >= this.max) {
if (this.queue.length >= this.maxQueue) {
throw new Error('Bulkhead full - rejected');
}
await new Promise((resolve) => this.queue.push(resolve));
}
this.active++;
try {
return await task();
} finally {
this.active--;
const next = this.queue.shift();
if (next) next();
}
}
}
const bh = new Bulkhead(2, 2);
const slow = (id) => () => new Promise((r) => setTimeout(() => { console.log('done', id); r(id); }, 50));
(async () => {
const results = await Promise.allSettled(
[1, 2, 3, 4, 5, 6].map((id) => bh.run(slow(id)))
);
results.forEach((r, i) =>
console.log(`task ${i + 1}: ${r.status}${r.reason ? ' - ' + r.reason.message : ''}`)
);
})();Per-Dependency Connection Pools
The most common real-world bulkhead is the HTTP connection pool. Node's default global agent shares sockets across all targets. Instead, give each downstream its own Agent with a capped maxSockets. A hung dependency can only exhaust its own pool.
This uses the http module's Agent and a live socket, so treat it as a configuration pattern rather than a judge-runnable program.
const http = require('http');
// One isolated pool per downstream service
const paymentsAgent = new http.Agent({
keepAlive: true,
maxSockets: 10, // at most 10 concurrent connections to payments
maxFreeSockets: 5,
});
const inventoryAgent = new http.Agent({
keepAlive: true,
maxSockets: 20, // inventory gets its own, independent budget
});
function callPayments(path) {
return new Promise((resolve, reject) => {
const req = http.request(
{ host: 'payments.internal', path, agent: paymentsAgent, timeout: 3000 },
(res) => { res.resume(); res.on('end', resolve); }
);
req.on('timeout', () => req.destroy(new Error('payments timeout')));
req.on('error', reject);
req.end();
});
}
module.exports = { callPayments, paymentsAgent, inventoryAgent };Combining Breaker + Bulkhead
Production resilience layers both patterns per dependency:
- Bulkhead bounds concurrency so a slow dependency can't saturate the process.
- Timeout ensures no single call hangs forever.
- Circuit breaker stops calling once the failure rate is high.
- Fallback returns a degraded-but-useful response.
Order matters: wrap the raw call with a timeout, run it through the bulkhead, and put the breaker on the outside so it short-circuits before you even acquire a bulkhead slot. With opossum, the breaker's timeout plus its capacity/volume options can cover most of this, but explicit per-dependency pools give you the strongest isolation.
// Compose: breaker(bulkhead(timeout(call)))
function withTimeout(fn, ms) {
return (...args) => Promise.race([
fn(...args),
new Promise((_, rej) => setTimeout(() => rej(new Error('timeout')), ms)),
]);
}
function resilient(rawCall, { bulkhead, breaker, timeoutMs }) {
const timed = withTimeout(rawCall, timeoutMs);
// breaker on the OUTSIDE: it can fail fast before we ever take a bulkhead slot
return (...args) => breaker.exec(() => bulkhead.run(() => timed(...args)));
}
async function demo() {
await withTimeout(() => Promise.resolve('fast'), 50)().then(console.log);
try { await withTimeout(() => new Promise(() => {}), 30)(); }
catch (e) { console.log('slow call ->', e.message); }
console.log('Compose order: breaker -> bulkhead -> timeout -> rawCall');
}
demo();Quick Check
Your payments dependency starts responding slowly (3-8s) but not erroring. Each request to it holds an event-loop slot and a socket. Which combination best prevents this one slow dependency from taking down your whole Node.js service?
Recap
To stop cascading failures in a Node.js backend, isolate and guard each dependency:
- Cascading failure starts with a slow dependency, not just a dead one — pending calls exhaust sockets and the event loop.
- Circuit breaker (CLOSED / OPEN / HALF_OPEN) fails fast once the recent error rate is high, giving the dependency room to recover. Use
opossumin production. - Fallbacks turn an outage into graceful degradation — cached values, safe defaults, or queued work.
- Bulkhead caps concurrency and gives each dependency its own connection pool, so one sick service can't starve the rest.
- Compose them per dependency: breaker around bulkhead around a timed call, with a fallback for the OPEN/rejected path.
Blind retries are the trap: they amplify load exactly when the system is weakest.
常见问题解答
「熔断器模式与舱壁隔离」课时是免费的吗?
是的 — 「熔断器模式与舱壁隔离」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Node.js Backend Development Bootcamp 课程的其余内容,请升级到 CoddyKit PRO。 Node.js Backend Development Bootcamp 课程共包含 4 节课。
「熔断器模式与舱壁隔离」这节课中我会学到什么?
通过触发熔断并按依赖隔离资源池,阻止故障级联。 你通过在浏览器中直接运行的动手代码来练习 Node.js Backend Development Bootcamp,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Node.js Backend Development Bootcamp 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Node.js Backend Development Bootcamp 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「熔断器模式与舱壁隔离」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Node.js Backend Development Bootcamp 课中编写并运行代码吗?
能。每节 Node.js Backend Development Bootcamp 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 超时、重试与带抖动的指数退避
- 熔断器模式与舱壁隔离
- 优雅关闭与在途请求排空
- 健康检查、就绪探针与负载丢弃