0Pricing
Node.js Backend Development Bootcamp · 강의

서킷 브레이커 패턴과 벌크헤드 격리

서킷을 차단하고 의존 서비스별 리소스 풀을 격리하여 연쇄 장애를 막는 방법을 배웁니다.

서킷 브레이커 패턴과 벌크헤드 격리은(는) CoddyKit의 무료 Node.js Backend Development Bootcamp 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 Node.js Backend Development Bootcamp 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. Node.js Backend Development Bootcamp 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Why Cascading Failures Happen

In a microservice backend, your Node.js API often depends on downstream services: a payments provider, an inventory service, a database. When one dependency becomes slow (not even down — just slow), every request that touches it piles up.

  • Each pending request holds an event-loop slot, a socket, and memory.
  • Callers retry, multiplying load on the already-struggling dependency.
  • Your healthy endpoints starve because the process is saturated waiting on the sick one.

This is a cascading failure: one slow dependency drags down the entire service. Two patterns fight this — the circuit breaker (stop calling a failing dependency) and the bulkhead (cap how many resources any one dependency can consume).

The Naive Retry Makes It Worse

A common first instinct is to wrap a flaky call in a retry loop. But blind retries amplify a partial outage into a full one — you triple the traffic at the exact moment the dependency can least handle it.

Below, a fragile downstream is hammered. Notice how retries turn 1 logical request into many physical calls.

async function callDependency(attempt) {
  // Simulate a dependency that fails 70% of the time
  if (Math.random() < 0.7) {
    throw new Error('dependency timeout');
  }
  return 'ok';
}

async function withBlindRetry(maxRetries) {
  let physicalCalls = 0;
  for (let i = 0; i <= maxRetries; i++) {
    physicalCalls++;
    try {
      const res = await callDependency(i);
      console.log(`Success after ${physicalCalls} physical call(s)`);
      return res;
    } catch (e) {
      console.log(`Attempt ${i + 1} failed: ${e.message}`);
    }
  }
  console.log(`Gave up after ${physicalCalls} physical calls`);
}

withBlindRetry(3);

The Circuit Breaker States

A circuit breaker wraps a call and tracks its health. It is a small state machine with three states:

  • CLOSED — calls flow through normally. Failures are counted.
  • OPEN — too many recent failures. Calls are rejected immediately without touching the dependency (fail fast).
  • HALF_OPEN — after a cooldown, a few trial calls are allowed. If they succeed, go back to CLOSED; if they fail, snap back to OPEN.

The key insight: when OPEN, you stop wasting resources on a dependency that is already failing. This gives it room to recover and keeps your event loop free.

A Minimal Circuit Breaker

Here is a self-contained breaker. It opens after a failure threshold, rejects fast while OPEN, then probes once the reset timeout passes. This is the core logic every production library implements.

class CircuitBreaker {
  constructor(fn, { threshold = 3, resetMs = 5000 } = {}) {
    this.fn = fn;
    this.threshold = threshold;
    this.resetMs = resetMs;
    this.failures = 0;
    this.state = 'CLOSED';
    this.nextTry = 0;
  }

  async exec(...args) {
    if (this.state === 'OPEN') {
      if (Date.now() < this.nextTry) {
        throw new Error('Circuit OPEN - failing fast');
      }
      this.state = 'HALF_OPEN';
    }
    try {
      const result = await this.fn(...args);
      this.failures = 0;
      this.state = 'CLOSED';
      return result;
    } catch (err) {
      this.failures++;
      if (this.failures >= this.threshold) {
        this.state = 'OPEN';
        this.nextTry = Date.now() + this.resetMs;
      }
      throw err;
    }
  }
}

let n = 0;
const flaky = async () => { n++; if (n <= 5) throw new Error('boom'); return 'ok'; };
const breaker = new CircuitBreaker(flaky, { threshold: 3, resetMs: 1000 });

(async () => {
  for (let i = 0; i < 4; i++) {
    try { console.log(await breaker.exec()); }
    catch (e) { console.log(`call ${i}: ${e.message} [state=${breaker.state}]`); }
  }
})();

Fail Fast Beats Slow Failure

The real win of OPEN state is latency. A timeout might be 10 seconds; a tripped breaker rejects in microseconds. Under load this is the difference between survival and collapse.

Compare the cost of 100 requests hitting a 2-second-timeout dependency versus a breaker that rejects instantly once tripped.

function estimate(reqs, timeoutMs, openAfter) {
  let totalMs = 0;
  for (let i = 0; i < reqs; i++) {
    if (i < openAfter) {
      totalMs += timeoutMs; // these waited for the full timeout
    } else {
      totalMs += 0.01; // breaker rejected instantly
    }
  }
  return totalMs;
}

const reqs = 100, timeout = 2000, openAfter = 5;
const withBreaker = estimate(reqs, timeout, openAfter);
const noBreaker = reqs * timeout;
console.log(`No breaker: ${noBreaker} ms of blocked time`);
console.log(`With breaker: ${withBreaker.toFixed(2)} ms of blocked time`);
console.log(`Saved: ${(noBreaker - withBreaker).toFixed(0)} ms`);

Using Opossum in Production

Don't ship a hand-rolled breaker. The de-facto Node.js library is opossum. It adds rolling-window error rates, a fallback, per-call timeouts, and rich events/metrics.

  • timeout — abort a call that hangs too long.
  • errorThresholdPercentage — trip when this % of calls in the window fail.
  • resetTimeout — how long to stay OPEN before probing.

This snippet needs the opossum package and a network call, so it is illustrative, not runnable here.

const CircuitBreaker = require('opossum');

async function getInventory(sku) {
  const res = await fetch(`http://inventory.internal/items/${sku}`);
  if (!res.ok) throw new Error(`HTTP ${res.status}`);
  return res.json();
}

const breaker = new CircuitBreaker(getInventory, {
  timeout: 3000,                 // fail a call after 3s
  errorThresholdPercentage: 50,  // open at 50% failures in the window
  resetTimeout: 10000,           // try again after 10s
  rollingCountTimeout: 10000,    // 10s stats window
});

// Serve stale/cached data instead of erroring out
breaker.fallback((sku) => ({ sku, stock: 'unknown', cached: true }));

breaker.on('open', () => console.warn('inventory breaker OPEN'));
breaker.on('halfOpen', () => console.info('inventory breaker probing'));

module.exports = (sku) => breaker.fire(sku);

Fallbacks: Degrade, Don't Die

A tripped breaker should usually return something useful rather than a 500. This is graceful degradation:

  • Serve a cached or stale value.
  • Return a safe default (empty recommendations, last-known price).
  • Queue the work for later (write to a buffer, process when the dependency recovers).

Pick the fallback per dependency based on business rules. A pricing service should fail closed (block the sale); a recommendations service should fail open (show nothing) so the page still loads.

The Bulkhead Pattern

A breaker stops calls after things go wrong. A bulkhead prevents one dependency from ever monopolizing resources in the first place. The name comes from ships: watertight compartments so one flooded section doesn't sink the whole vessel.

In a backend, you give each dependency its own bounded pool:

  • A capped number of concurrent in-flight calls.
  • A separate connection pool per database or HTTP target.
  • Optionally a bounded queue; overflow is rejected fast.

So if the payments API hangs, at most N slots are stuck there — your inventory and auth calls keep their own slots and stay healthy.

Implementing a Concurrency Bulkhead

The simplest bulkhead is a concurrency limiter (a semaphore). It caps how many calls to a given dependency run at once and rejects (or queues) the rest, so a slow dependency can only ever tie up maxConcurrent slots.

class Bulkhead {
  constructor(maxConcurrent, maxQueue = 0) {
    this.max = maxConcurrent;
    this.maxQueue = maxQueue;
    this.active = 0;
    this.queue = [];
  }

  async run(task) {
    if (this.active >= this.max) {
      if (this.queue.length >= this.maxQueue) {
        throw new Error('Bulkhead full - rejected');
      }
      await new Promise((resolve) => this.queue.push(resolve));
    }
    this.active++;
    try {
      return await task();
    } finally {
      this.active--;
      const next = this.queue.shift();
      if (next) next();
    }
  }
}

const bh = new Bulkhead(2, 2);
const slow = (id) => () => new Promise((r) => setTimeout(() => { console.log('done', id); r(id); }, 50));

(async () => {
  const results = await Promise.allSettled(
    [1, 2, 3, 4, 5, 6].map((id) => bh.run(slow(id)))
  );
  results.forEach((r, i) =>
    console.log(`task ${i + 1}: ${r.status}${r.reason ? ' - ' + r.reason.message : ''}`)
  );
})();

Per-Dependency Connection Pools

The most common real-world bulkhead is the HTTP connection pool. Node's default global agent shares sockets across all targets. Instead, give each downstream its own Agent with a capped maxSockets. A hung dependency can only exhaust its own pool.

This uses the http module's Agent and a live socket, so treat it as a configuration pattern rather than a judge-runnable program.

const http = require('http');

// One isolated pool per downstream service
const paymentsAgent = new http.Agent({
  keepAlive: true,
  maxSockets: 10,        // at most 10 concurrent connections to payments
  maxFreeSockets: 5,
});

const inventoryAgent = new http.Agent({
  keepAlive: true,
  maxSockets: 20,        // inventory gets its own, independent budget
});

function callPayments(path) {
  return new Promise((resolve, reject) => {
    const req = http.request(
      { host: 'payments.internal', path, agent: paymentsAgent, timeout: 3000 },
      (res) => { res.resume(); res.on('end', resolve); }
    );
    req.on('timeout', () => req.destroy(new Error('payments timeout')));
    req.on('error', reject);
    req.end();
  });
}

module.exports = { callPayments, paymentsAgent, inventoryAgent };

Combining Breaker + Bulkhead

Production resilience layers both patterns per dependency:

  • Bulkhead bounds concurrency so a slow dependency can't saturate the process.
  • Timeout ensures no single call hangs forever.
  • Circuit breaker stops calling once the failure rate is high.
  • Fallback returns a degraded-but-useful response.

Order matters: wrap the raw call with a timeout, run it through the bulkhead, and put the breaker on the outside so it short-circuits before you even acquire a bulkhead slot. With opossum, the breaker's timeout plus its capacity/volume options can cover most of this, but explicit per-dependency pools give you the strongest isolation.

// Compose: breaker(bulkhead(timeout(call)))
function withTimeout(fn, ms) {
  return (...args) => Promise.race([
    fn(...args),
    new Promise((_, rej) => setTimeout(() => rej(new Error('timeout')), ms)),
  ]);
}

function resilient(rawCall, { bulkhead, breaker, timeoutMs }) {
  const timed = withTimeout(rawCall, timeoutMs);
  // breaker on the OUTSIDE: it can fail fast before we ever take a bulkhead slot
  return (...args) => breaker.exec(() => bulkhead.run(() => timed(...args)));
}

async function demo() {
  await withTimeout(() => Promise.resolve('fast'), 50)().then(console.log);
  try { await withTimeout(() => new Promise(() => {}), 30)(); }
  catch (e) { console.log('slow call ->', e.message); }
  console.log('Compose order: breaker -> bulkhead -> timeout -> rawCall');
}
demo();

Quick Check

Your payments dependency starts responding slowly (3-8s) but not erroring. Each request to it holds an event-loop slot and a socket. Which combination best prevents this one slow dependency from taking down your whole Node.js service?

Recap

To stop cascading failures in a Node.js backend, isolate and guard each dependency:

  • Cascading failure starts with a slow dependency, not just a dead one — pending calls exhaust sockets and the event loop.
  • Circuit breaker (CLOSED / OPEN / HALF_OPEN) fails fast once the recent error rate is high, giving the dependency room to recover. Use opossum in production.
  • Fallbacks turn an outage into graceful degradation — cached values, safe defaults, or queued work.
  • Bulkhead caps concurrency and gives each dependency its own connection pool, so one sick service can't starve the rest.
  • Compose them per dependency: breaker around bulkhead around a timed call, with a fallback for the OPEN/rejected path.

Blind retries are the trap: they amplify load exactly when the system is weakest.

자주 묻는 질문

“서킷 브레이커 패턴과 벌크헤드 격리” 강의는 무료인가요?

네 — “서킷 브레이커 패턴과 벌크헤드 격리” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 Node.js Backend Development Bootcamp 강의 전체를 잠금 해제할 수 있습니다. Node.js Backend Development Bootcamp 강의에는 총 4개의 강의가 포함되어 있습니다.

“서킷 브레이커 패턴과 벌크헤드 격리”에서 뭘 배우나요?

서킷을 차단하고 의존 서비스별 리소스 풀을 격리하여 연쇄 장애를 막는 방법을 배웁니다. 브라우저에서 직접 실행하는 실습 코드로 Node.js Backend Development Bootcamp을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

Node.js Backend Development Bootcamp을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 Node.js Backend Development Bootcamp은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“서킷 브레이커 패턴과 벌크헤드 격리” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 Node.js Backend Development Bootcamp 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 Node.js Backend Development Bootcamp 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 타임아웃, 재시도와 지터를 적용한 지수 백오프
  2. 서킷 브레이커 패턴과 벌크헤드 격리
  3. 우아한 종료와 진행 중인 요청 처리
  4. 상태 확인, 준비 상태 프로브와 부하 차단
← Node.js Backend Development Bootcamp(으)로 돌아가기