0Pricing
Node.js Backend Development Bootcamp · レッスン

アプリケーションメトリクスの公開とREDメソッド

レート、エラー、期間のメトリクスをPrometheus形式で公開し、意味のあるSLIを定義します。

「アプリケーションメトリクスの公開とREDメソッド」はCoddyKit上の無料Node.js Backend Development Bootcampレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはNode.js Backend Development Bootcamp学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Node.js Backend Development Bootcampコースには全4レッスンが含まれています。

このレッスンの一部はまだ翻訳されておらず、英語で表示されています。

Why Metrics Matter

Logs tell you what happened in a single request. Metrics tell you how your whole service is behaving right now and over time. They are cheap to store, aggregate well, and power dashboards and alerts.

A metric is a numeric measurement sampled over time. In a Node.js backend you typically export metrics in Prometheus text format over an HTTP endpoint (conventionally /metrics). Prometheus scrapes that endpoint every few seconds and stores the time series.

  • Counter — only goes up (total requests, total errors).
  • Gauge — goes up and down (active connections, queue depth).
  • Histogram — buckets observations (request durations) to compute quantiles.

The Prometheus Text Format

Before wiring up libraries, it helps to see what the /metrics endpoint actually returns. Each line is metric_name{label="value"} number. The # HELP and # TYPE comment lines describe each metric.

Labels let one metric name carry many dimensions (per route, per status code). Below is a tiny generator that builds this exposition format by hand so you can see exactly what Prometheus parses.

function renderMetrics(samples) {
  const lines = [
    '# HELP http_requests_total Total HTTP requests',
    '# TYPE http_requests_total counter',
  ];
  for (const s of samples) {
    const labels = Object.entries(s.labels)
      .map(([k, v]) => k + '="' + v + '"')
      .join(',');
    lines.push('http_requests_total{' + labels + '} ' + s.value);
  }
  return lines.join('\n');
}

const out = renderMetrics([
  { labels: { method: 'GET', route: '/users', status: '200' }, value: 1024 },
  { labels: { method: 'GET', route: '/users', status: '500' }, value: 7 },
]);
console.log(out);

Introducing the RED Method

The RED method is a focused recipe for monitoring request-driven services. For every service you track three signals:

  • R — Rate: requests per second the service is handling.
  • E — Errors: rate of failed requests (typically HTTP 5xx, sometimes 4xx).
  • D — Duration: distribution of how long requests take (latency).

RED is the request-side complement to USE (Utilization, Saturation, Errors) which targets resources like CPU and disk. For an HTTP API, RED maps almost perfectly onto user-facing experience: throughput, failure rate, and latency.

The elegant part: all three derive from instrumenting just one thing — the request lifecycle.

Installing prom-client

In Node.js the de-facto library is prom-client. It manages a registry of metrics and renders them in Prometheus format. It also ships default metrics (event loop lag, heap usage, GC) that you enable with one call.

Create a single shared registry so every part of your app registers into the same place.

const client = require('prom-client');

const register = new client.Registry();
register.setDefaultLabels({ app: 'orders-api' });

// Node.js runtime metrics: event loop lag, heap, GC, fd count...
client.collectDefaultMetrics({ register });

module.exports = { client, register };

R — A Counter for Request Rate

Rate is not stored directly. You store a monotonic counter of total requests, and Prometheus computes the per-second rate at query time with rate(http_requests_total[1m]).

Give the counter labels for method, route, and status_code. Crucially, use the route template (/users/:id) and not the raw URL (/users/42) — otherwise every id becomes a new time series and your cardinality explodes.

const client = require('prom-client');

const httpRequestsTotal = new client.Counter({
  name: 'http_requests_total',
  help: 'Total number of HTTP requests',
  labelNames: ['method', 'route', 'status_code'],
});

// On each finished request:
httpRequestsTotal.inc({
  method: 'GET',
  route: '/users/:id',   // template, NOT /users/42
  status_code: '200',
});

D — A Histogram for Duration

Duration needs a histogram so you can compute percentiles (p50, p95, p99). A histogram counts observations into predefined buckets (upper bounds in seconds). Choose buckets that span your expected latency range.

A histogram automatically exposes three series per label set: _bucket, _sum, and _count. From _sum and _count you also get the average; from _bucket you estimate quantiles with histogram_quantile().

const client = require('prom-client');

const httpRequestDuration = new client.Histogram({
  name: 'http_request_duration_seconds',
  help: 'HTTP request duration in seconds',
  labelNames: ['method', 'route', 'status_code'],
  // Tuned for a typical web API (5ms .. 5s)
  buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5],
});

// Observe a completed request that took 0.137s
httpRequestDuration.observe(
  { method: 'GET', route: '/users/:id', status_code: '200' },
  0.137
);

One Middleware to Capture R, E, and D

The whole point of RED is that a single piece of instrumentation feeds all three signals. An Express-style middleware starts a timer, lets the request run, then on res.finish records the count and duration with the status code as a label.

Errors need no separate metric — they are just the counter filtered by status_code=~"5.." at query time. This is framework code (Express), so treat it as a reference pattern rather than a standalone program.

function metricsMiddleware(req, res, next) {
  const end = httpRequestDuration.startTimer({ method: req.method });

  res.on('finish', () => {
    // req.route?.path gives the template, e.g. '/users/:id'
    const route = (req.route && req.route.path) || req.path;
    const labels = {
      method: req.method,
      route,
      status_code: String(res.statusCode),
    };
    httpRequestsTotal.inc(labels);
    end(labels); // stops timer and records duration with these labels
  });

  next();
}

app.use(metricsMiddleware);

Exposing the /metrics Endpoint

Prometheus pulls metrics, so you must expose them over HTTP. Add a GET /metrics route that returns register.metrics() with the correct content type. This is the endpoint you list as a scrape target in prometheus.yml.

Keep this endpoint internal (bind to a private interface or protect it) — it can leak operational detail and is a tempting DoS target.

const { register } = require('./metrics');

app.get('/metrics', async (req, res) => {
  res.set('Content-Type', register.contentType);
  res.end(await register.metrics());
});

// prometheus.yml scrape config:
// scrape_configs:
//   - job_name: 'orders-api'
//     scrape_interval: 15s
//     static_configs:
//       - targets: ['orders-api:3000']

Querying RED with PromQL

Once data is flowing, you express each RED signal as a PromQL query for dashboards and alerts:

  • Rate: sum(rate(http_requests_total[5m])) by (route)
  • Errors (ratio): sum(rate(http_requests_total{status_code=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
  • Duration (p95): histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))

Note the error ratio is a fraction of total requests, not a raw count. An error rate of 0.02 means 2% of requests failed — far more meaningful than "40 errors" with no denominator.

From Metrics to SLIs and SLOs

A Service Level Indicator (SLI) is a precisely defined measure of service health, usually expressed as good events / valid events. RED metrics are exactly the raw material for SLIs:

  • Availability SLI: non-5xx responses / total responses.
  • Latency SLI: requests faster than 300ms / total requests.

A Service Level Objective (SLO) is a target for an SLI over a window, e.g. "99.9% of requests succeed over 30 days." The complement of the SLO is your error budget (0.1%), which tells you how much failure you can tolerate before you must stop shipping features and fix reliability.

Computing a Latency SLI From Buckets

A latency SLI like "fraction of requests under 300ms" comes straight from histogram bucket counts: take the cumulative count in the le="0.3" bucket and divide by the total count. Here is a standalone simulation of that computation over two scrape samples.

// Cumulative bucket counts from a histogram, le = upper bound (seconds)
const sample = {
  total: 10000,
  buckets: {
    '0.1': 6000,
    '0.3': 9700,  // 9700 requests finished within 300ms
    '1.0': 9990,
    '+Inf': 10000,
  },
};

function latencySLI(s, thresholdLabel) {
  const good = s.buckets[thresholdLabel];
  return good / s.total;
}

const sli = latencySLI(sample, '0.3');
console.log('Latency SLI (<300ms): ' + (sli * 100).toFixed(2) + '%');
console.log('Meets 99% SLO? ' + (sli >= 0.99));

Quick Check

Test your understanding of cardinality and the RED method.

Recap

You can now expose meaningful application metrics in Prometheus format and reason about them with the RED method.

  • Counter, Gauge, Histogram are the core metric types; prom-client manages a registry and renders the exposition format.
  • RED = Rate, Errors, Duration — all three fall out of instrumenting the request lifecycle once, in a single middleware.
  • Rate comes from a counter via rate(); Errors are that counter filtered by status_code=~"5.."; Duration comes from a histogram via histogram_quantile().
  • Always label by route template, never raw URLs, to control cardinality.
  • Expose metrics at /metrics and let Prometheus scrape it; keep the endpoint internal.
  • Turn RED metrics into SLIs (good/valid events), set SLOs, and track your error budget.

よくある質問

「アプリケーションメトリクスの公開とREDメソッド」レッスンは無料ですか?

はい。「アプリケーションメトリクスの公開とREDメソッド」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Node.js Backend Development Bootcampコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Node.js Backend Development Bootcampコースには全4レッスンが含まれています。

「アプリケーションメトリクスの公開とREDメソッド」で何を学びますか?

レート、エラー、期間のメトリクスをPrometheus形式で公開し、意味のあるSLIを定義します。 ブラウザで直接実行するハンズオンコードでNode.js Backend Development Bootcampを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。

Node.js Backend Development Bootcampを始めるのに経験は必要ですか?

事前経験は必要ありません。CoddyKitのNode.js Backend Development Bootcampは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。

「アプリケーションメトリクスの公開とREDメソッド」レッスンにはどのくらい時間がかかりますか?

ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。

このNode.js Backend Development Bootcampレッスンでコードを書いて実行できますか?

はい。すべてのNode.js Backend Development Bootcampレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。

このコースのすべてのレッスン

  1. 相関IDを使った構造化ロギング
  2. OpenTelemetryスパンによる分散トレーシング
  3. アプリケーションメトリクスの公開とREDメソッド
  4. AsyncLocalStorageによるコンテキスト伝播
← Node.js Backend Development Bootcampに戻る