Node.js Backend Development Bootcamp · 课时

使用 OpenTelemetry 跨服务跟踪

自动或手动为服务添加插桩,生成能够揭示服务边界间延迟的跨度。

第 2 / 4 课13 个步骤

使用 OpenTelemetry 跨服务跟踪 是 CoddyKit 上的免费 Node.js Backend Development Bootcamp 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Node.js Backend Development Bootcamp 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Node.js Backend Development Bootcamp 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Distributed Tracing?

In a microservices system one user request can hop across an API gateway, an orders service, a payments service, and a database. When that request is slow, a single service's logs cannot tell you where the time went.

Distributed tracing stitches the whole journey together. Each unit of work becomes a span, spans are linked into a trace, and the trace reveals latency across every service boundary.

  • Trace: the entire end-to-end request, identified by a traceId.
  • Span: one operation (an HTTP call, a DB query) with a start time, duration, and parent.
  • Context propagation: passing traceId and spanId across service boundaries, usually via HTTP headers.

OpenTelemetry (OTel) is the vendor-neutral standard for producing these spans in Node.js.

Anatomy of a Span

A span is the atomic building block of a trace. Every span carries the same trace identity but its own identity and timing.

  • traceId: 16 bytes, shared by every span in the trace.
  • spanId: 8 bytes, unique to this span.
  • parentSpanId: links this span to the operation that caused it.
  • name, startTime, endTime (duration = end - start).
  • Attributes: key/value tags like http.method or db.system.
  • Status: OK, ERROR, or UNSET.

Parent/child links form a tree. The root span is the whole request; child spans are the calls it makes. Visualized, the tree becomes the familiar waterfall you see in Jaeger or Tempo.

Auto-Instrumentation with the Node SDK

The fastest way to get spans is auto-instrumentation. The OTel Node SDK monkey-patches popular libraries (http, Express, pg, ioredis, etc.) so they emit spans without you writing tracing code.

Create a tracing.js file that starts the SDK before anything else, then run your app with node -r ./tracing.js app.js so it loads first.

// tracing.js
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-http');
const { Resource } = require('@opentelemetry/resources');
const { SemanticResourceAttributes } = require('@opentelemetry/semantic-conventions');

const sdk = new NodeSDK({
  resource: new Resource({
    [SemanticResourceAttributes.SERVICE_NAME]: 'orders-service',
  }),
  traceExporter: new OTLPTraceExporter({
    url: 'http://localhost:4318/v1/traces',
  }),
  instrumentations: [getNodeAutoInstrumentations()],
});

sdk.start();

What Auto-Instrumentation Gives You

With the SDK loaded, an incoming HTTP request automatically becomes a root span, and any outgoing http/fetch call or pg query becomes a child span underneath it.

  • Inbound Express route → server span with http.method, http.route, http.status_code.
  • Outbound HTTP call → client span, and the headers are injected automatically.
  • Database query → client span with db.system and the statement.

This covers the boundaries for free. But auto-instrumentation does not understand your business logic — pricing rules, cache decisions, batch loops. For those you add spans manually.

Getting a Tracer

To create spans manually you first obtain a tracer from the global trace API. Name it after the module or library producing the spans; the version is optional but helps when debugging instrumentation.

The tracer is the factory for all your manual spans.

const { trace } = require('@opentelemetry/api');

// Name + version identify the instrumentation scope
const tracer = trace.getTracer('orders-service', '1.0.0');

// Later, anywhere in the code:
// const span = tracer.startSpan('chargeCustomer');

startActiveSpan: The Idiomatic Pattern

Prefer tracer.startActiveSpan() over startSpan(). startActiveSpan makes the new span the active span for the duration of its callback, so any child spans created inside (including auto-instrumented ones) automatically attach as children.

The golden rules: always span.end() in a finally block, and record errors plus an ERROR status on failure.

const { trace, SpanStatusCode } = require('@opentelemetry/api');
const tracer = trace.getTracer('orders-service');

async function processOrder(order) {
  return tracer.startActiveSpan('processOrder', async (span) => {
    try {
      span.setAttribute('order.id', order.id);
      span.setAttribute('order.items', order.items.length);
      const result = await chargeAndShip(order); // child spans nest here
      span.setStatus({ code: SpanStatusCode.OK });
      return result;
    } catch (err) {
      span.recordException(err);
      span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
      throw err;
    } finally {
      span.end();
    }
  });
}

Attributes, Events, and Status

Spans become useful when you enrich them. Three tools:

  • Attributes — searchable key/value tags. Use semantic conventions (http.method, db.system, messaging.system) so backends understand them.
  • Events — timestamped log lines anchored inside the span, e.g. span.addEvent('cache.miss').
  • Status — set ERROR only on real failures; leave success as UNSET or OK.

Keep cardinality sane: never put a raw user ID or full SQL with literals into a high-traffic attribute if your backend indexes it — it can explode storage.

function readFromCache(span, key) {
  const hit = cache.has(key);
  if (hit) {
    span.addEvent('cache.hit', { 'cache.key': key });
  } else {
    span.addEvent('cache.miss', { 'cache.key': key });
  }
  span.setAttribute('cache.hit', hit);
  return hit ? cache.get(key) : null;
}

Context Propagation Across Services

A trace only spans services if the trace context travels with the request. The W3C traceparent header carries the traceId, parent spanId, and sampling flag.

Auto-instrumentation injects and extracts this header for you on standard HTTP. When you do something non-standard (a custom transport, a message queue), you inject/extract manually with the propagation API.

const { context, propagation, trace } = require('@opentelemetry/api');

// SENDER: inject current context into outgoing carrier (e.g. message headers)
function publish(queue, payload) {
  const headers = {};
  propagation.inject(context.active(), headers);
  queue.send({ payload, headers }); // traceparent now travels with the message
}

// RECEIVER: extract context and continue the trace
function onMessage(msg) {
  const parentCtx = propagation.extract(context.active(), msg.headers);
  const tracer = trace.getTracer('worker');
  context.with(parentCtx, () => {
    tracer.startActiveSpan('handleMessage', (span) => {
      handle(msg.payload);
      span.end();
    });
  });
}

Reading the traceparent Header

The traceparent header has a fixed, parseable shape. Understanding it helps you debug broken traces (a missing child usually means a dropped header).

Format: version-traceId-parentId-flags, e.g.
00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

  • 00 — version
  • 32 hex chars — the traceId
  • 16 hex chars — the parent spanId
  • 01 — flags (bit 0 = sampled)

Here is a tiny standalone parser to make the structure concrete.

function parseTraceparent(header) {
  const parts = header.split('-');
  if (parts.length !== 4) throw new Error('invalid traceparent');
  const [version, traceId, parentId, flags] = parts;
  return {
    version,
    traceId,
    parentId,
    sampled: (parseInt(flags, 16) & 1) === 1,
  };
}

const h = '00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01';
console.log(parseTraceparent(h));
// { version: '00', traceId: '4bf9...4736', parentId: '00f0...02b7', sampled: true }

Sampling to Control Cost

Tracing every request at high traffic is expensive. Samplers decide which traces to keep. The decision propagates via the traceparent sampled flag, so a trace is kept or dropped consistently across all services.

  • AlwaysOnSampler — keep everything (dev/low traffic).
  • TraceIdRatioBasedSampler — keep a fixed fraction, e.g. 10%.
  • ParentBasedSampler — respect the upstream decision; sample new roots by ratio. This is the production default.

Use head sampling (decide at the start) for simplicity, or tail sampling in a collector to always keep errors and slow traces.

const { ParentBasedSampler, TraceIdRatioBasedSampler } = require('@opentelemetry/sdk-trace-base');

// Keep 10% of new root traces; honor upstream decisions for the rest
const sampler = new ParentBasedSampler({
  root: new TraceIdRatioBasedSampler(0.1),
});

// Pass to the NodeSDK: new NodeSDK({ sampler, ... });

Reading the Waterfall to Find Latency

Once spans reach a backend (Jaeger, Tempo, Honeycomb), you read the trace as a waterfall. Each bar is a span; its width is its duration; indentation shows parent/child.

How to find the bottleneck:

  • Look for the widest child bar — that operation dominates the request.
  • Watch for gaps between a parent and its first child — usually queueing, GC pauses, or un-instrumented work.
  • Sequential bars that could run in parallel reveal a chance to use Promise.all.
  • A red span with ERROR status points straight at the failing boundary.

The cross-service value: you can see that 80% of a 900ms request was spent inside the downstream payments service, not your own code.

Quick Check: Active Span Nesting

You manually wrap a function in tracer.startActiveSpan('outer', cb). Inside the callback, your auto-instrumented HTTP client makes an outbound call. Which statement is correct?

Recap & Takeaways

You can now produce spans that expose latency across service boundaries:

  • Trace = many spans sharing a traceId; each span has its own spanId and a parentSpanId.
  • Auto-instrumentation (NodeSDK + auto-instrumentations-node, loaded with node -r) covers HTTP, DB, and queue boundaries for free.
  • Manual spans with tracer.startActiveSpan() capture business logic; always end() in finally and set ERROR status on exceptions.
  • Enrich spans with attributes and events, watching cardinality.
  • Context propagation via the W3C traceparent header makes traces cross services; inject/extract manually for non-HTTP transports.
  • Sampling (ParentBased + ratio) controls cost while keeping decisions consistent across services.
  • Read the waterfall: widest bars, gaps, and serial calls reveal the real bottleneck.
免费开始

用 AI 导师学习 JavaScript — 免费

在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。

课程
22
课程
92

常见问题解答

「使用 OpenTelemetry 跨服务跟踪」课时是免费的吗?

是的 — 「使用 OpenTelemetry 跨服务跟踪」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Node.js Backend Development Bootcamp 课程的其余内容,请升级到 CoddyKit PRO。 Node.js Backend Development Bootcamp 课程共包含 4 节课。

「使用 OpenTelemetry 跨服务跟踪」这节课中我会学到什么?

自动或手动为服务添加插桩,生成能够揭示服务边界间延迟的跨度。 你通过在浏览器中直接运行的动手代码来练习 Node.js Backend Development Bootcamp,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 Node.js Backend Development Bootcamp 需要有经验吗?

无需任何先前经验。CoddyKit 上的 Node.js Backend Development Bootcamp 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「使用 OpenTelemetry 跨服务跟踪」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 Node.js Backend Development Bootcamp 课中编写并运行代码吗?

能。每节 Node.js Backend Development Bootcamp 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 带关联 ID 的结构化日志
  2. 使用 OpenTelemetry 跨服务跟踪
  3. 暴露应用指标与 RED 方法
  4. 使用 AsyncLocalStorage 传播上下文
← 返回 Node.js Backend Development Bootcamp