使用 OpenTelemetry 跨服务跟踪
自动或手动为服务添加插桩,生成能够揭示服务边界间延迟的跨度。
使用 OpenTelemetry 跨服务跟踪 是 CoddyKit 上的免费 Node.js Backend Development Bootcamp 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 Node.js Backend Development Bootcamp 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 Node.js Backend Development Bootcamp 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Why Distributed Tracing?
In a microservices system one user request can hop across an API gateway, an orders service, a payments service, and a database. When that request is slow, a single service's logs cannot tell you where the time went.
Distributed tracing stitches the whole journey together. Each unit of work becomes a span, spans are linked into a trace, and the trace reveals latency across every service boundary.
- Trace: the entire end-to-end request, identified by a
traceId. - Span: one operation (an HTTP call, a DB query) with a start time, duration, and parent.
- Context propagation: passing
traceIdandspanIdacross service boundaries, usually via HTTP headers.
OpenTelemetry (OTel) is the vendor-neutral standard for producing these spans in Node.js.
Anatomy of a Span
A span is the atomic building block of a trace. Every span carries the same trace identity but its own identity and timing.
traceId: 16 bytes, shared by every span in the trace.spanId: 8 bytes, unique to this span.parentSpanId: links this span to the operation that caused it.name,startTime,endTime(duration = end - start).- Attributes: key/value tags like
http.methodordb.system. - Status:
OK,ERROR, orUNSET.
Parent/child links form a tree. The root span is the whole request; child spans are the calls it makes. Visualized, the tree becomes the familiar waterfall you see in Jaeger or Tempo.
Auto-Instrumentation with the Node SDK
The fastest way to get spans is auto-instrumentation. The OTel Node SDK monkey-patches popular libraries (http, Express, pg, ioredis, etc.) so they emit spans without you writing tracing code.
Create a tracing.js file that starts the SDK before anything else, then run your app with node -r ./tracing.js app.js so it loads first.
// tracing.js
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-http');
const { Resource } = require('@opentelemetry/resources');
const { SemanticResourceAttributes } = require('@opentelemetry/semantic-conventions');
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'orders-service',
}),
traceExporter: new OTLPTraceExporter({
url: 'http://localhost:4318/v1/traces',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();What Auto-Instrumentation Gives You
With the SDK loaded, an incoming HTTP request automatically becomes a root span, and any outgoing http/fetch call or pg query becomes a child span underneath it.
- Inbound Express route → server span with
http.method,http.route,http.status_code. - Outbound HTTP call → client span, and the headers are injected automatically.
- Database query → client span with
db.systemand the statement.
This covers the boundaries for free. But auto-instrumentation does not understand your business logic — pricing rules, cache decisions, batch loops. For those you add spans manually.
Getting a Tracer
To create spans manually you first obtain a tracer from the global trace API. Name it after the module or library producing the spans; the version is optional but helps when debugging instrumentation.
The tracer is the factory for all your manual spans.
const { trace } = require('@opentelemetry/api');
// Name + version identify the instrumentation scope
const tracer = trace.getTracer('orders-service', '1.0.0');
// Later, anywhere in the code:
// const span = tracer.startSpan('chargeCustomer');startActiveSpan: The Idiomatic Pattern
Prefer tracer.startActiveSpan() over startSpan(). startActiveSpan makes the new span the active span for the duration of its callback, so any child spans created inside (including auto-instrumented ones) automatically attach as children.
The golden rules: always span.end() in a finally block, and record errors plus an ERROR status on failure.
const { trace, SpanStatusCode } = require('@opentelemetry/api');
const tracer = trace.getTracer('orders-service');
async function processOrder(order) {
return tracer.startActiveSpan('processOrder', async (span) => {
try {
span.setAttribute('order.id', order.id);
span.setAttribute('order.items', order.items.length);
const result = await chargeAndShip(order); // child spans nest here
span.setStatus({ code: SpanStatusCode.OK });
return result;
} catch (err) {
span.recordException(err);
span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
throw err;
} finally {
span.end();
}
});
}Attributes, Events, and Status
Spans become useful when you enrich them. Three tools:
- Attributes — searchable key/value tags. Use semantic conventions (
http.method,db.system,messaging.system) so backends understand them. - Events — timestamped log lines anchored inside the span, e.g.
span.addEvent('cache.miss'). - Status — set
ERRORonly on real failures; leave success asUNSETorOK.
Keep cardinality sane: never put a raw user ID or full SQL with literals into a high-traffic attribute if your backend indexes it — it can explode storage.
function readFromCache(span, key) {
const hit = cache.has(key);
if (hit) {
span.addEvent('cache.hit', { 'cache.key': key });
} else {
span.addEvent('cache.miss', { 'cache.key': key });
}
span.setAttribute('cache.hit', hit);
return hit ? cache.get(key) : null;
}Context Propagation Across Services
A trace only spans services if the trace context travels with the request. The W3C traceparent header carries the traceId, parent spanId, and sampling flag.
Auto-instrumentation injects and extracts this header for you on standard HTTP. When you do something non-standard (a custom transport, a message queue), you inject/extract manually with the propagation API.
const { context, propagation, trace } = require('@opentelemetry/api');
// SENDER: inject current context into outgoing carrier (e.g. message headers)
function publish(queue, payload) {
const headers = {};
propagation.inject(context.active(), headers);
queue.send({ payload, headers }); // traceparent now travels with the message
}
// RECEIVER: extract context and continue the trace
function onMessage(msg) {
const parentCtx = propagation.extract(context.active(), msg.headers);
const tracer = trace.getTracer('worker');
context.with(parentCtx, () => {
tracer.startActiveSpan('handleMessage', (span) => {
handle(msg.payload);
span.end();
});
});
}Reading the traceparent Header
The traceparent header has a fixed, parseable shape. Understanding it helps you debug broken traces (a missing child usually means a dropped header).
Format: version-traceId-parentId-flags, e.g.00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
00— version- 32 hex chars — the
traceId - 16 hex chars — the parent
spanId 01— flags (bit 0 = sampled)
Here is a tiny standalone parser to make the structure concrete.
function parseTraceparent(header) {
const parts = header.split('-');
if (parts.length !== 4) throw new Error('invalid traceparent');
const [version, traceId, parentId, flags] = parts;
return {
version,
traceId,
parentId,
sampled: (parseInt(flags, 16) & 1) === 1,
};
}
const h = '00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01';
console.log(parseTraceparent(h));
// { version: '00', traceId: '4bf9...4736', parentId: '00f0...02b7', sampled: true }Sampling to Control Cost
Tracing every request at high traffic is expensive. Samplers decide which traces to keep. The decision propagates via the traceparent sampled flag, so a trace is kept or dropped consistently across all services.
AlwaysOnSampler— keep everything (dev/low traffic).TraceIdRatioBasedSampler— keep a fixed fraction, e.g. 10%.ParentBasedSampler— respect the upstream decision; sample new roots by ratio. This is the production default.
Use head sampling (decide at the start) for simplicity, or tail sampling in a collector to always keep errors and slow traces.
const { ParentBasedSampler, TraceIdRatioBasedSampler } = require('@opentelemetry/sdk-trace-base');
// Keep 10% of new root traces; honor upstream decisions for the rest
const sampler = new ParentBasedSampler({
root: new TraceIdRatioBasedSampler(0.1),
});
// Pass to the NodeSDK: new NodeSDK({ sampler, ... });Reading the Waterfall to Find Latency
Once spans reach a backend (Jaeger, Tempo, Honeycomb), you read the trace as a waterfall. Each bar is a span; its width is its duration; indentation shows parent/child.
How to find the bottleneck:
- Look for the widest child bar — that operation dominates the request.
- Watch for gaps between a parent and its first child — usually queueing, GC pauses, or un-instrumented work.
- Sequential bars that could run in parallel reveal a chance to use
Promise.all. - A red span with ERROR status points straight at the failing boundary.
The cross-service value: you can see that 80% of a 900ms request was spent inside the downstream payments service, not your own code.
Quick Check: Active Span Nesting
You manually wrap a function in tracer.startActiveSpan('outer', cb). Inside the callback, your auto-instrumented HTTP client makes an outbound call. Which statement is correct?
Recap & Takeaways
You can now produce spans that expose latency across service boundaries:
- Trace = many spans sharing a
traceId; each span has its ownspanIdand aparentSpanId. - Auto-instrumentation (NodeSDK + auto-instrumentations-node, loaded with
node -r) covers HTTP, DB, and queue boundaries for free. - Manual spans with
tracer.startActiveSpan()capture business logic; alwaysend()infinallyand set ERROR status on exceptions. - Enrich spans with attributes and events, watching cardinality.
- Context propagation via the W3C
traceparentheader makes traces cross services; inject/extract manually for non-HTTP transports. - Sampling (ParentBased + ratio) controls cost while keeping decisions consistent across services.
- Read the waterfall: widest bars, gaps, and serial calls reveal the real bottleneck.
用 AI 导师学习 JavaScript — 免费
在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。
- 课程
- 22
- 课程
- 92
常见问题解答
「使用 OpenTelemetry 跨服务跟踪」课时是免费的吗?
是的 — 「使用 OpenTelemetry 跨服务跟踪」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 Node.js Backend Development Bootcamp 课程的其余内容,请升级到 CoddyKit PRO。 Node.js Backend Development Bootcamp 课程共包含 4 节课。
「使用 OpenTelemetry 跨服务跟踪」这节课中我会学到什么?
自动或手动为服务添加插桩,生成能够揭示服务边界间延迟的跨度。 你通过在浏览器中直接运行的动手代码来练习 Node.js Backend Development Bootcamp,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 Node.js Backend Development Bootcamp 需要有经验吗?
无需任何先前经验。CoddyKit 上的 Node.js Backend Development Bootcamp 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。
「使用 OpenTelemetry 跨服务跟踪」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 Node.js Backend Development Bootcamp 课中编写并运行代码吗?
能。每节 Node.js Backend Development Bootcamp 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 带关联 ID 的结构化日志
- 使用 OpenTelemetry 跨服务跟踪
- 暴露应用指标与 RED 方法
- 使用 AsyncLocalStorage 传播上下文