0Pricing
Node.js Backend Development Bootcamp · Lección

Implementación de streams Transform personalizados con _transform y _flush

Construya streams Transform reutilizables que modifiquen, filtren y agreguen fragmentos a medida que fluyen los datos.

Implementación de streams Transform personalizados con _transform y _flush es una lección gratuita de Node.js Backend Development Bootcamp en CoddyKit. Esta es la lección 2 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Node.js Backend Development Bootcamp, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Node.js Backend Development Bootcamp incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

Why Transform Streams?

A Transform stream is both readable and writable: it consumes input chunks, processes them, and pushes output chunks. It is the right tool whenever data must be mutated as it flows rather than buffered fully in memory.

  • Writable side accepts data via write() / pipe() from upstream.
  • Readable side emits processed data that downstream consumers read.

Typical backend uses: uppercasing/normalizing a payload, gzip-style encoding, CSV-to-JSON conversion, redacting secrets in a log pipeline, or counting bytes — all without loading the whole file or HTTP body into RAM.

The Two Hooks: _transform and _flush

A custom Transform stream is defined by implementing two internal methods. Node calls them for you — you never call them directly.

  • _transform(chunk, encoding, callback) — invoked once per incoming chunk. Do your work, push() any output, then signal completion with callback().
  • _flush(callback) — invoked once, after the last chunk, just before the stream ends. Use it to emit any trailing/aggregated data.

The leading underscore marks them as the framework-facing implementation. Consumers still use the public write, read, and pipe API.

A Minimal Uppercase Transform

The classic starting point: extend the Transform class and override _transform. Each chunk is a Buffer (unless object mode), so convert to a string, transform it, and push the result.

Calling callback() with no error tells Node this chunk is fully processed and it may deliver the next one. Passing the value as the second arg to callback is shorthand for push + callback().

const { Transform } = require('node:stream');

class Upper extends Transform {
  _transform(chunk, encoding, callback) {
    const out = chunk.toString().toUpperCase();
    callback(null, out); // shorthand for this.push(out); callback();
  }
}

const up = new Upper();
up.on('data', (d) => process.stdout.write(d));
up.write('hello ');
up.write('streams\n');
up.end();

callback() Is a Contract

The callback in _transform is mandatory. Until you call it, Node assumes the chunk is still in progress and will not hand you the next one. This is how backpressure flows through your transform.

  • callback() — success, ready for next chunk.
  • callback(err) — emits an 'error' event and destroys the stream.
  • callback(null, data) — pushes data and signals success.

Forgetting to call callback is the #1 bug: the pipeline silently stalls forever with no error.

push() Multiple Times Per Chunk

One input chunk can produce zero, one, or many output chunks. Call this.push() as many times as needed before invoking callback(). This is what makes splitting (e.g. line-by-line) possible.

Below, a single write containing several lines is fanned out into one push per line.

const { Transform } = require('node:stream');

class LineSplitter extends Transform {
  _transform(chunk, encoding, callback) {
    const lines = chunk.toString().split('\n');
    for (const line of lines) {
      if (line.length) this.push(line + ' <<\n');
    }
    callback();
  }
}

const s = new LineSplitter();
s.on('data', (d) => process.stdout.write(d));
s.end('alpha\nbeta\ngamma\n');

Filtering: Drop Chunks by Pushing Nothing

To filter, simply decide not to push. If a chunk should be discarded, call callback() without pushing anything — the data never reaches the readable side.

This pattern is ideal for redacting or dropping records mid-pipeline, e.g. removing log lines that contain a secret token.

const { Transform } = require('node:stream');

class DropSecrets extends Transform {
  _transform(chunk, encoding, callback) {
    const line = chunk.toString();
    if (line.includes('SECRET')) {
      return callback(); // filtered out, nothing pushed
    }
    callback(null, line);
  }
}

const f = new DropSecrets();
f.on('data', (d) => process.stdout.write(d));
f.write('ok line 1\n');
f.write('this has a SECRET token\n');
f.write('ok line 2\n');
f.end();

Object Mode for Structured Records

By default chunks are Buffer/string. Set objectMode: true to push and receive JavaScript objects instead — essential for record-oriented pipelines (JSON rows, DB results, parsed events).

  • writableObjectMode / readableObjectMode can be set independently if input and output types differ.
  • In object mode, each push emits exactly one object regardless of size.
const { Transform } = require('node:stream');

class AddTax extends Transform {
  constructor() { super({ objectMode: true }); }
  _transform(order, encoding, callback) {
    callback(null, { ...order, total: order.price * 1.2 });
  }
}

const t = new AddTax();
t.on('data', (o) => console.log(o));
t.write({ id: 1, price: 100 });
t.write({ id: 2, price: 250 });
t.end();

_flush: Emit Trailing/Aggregated Data

_flush(callback) runs once after the final chunk, before 'end'. It is where you push anything you have been accumulating — a running total, a buffered partial line, or a closing delimiter.

You can push inside _flush just like in _transform. You must call its callback() so the stream can finish.

const { Transform } = require('node:stream');

class Summer extends Transform {
  constructor() { super({ objectMode: true }); this.sum = 0; }
  _transform(num, encoding, callback) {
    this.sum += num;
    callback(); // aggregate, emit nothing yet
  }
  _flush(callback) {
    this.push({ total: this.sum }); // emit once at the end
    callback();
  }
}

const agg = new Summer();
agg.on('data', (o) => console.log(o));
[10, 20, 30, 40].forEach((n) => agg.write(n));
agg.end();

Buffering Partial Lines Across Chunk Boundaries

Chunks do not align with logical records. A line may be split across two chunks, so a robust line-parser keeps a leftover buffer between calls and flushes the remainder in _flush.

This combine-in-_transform, drain-in-_flush pattern is the backbone of real CSV/NDJSON parsers.

const { Transform } = require('node:stream');

class LineParser extends Transform {
  constructor() { super({ readableObjectMode: true }); this.tail = ''; }
  _transform(chunk, encoding, callback) {
    const data = this.tail + chunk.toString();
    const parts = data.split('\n');
    this.tail = parts.pop(); // keep incomplete last segment
    for (const line of parts) this.push(line);
    callback();
  }
  _flush(callback) {
    if (this.tail) this.push(this.tail);
    callback();
  }
}

const p = new LineParser();
p.on('data', (l) => console.log('LINE:', l));
p.write('he');
p.write('llo\nwor');
p.write('ld\nlast');
p.end();

The Functional Shorthand: stream.Transform options

You don't always need a class. The Transform constructor accepts transform and flush functions directly — handy for small, one-off transforms.

Inside these functions, this is still the stream, so this.push() works. The class form is preferred when you want reusable, named, instantiable components; the inline form is great for quick glue.

const { Transform } = require('node:stream');

const csvToRows = new Transform({
  readableObjectMode: true,
  transform(chunk, encoding, callback) {
    for (const line of chunk.toString().trim().split('\n')) {
      const [id, name] = line.split(',');
      this.push({ id: Number(id), name });
    }
    callback();
  },
});

csvToRows.on('data', (row) => console.log(row));
csvToRows.end('1,Ada\n2,Linus\n3,Grace');

Composing in a Pipeline

Transform streams shine when chained. Use stream.pipeline() (callback or promise form) instead of raw .pipe() — it propagates errors and destroys every stream on failure, preventing leaks.

Here a source feeds an uppercaser, then a suffix-adder, then stdout. Each transform stays small and reusable.

const { Transform, Readable, pipeline } = require('node:stream');

const make = (fn) => new Transform({
  transform(chunk, enc, cb) { cb(null, fn(chunk.toString())); },
});

const upper = make((s) => s.toUpperCase());
const bang = make((s) => s + '!\n');

pipeline(
  Readable.from(['log a\n', 'log b\n']),
  upper,
  bang,
  process.stdout,
  (err) => { if (err) console.error('failed', err); else console.error('done'); }
);

Quick Check: Where Do Trailing Aggregates Go?

You are building a Transform that counts total bytes seen and must emit a single summary object after all input is processed. Which method should push that summary?

Recap: Transform Stream Essentials

You can now build reusable Transform streams that mutate, filter, and aggregate flowing data:

  • _transform(chunk, enc, cb) — process each chunk; push zero or more outputs; always call cb() (or cb(null, data)).
  • _flush(cb) — runs once at the end to emit trailing or aggregated data; must call cb().
  • Filter by pushing nothing; split by pushing many times; aggregate by accumulating state and flushing.
  • objectMode (and the independent readable/writable variants) carries structured records.
  • Buffer partial records in _transform and drain them in _flush.
  • Compose with stream.pipeline() for safe error handling and cleanup.

Never forget the callback — a missing cb() silently stalls the entire pipeline.

Preguntas frecuentes

¿La lección «Implementación de streams Transform personalizados con _transform y _flush» es gratis?

Sí — el texto completo de «Implementación de streams Transform personalizados con _transform y _flush» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Node.js Backend Development Bootcamp, actualiza a CoddyKit PRO. El curso de Node.js Backend Development Bootcamp incluye 4 lecciones en total.

¿Qué aprenderé en «Implementación de streams Transform personalizados con _transform y _flush»?

Construya streams Transform reutilizables que modifiquen, filtren y agreguen fragmentos a medida que fluyen los datos. Practicas Node.js Backend Development Bootcamp con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar Node.js Backend Development Bootcamp?

No se requiere experiencia previa. Node.js Backend Development Bootcamp en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 2 de 4.

¿Cuánto tiempo toma la lección «Implementación de streams Transform personalizados con _transform y _flush»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de Node.js Backend Development Bootcamp?

Sí. Cada lección de Node.js Backend Development Bootcamp incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. Aspectos internos de los streams Readable, Writable, Duplex y Transform
  2. Implementación de streams Transform personalizados con _transform y _flush
  3. Backpressure, pipe() y la utilidad pipeline()
  4. Iteradores asíncronos y for-await-of sobre streams
← Volver a Node.js Backend Development Bootcamp