0Pricing
MongoDB Academy · Lección

Ejecución de pipelines de agregación entre fuentes

Escribirá pipelines de agregación que combinen datos de colecciones de Atlas con archivos JSON o Parquet almacenados en S3 en una sola consulta.

Ejecución de pipelines de agregación entre fuentes es una lección gratuita de MongoDB Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de MongoDB Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de MongoDB Academy incluye 4 lecciones en total.

Partes de esta lección aún no han sido traducidas y se muestran en inglés.

What Makes Cross-Source Pipelines Special?

A cross-source aggregation pipeline in Atlas Data Federation runs the same aggregation stages you know from MongoDB, but the data under each stage may come from different physical systems — an S3 bucket, a live Atlas cluster, or both. The federated query engine handles all the routing, fan-out, and result merging transparently. From your application's perspective, it looks like a single MongoDB collection query.

Simple Cross-Source Find

The simplest cross-source query is a find() on a virtual collection backed by S3 files. The query engine reads the files, parses them, and applies the filter. Fields in the filter that match partition attributes in the path cause automatic file pruning. Fields that do not match partition attributes are applied as a post-read filter.

// Virtual collection 'events' backed by S3 JSON files
// Path: /data/events/{year int}/{month int}/*.json

// This query prunes to /data/events/2025/1/ only
const jan2025 = await db.collection('events').find({
  year: 2025,
  month: 1,
  eventType: 'purchase'   // post-read filter (not a partition attr)
}).toArray()

Aggregating S3 Data Like a Live Collection

You can run any aggregation stage against S3-backed virtual collections: $match, $group, $project, $sort, $limit. The query engine pushes down stages where possible (especially $match for partition pruning and column pruning in Parquet) and executes remaining stages in its own compute layer after reading the data.

// Group S3-archived events by region and count
const summary = await db.collection('events_2024').aggregate([
  { $match: { year: 2024, month: { $in: [10, 11, 12] } } },  // pruning
  { $group: { _id: '$region', total: { $sum: 1 }, revenue: { $sum: '$amount' } } },
  { $sort: { revenue: -1 } },
  { $limit: 10 }
]).toArray()

Joining Atlas and S3 With $lookup

The most powerful cross-source pattern is using $lookup to join a live Atlas collection with an S3-archived collection. Start the pipeline from the live collection (the 'driver') and look up into the virtual S3-backed collection. Always place a $match early to minimize the number of lookups performed.

// Join live customers (Atlas) with archived orders (S3)
const result = await db.collection('customers').aggregate([
  { $match: { tier: 'gold', region: 'EU' } },   // filter live data first
  { $lookup: {
    from: 'orders_archive',   // virtual S3-backed collection
    let: { custId: '$_id' },
    pipeline: [
      { $match: { $expr: { $eq: ['$customerId', '$$custId'] } } },
      { $project: { orderId: 1, amount: 1, date: 1 } }
    ],
    as: 'orderHistory'
  }},
  { $addFields: { totalSpend: { $sum: '$orderHistory.amount' } } },
  { $sort: { totalSpend: -1 } },
  { $limit: 50 }
]).toArray()

Aggregating Across Multiple Atlas Clusters

If your federated instance has multiple Atlas cluster stores, you can join collections from different Atlas clusters in a single pipeline. This is useful for multi-tenant or multi-region deployments where data is sharded across separate clusters and you need cross-cluster reports without merging clusters or building a separate reporting database.

// Virtual collections pointing to different Atlas clusters
// 'orders_us' -> Atlas cluster in US
// 'orders_eu' -> Atlas cluster in EU

// Union results from two clusters
db.orders_us.aggregate([
  { $match: { date: { $gte: ISODate('2025-01-01') } } },
  { $unionWith: {
    coll: 'orders_eu',
    pipeline: [{ $match: { date: { $gte: ISODate('2025-01-01') } } }]
  }},
  { $group: { _id: '$status', count: { $sum: 1 } } }
])

Writing Results to Atlas With $out / $merge

After running a cross-source aggregation, you can write the results back to a live Atlas collection using $out or $merge. This is the ETL pattern: read historical data from S3, join with live data, compute aggregates, and write the results to a materialised view collection in Atlas that application queries can then read cheaply and quickly.

// ETL: aggregate S3 archive + Atlas, write result to Atlas
db.events_2024.aggregate([
  { $match: { year: 2024 } },
  { $group: {
    _id: { region: '$region', month: '$month' },
    sessions: { $sum: 1 },
    revenue: { $sum: '$amount' }
  }},
  { $merge: {
    into: { db: 'reporting', coll: 'monthly_summary' },
    whenMatched: 'replace',
    whenNotMatched: 'insert'
  }}
])

Parquet Column Pruning: Only Read What You Need

When querying Parquet files, Data Federation applies column pruning: if your $project stage specifies only certain fields, the query engine reads only those columns from the Parquet file (which stores data column-by-column). This can reduce bytes read by 90%+ compared to reading every column. Place your $project as early as possible in the pipeline for maximum column pruning benefit.

// Column pruning: only reads 'region', 'amount', 'date' columns from Parquet
db.events_2024.aggregate([
  { $project: { region: 1, amount: 1, date: 1, _id: 0 } },  // early project
  { $match: { region: 'EU' } },
  { $group: { _id: '$region', totalRevenue: { $sum: '$amount' } } }
])
// Other columns (userId, sessionId, metadata, etc.) are never read from disk

Federated Query Performance Monitoring

Atlas Data Federation logs query execution details in the Atlas UI under the Query History tab. Each query shows: bytes processed, execution time, and the number of files/partitions scanned. High bytes-processed numbers usually mean partition attributes are missing or the query does not match any partition keys. Use this log to tune your storage configuration and query patterns.

// Get query stats via the admin DB on the federated instance
db.adminCommand({ currentOp: 1 })
// Shows active federated queries with bytes read, duration

// In Atlas UI: Data Federation > Query History
// Shows past queries, duration, data processed, and cost estimate

Handling Schema Differences Across Sources

S3 files from different time periods or systems may have different schemas (field names, types, structure). Data Federation handles this gracefully — missing fields return null, extra fields are included. You can use $ifNull, $cond, and $convert in your pipeline to normalise varying schemas before grouping or joining.

// Normalise schema variations across old and new S3 file formats
db.events.aggregate([
  { $addFields: {
    // Old format: 'user_id', New format: 'userId'
    userId: { $ifNull: ['$userId', '$user_id'] },
    // Old format: string amount, New format: number
    amount: { $convert: { input: '$amount', to: 'double', onError: 0 } }
  }},
  { $group: { _id: '$userId', total: { $sum: '$amount' } } }
])

Caching Federated Query Results

Data Federation does not cache results between queries — each query re-reads the underlying sources. For dashboards that run the same report repeatedly, use the ETL pattern: schedule an aggregation that writes results to an Atlas collection via $merge, then have your dashboard query the fast Atlas collection. Atlas Triggers can schedule this refresh on any cron interval.

Limitations of Cross-Source Pipelines

Be aware of current limitations: 1) Transactions are not supported on federated instances. 2) Index usage only applies to Atlas-backed collections, not S3 files. 3) Very large result sets may time out — use $out/$merge to write results instead of streaming them back. 4) Latency is higher than a live Atlas query due to S3 I/O — not suitable for user-facing, real-time queries.

Quick Check

Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.

Lesson Recap

In this lesson you learned: cross-source aggregation pipelines use the same MongoDB stages against virtual collections backed by S3 or Atlas clusters, $lookup enables joining live Atlas data with S3 archives in a single pipeline, and early $project enables column pruning in Parquet files to dramatically reduce bytes scanned. Next up we explore S3 data partitioning for query performance.

Preguntas frecuentes

¿La lección «Ejecución de pipelines de agregación entre fuentes» es gratis?

Sí — el texto completo de «Ejecución de pipelines de agregación entre fuentes» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de MongoDB Academy, actualiza a CoddyKit PRO. El curso de MongoDB Academy incluye 4 lecciones en total.

¿Qué aprenderé en «Ejecución de pipelines de agregación entre fuentes»?

Escribirá pipelines de agregación que combinen datos de colecciones de Atlas con archivos JSON o Parquet almacenados en S3 en una sola consulta. Practicas MongoDB Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.

¿Necesito experiencia previa para empezar MongoDB Academy?

No se requiere experiencia previa. MongoDB Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.

¿Cuánto tiempo toma la lección «Ejecución de pipelines de agregación entre fuentes»?

La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.

¿Puedo escribir y ejecutar código en esta lección de MongoDB Academy?

Sí. Cada lección de MongoDB Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.

Todas las lecciones de este curso

  1. ¿Qué es Atlas Data Federation?
  2. Asignación de fuentes de S3 y Atlas a un espacio de nombres virtual
  3. Ejecución de pipelines de agregación entre fuentes
  4. Partición de datos de S3 para mejorar el rendimiento de las consultas
← Volver a MongoDB Academy