MongoDB Academy · Lekcja

Uruchamianie potoków agregacji dla wielu źródeł

Nauczą się Państwo pisać potoki agregacji, które w jednym zapytaniu łączą dane kolekcji Atlas z plikami JSON lub Parquet przechowywanymi w S3.

Lekcja 3 z 413 kroki

Uruchamianie potoków agregacji dla wielu źródeł to bezpłatna lekcja MongoDB Academy na CoddyKit. To lekcja 3 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej MongoDB Academy, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs MongoDB Academy zawiera 4 lekcji w sumie.

Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.

What Makes Cross-Source Pipelines Special?

A cross-source aggregation pipeline in Atlas Data Federation runs the same aggregation stages you know from MongoDB, but the data under each stage may come from different physical systems — an S3 bucket, a live Atlas cluster, or both. The federated query engine handles all the routing, fan-out, and result merging transparently. From your application's perspective, it looks like a single MongoDB collection query.

Simple Cross-Source Find

The simplest cross-source query is a find() on a virtual collection backed by S3 files. The query engine reads the files, parses them, and applies the filter. Fields in the filter that match partition attributes in the path cause automatic file pruning. Fields that do not match partition attributes are applied as a post-read filter.

// Virtual collection 'events' backed by S3 JSON files
// Path: /data/events/{year int}/{month int}/*.json

// This query prunes to /data/events/2025/1/ only
const jan2025 = await db.collection('events').find({
  year: 2025,
  month: 1,
  eventType: 'purchase'   // post-read filter (not a partition attr)
}).toArray()

Aggregating S3 Data Like a Live Collection

You can run any aggregation stage against S3-backed virtual collections: $match, $group, $project, $sort, $limit. The query engine pushes down stages where possible (especially $match for partition pruning and column pruning in Parquet) and executes remaining stages in its own compute layer after reading the data.

// Group S3-archived events by region and count
const summary = await db.collection('events_2024').aggregate([
  { $match: { year: 2024, month: { $in: [10, 11, 12] } } },  // pruning
  { $group: { _id: '$region', total: { $sum: 1 }, revenue: { $sum: '$amount' } } },
  { $sort: { revenue: -1 } },
  { $limit: 10 }
]).toArray()

Joining Atlas and S3 With $lookup

The most powerful cross-source pattern is using $lookup to join a live Atlas collection with an S3-archived collection. Start the pipeline from the live collection (the 'driver') and look up into the virtual S3-backed collection. Always place a $match early to minimize the number of lookups performed.

// Join live customers (Atlas) with archived orders (S3)
const result = await db.collection('customers').aggregate([
  { $match: { tier: 'gold', region: 'EU' } },   // filter live data first
  { $lookup: {
    from: 'orders_archive',   // virtual S3-backed collection
    let: { custId: '$_id' },
    pipeline: [
      { $match: { $expr: { $eq: ['$customerId', '$$custId'] } } },
      { $project: { orderId: 1, amount: 1, date: 1 } }
    ],
    as: 'orderHistory'
  }},
  { $addFields: { totalSpend: { $sum: '$orderHistory.amount' } } },
  { $sort: { totalSpend: -1 } },
  { $limit: 50 }
]).toArray()

Aggregating Across Multiple Atlas Clusters

If your federated instance has multiple Atlas cluster stores, you can join collections from different Atlas clusters in a single pipeline. This is useful for multi-tenant or multi-region deployments where data is sharded across separate clusters and you need cross-cluster reports without merging clusters or building a separate reporting database.

// Virtual collections pointing to different Atlas clusters
// 'orders_us' -> Atlas cluster in US
// 'orders_eu' -> Atlas cluster in EU

// Union results from two clusters
db.orders_us.aggregate([
  { $match: { date: { $gte: ISODate('2025-01-01') } } },
  { $unionWith: {
    coll: 'orders_eu',
    pipeline: [{ $match: { date: { $gte: ISODate('2025-01-01') } } }]
  }},
  { $group: { _id: '$status', count: { $sum: 1 } } }
])

Writing Results to Atlas With $out / $merge

After running a cross-source aggregation, you can write the results back to a live Atlas collection using $out or $merge. This is the ETL pattern: read historical data from S3, join with live data, compute aggregates, and write the results to a materialised view collection in Atlas that application queries can then read cheaply and quickly.

// ETL: aggregate S3 archive + Atlas, write result to Atlas
db.events_2024.aggregate([
  { $match: { year: 2024 } },
  { $group: {
    _id: { region: '$region', month: '$month' },
    sessions: { $sum: 1 },
    revenue: { $sum: '$amount' }
  }},
  { $merge: {
    into: { db: 'reporting', coll: 'monthly_summary' },
    whenMatched: 'replace',
    whenNotMatched: 'insert'
  }}
])

Parquet Column Pruning: Only Read What You Need

When querying Parquet files, Data Federation applies column pruning: if your $project stage specifies only certain fields, the query engine reads only those columns from the Parquet file (which stores data column-by-column). This can reduce bytes read by 90%+ compared to reading every column. Place your $project as early as possible in the pipeline for maximum column pruning benefit.

// Column pruning: only reads 'region', 'amount', 'date' columns from Parquet
db.events_2024.aggregate([
  { $project: { region: 1, amount: 1, date: 1, _id: 0 } },  // early project
  { $match: { region: 'EU' } },
  { $group: { _id: '$region', totalRevenue: { $sum: '$amount' } } }
])
// Other columns (userId, sessionId, metadata, etc.) are never read from disk

Federated Query Performance Monitoring

Atlas Data Federation logs query execution details in the Atlas UI under the Query History tab. Each query shows: bytes processed, execution time, and the number of files/partitions scanned. High bytes-processed numbers usually mean partition attributes are missing or the query does not match any partition keys. Use this log to tune your storage configuration and query patterns.

// Get query stats via the admin DB on the federated instance
db.adminCommand({ currentOp: 1 })
// Shows active federated queries with bytes read, duration

// In Atlas UI: Data Federation > Query History
// Shows past queries, duration, data processed, and cost estimate

Handling Schema Differences Across Sources

S3 files from different time periods or systems may have different schemas (field names, types, structure). Data Federation handles this gracefully — missing fields return null, extra fields are included. You can use $ifNull, $cond, and $convert in your pipeline to normalise varying schemas before grouping or joining.

// Normalise schema variations across old and new S3 file formats
db.events.aggregate([
  { $addFields: {
    // Old format: 'user_id', New format: 'userId'
    userId: { $ifNull: ['$userId', '$user_id'] },
    // Old format: string amount, New format: number
    amount: { $convert: { input: '$amount', to: 'double', onError: 0 } }
  }},
  { $group: { _id: '$userId', total: { $sum: '$amount' } } }
])

Caching Federated Query Results

Data Federation does not cache results between queries — each query re-reads the underlying sources. For dashboards that run the same report repeatedly, use the ETL pattern: schedule an aggregation that writes results to an Atlas collection via $merge, then have your dashboard query the fast Atlas collection. Atlas Triggers can schedule this refresh on any cron interval.

Limitations of Cross-Source Pipelines

Be aware of current limitations: 1) Transactions are not supported on federated instances. 2) Index usage only applies to Atlas-backed collections, not S3 files. 3) Very large result sets may time out — use $out/$merge to write results instead of streaming them back. 4) Latency is higher than a live Atlas query due to S3 I/O — not suitable for user-facing, real-time queries.

Quick Check

Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.

Lesson Recap

In this lesson you learned: cross-source aggregation pipelines use the same MongoDB stages against virtual collections backed by S3 or Atlas clusters, $lookup enables joining live Atlas data with S3 archives in a single pipeline, and early $project enables column pruning in Parquet files to dramatically reduce bytes scanned. Next up we explore S3 data partitioning for query performance.

Bezpłatny start

Ucz się JavaScript dzięki korepetycjom AI — za darmo

Pisz i uruchamiaj kod w przeglądarce, otrzymuj natychmiastową pomoc od korepetytora AI dostępnego 24/7 i kontynuuj naukę w sieci lub w aplikacji.

Kursy
30
Lekcje
120

Często zadawane pytania

Czy lekcja „Uruchamianie potoków agregacji dla wielu źródeł” jest bezpłatna?

Tak — pełny tekst „Uruchamianie potoków agregacji dla wielu źródeł” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu MongoDB Academy, przejdź na CoddyKit PRO. Kurs MongoDB Academy zawiera 4 lekcji w sumie.

Co nauczysz się w „Uruchamianie potoków agregacji dla wielu źródeł”?

Nauczą się Państwo pisać potoki agregacji, które w jednym zapytaniu łączą dane kolekcji Atlas z plikami JSON lub Parquet przechowywanymi w S3. Ćwiczysz MongoDB Academy z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.

Czy potrzebuję doświadczenia, aby zacząć MongoDB Academy?

Nie wymagamy żadnego doświadczenia. MongoDB Academy w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 3 z 4.

Ile czasu zajmuje lekcja „Uruchamianie potoków agregacji dla wielu źródeł”?

Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.

Czy mogę pisać i uruchamiać kod w tej lekcji MongoDB Academy?

Tak. Każda lekcja MongoDB Academy zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.

Wszystkie lekcje w tym kursie

  1. Czym jest Atlas Data Federation
  2. Mapowanie źródeł S3 i Atlas na wirtualną przestrzeń nazw
  3. Uruchamianie potoków agregacji dla wielu źródeł
  4. Partycjonowanie danych S3 na potrzeby wydajności zapytań
← Powrót do MongoDB Academy