複数ソースにまたがる集約パイプラインの実行
1つのクエリで、AtlasコレクションのデータとS3に保存されたJSONまたはParquetファイルを結合する集約パイプラインを記述します。
「複数ソースにまたがる集約パイプラインの実行」はCoddyKit上の無料MongoDB Academyレッスンです。 これはレッスン3/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはMongoDB Academy学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 MongoDB Academyコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
What Makes Cross-Source Pipelines Special?
A cross-source aggregation pipeline in Atlas Data Federation runs the same aggregation stages you know from MongoDB, but the data under each stage may come from different physical systems — an S3 bucket, a live Atlas cluster, or both. The federated query engine handles all the routing, fan-out, and result merging transparently. From your application's perspective, it looks like a single MongoDB collection query.
Simple Cross-Source Find
The simplest cross-source query is a find() on a virtual collection backed by S3 files. The query engine reads the files, parses them, and applies the filter. Fields in the filter that match partition attributes in the path cause automatic file pruning. Fields that do not match partition attributes are applied as a post-read filter.
// Virtual collection 'events' backed by S3 JSON files
// Path: /data/events/{year int}/{month int}/*.json
// This query prunes to /data/events/2025/1/ only
const jan2025 = await db.collection('events').find({
year: 2025,
month: 1,
eventType: 'purchase' // post-read filter (not a partition attr)
}).toArray()Aggregating S3 Data Like a Live Collection
You can run any aggregation stage against S3-backed virtual collections: $match, $group, $project, $sort, $limit. The query engine pushes down stages where possible (especially $match for partition pruning and column pruning in Parquet) and executes remaining stages in its own compute layer after reading the data.
// Group S3-archived events by region and count
const summary = await db.collection('events_2024').aggregate([
{ $match: { year: 2024, month: { $in: [10, 11, 12] } } }, // pruning
{ $group: { _id: '$region', total: { $sum: 1 }, revenue: { $sum: '$amount' } } },
{ $sort: { revenue: -1 } },
{ $limit: 10 }
]).toArray()Joining Atlas and S3 With $lookup
The most powerful cross-source pattern is using $lookup to join a live Atlas collection with an S3-archived collection. Start the pipeline from the live collection (the 'driver') and look up into the virtual S3-backed collection. Always place a $match early to minimize the number of lookups performed.
// Join live customers (Atlas) with archived orders (S3)
const result = await db.collection('customers').aggregate([
{ $match: { tier: 'gold', region: 'EU' } }, // filter live data first
{ $lookup: {
from: 'orders_archive', // virtual S3-backed collection
let: { custId: '$_id' },
pipeline: [
{ $match: { $expr: { $eq: ['$customerId', '$$custId'] } } },
{ $project: { orderId: 1, amount: 1, date: 1 } }
],
as: 'orderHistory'
}},
{ $addFields: { totalSpend: { $sum: '$orderHistory.amount' } } },
{ $sort: { totalSpend: -1 } },
{ $limit: 50 }
]).toArray()Aggregating Across Multiple Atlas Clusters
If your federated instance has multiple Atlas cluster stores, you can join collections from different Atlas clusters in a single pipeline. This is useful for multi-tenant or multi-region deployments where data is sharded across separate clusters and you need cross-cluster reports without merging clusters or building a separate reporting database.
// Virtual collections pointing to different Atlas clusters
// 'orders_us' -> Atlas cluster in US
// 'orders_eu' -> Atlas cluster in EU
// Union results from two clusters
db.orders_us.aggregate([
{ $match: { date: { $gte: ISODate('2025-01-01') } } },
{ $unionWith: {
coll: 'orders_eu',
pipeline: [{ $match: { date: { $gte: ISODate('2025-01-01') } } }]
}},
{ $group: { _id: '$status', count: { $sum: 1 } } }
])Writing Results to Atlas With $out / $merge
After running a cross-source aggregation, you can write the results back to a live Atlas collection using $out or $merge. This is the ETL pattern: read historical data from S3, join with live data, compute aggregates, and write the results to a materialised view collection in Atlas that application queries can then read cheaply and quickly.
// ETL: aggregate S3 archive + Atlas, write result to Atlas
db.events_2024.aggregate([
{ $match: { year: 2024 } },
{ $group: {
_id: { region: '$region', month: '$month' },
sessions: { $sum: 1 },
revenue: { $sum: '$amount' }
}},
{ $merge: {
into: { db: 'reporting', coll: 'monthly_summary' },
whenMatched: 'replace',
whenNotMatched: 'insert'
}}
])Parquet Column Pruning: Only Read What You Need
When querying Parquet files, Data Federation applies column pruning: if your $project stage specifies only certain fields, the query engine reads only those columns from the Parquet file (which stores data column-by-column). This can reduce bytes read by 90%+ compared to reading every column. Place your $project as early as possible in the pipeline for maximum column pruning benefit.
// Column pruning: only reads 'region', 'amount', 'date' columns from Parquet
db.events_2024.aggregate([
{ $project: { region: 1, amount: 1, date: 1, _id: 0 } }, // early project
{ $match: { region: 'EU' } },
{ $group: { _id: '$region', totalRevenue: { $sum: '$amount' } } }
])
// Other columns (userId, sessionId, metadata, etc.) are never read from diskFederated Query Performance Monitoring
Atlas Data Federation logs query execution details in the Atlas UI under the Query History tab. Each query shows: bytes processed, execution time, and the number of files/partitions scanned. High bytes-processed numbers usually mean partition attributes are missing or the query does not match any partition keys. Use this log to tune your storage configuration and query patterns.
// Get query stats via the admin DB on the federated instance
db.adminCommand({ currentOp: 1 })
// Shows active federated queries with bytes read, duration
// In Atlas UI: Data Federation > Query History
// Shows past queries, duration, data processed, and cost estimateHandling Schema Differences Across Sources
S3 files from different time periods or systems may have different schemas (field names, types, structure). Data Federation handles this gracefully — missing fields return null, extra fields are included. You can use $ifNull, $cond, and $convert in your pipeline to normalise varying schemas before grouping or joining.
// Normalise schema variations across old and new S3 file formats
db.events.aggregate([
{ $addFields: {
// Old format: 'user_id', New format: 'userId'
userId: { $ifNull: ['$userId', '$user_id'] },
// Old format: string amount, New format: number
amount: { $convert: { input: '$amount', to: 'double', onError: 0 } }
}},
{ $group: { _id: '$userId', total: { $sum: '$amount' } } }
])Caching Federated Query Results
Data Federation does not cache results between queries — each query re-reads the underlying sources. For dashboards that run the same report repeatedly, use the ETL pattern: schedule an aggregation that writes results to an Atlas collection via $merge, then have your dashboard query the fast Atlas collection. Atlas Triggers can schedule this refresh on any cron interval.
Limitations of Cross-Source Pipelines
Be aware of current limitations: 1) Transactions are not supported on federated instances. 2) Index usage only applies to Atlas-backed collections, not S3 files. 3) Very large result sets may time out — use $out/$merge to write results instead of streaming them back. 4) Latency is higher than a live Atlas query due to S3 I/O — not suitable for user-facing, real-time queries.
Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: cross-source aggregation pipelines use the same MongoDB stages against virtual collections backed by S3 or Atlas clusters, $lookup enables joining live Atlas data with S3 archives in a single pipeline, and early $project enables column pruning in Parquet files to dramatically reduce bytes scanned. Next up we explore S3 data partitioning for query performance.
よくある質問
「複数ソースにまたがる集約パイプラインの実行」レッスンは無料ですか?
はい。「複数ソースにまたがる集約パイプラインの実行」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、MongoDB Academyコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 MongoDB Academyコースには全4レッスンが含まれています。
「複数ソースにまたがる集約パイプラインの実行」で何を学びますか?
1つのクエリで、AtlasコレクションのデータとS3に保存されたJSONまたはParquetファイルを結合する集約パイプラインを記述します。 ブラウザで直接実行するハンズオンコードでMongoDB Academyを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
MongoDB Academyを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのMongoDB Academyは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン3/4です。
「複数ソースにまたがる集約パイプラインの実行」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このMongoDB Academyレッスンでコードを書いて実行できますか?
はい。すべてのMongoDB Academyレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- Atlas Data Federationとは
- S3とAtlasのソースを仮想名前空間にマッピングする
- 複数ソースにまたがる集約パイプラインの実行
- クエリ性能のためのS3データのパーティショニング