什么是 Atlas Data Federation
您将描述 Data Federation 的架构、支持的数据源类型,以及统一这些数据源的查询引擎。
什么是 Atlas Data Federation 是 CoddyKit 上的免费 MongoDB Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 MongoDB Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 MongoDB Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
The Problem: Data Lives Everywhere
Modern applications generate data across multiple systems: live operational data in MongoDB Atlas, historical archives in Amazon S3, and analytics exports in data lakes. Querying across these silos traditionally requires data pipelines, ETL jobs, and separate query engines. Atlas Data Federation solves this by letting you query all these sources with a single MongoDB connection and the familiar aggregation pipeline.
What Is Atlas Data Federation?
Atlas Data Federation is a fully managed query engine built into MongoDB Atlas. It creates a federated database instance — a virtual MongoDB deployment that maps data from multiple sources (Atlas clusters, S3 buckets, Atlas Data Lake, HTTP endpoints) to virtual collections. You connect with a standard MongoDB connection string and use the same aggregation pipeline you already know.
// Connect to a federated database instance
// The URI looks like a regular Atlas connection string
// mongodb://...@data.mongodb-api.com/federated
const client = new MongoClient(
'mongodb+srv://federated-instance.mongodb.net/myFederatedDB'
)Supported Data Sources
Atlas Data Federation can query: Atlas clusters — live MongoDB collections. Amazon S3 — JSON, BSON, CSV, TSV, Avro, ORC, and Parquet files stored in S3 buckets. Atlas Data Lake — processed and enriched datasets. HTTP/HTTPS endpoints — external REST APIs that return JSON. All sources are mapped to virtual namespaces within the federated instance.
The Federated Database Architecture
A federated database has three layers: Storage configuration — defines which sources map to which virtual databases and collections. Query engine — the distributed SQL/MQL processor that reads from multiple sources, applies pipeline stages, and merges results. Connection layer — a mongos-compatible endpoint you connect to with any MongoDB driver or mongosh.
// Storage configuration (simplified JSON structure)
{
'databases': [{
'name': 'analytics',
'collections': [{
'name': 'orders_archive',
'dataSources': [{
'storeName': 's3Store',
'path': '/data/orders/2024/'
}]
}]
}]
}Creating a Federated Database Instance
You create a federated database instance through the Atlas UI, Atlas Admin API, or Atlas CLI. During setup you: 1) Name the instance. 2) Add stores (S3 buckets with IAM credentials, Atlas clusters, etc.). 3) Define virtual databases and collections that point to those stores. 4) Copy the connection string and connect with your MongoDB driver.
// Using Atlas CLI to create a data federation instance
// atlas dataFederation create myFederation --region US_EAST_1
// Then add a store via Atlas UI or API
// POST /api/atlas/v1.0/groups/{groupId}/dataFederation/{name}/dataStores
// { 'name': 's3Store', 'provider': 'S3', 'region': 'us-east-1', 'bucket': 'my-data' }Virtual Namespaces: Collections Without Schemas
Virtual collections in a federated database do not store data — they are logical views over the underlying source files or collections. You can query a virtual collection named analytics.orders that actually reads S3 Parquet files at s3://my-bucket/orders/2024/. To MongoDB drivers and tools, the virtual collection looks and behaves like a regular MongoDB collection.
// Query a virtual collection backed by S3 files
const ordersArchive = db.collection('orders_archive')
const result = await ordersArchive.aggregate([
{ $match: { year: 2024, region: 'EU' } },
{ $group: { _id: '$category', total: { $sum: '$revenue' } } },
{ $sort: { total: -1 } }
]).toArray()Cross-Source Joins With $lookup
One of the most powerful features is joining a live Atlas collection with archived S3 data in a single pipeline. For example: look up active customer details from a live Atlas cluster and join them with their 3-year purchase history stored in S3 Parquet files — all in one aggregation with no ETL job required.
// Join live Atlas collection with S3 archive
db.customers.aggregate([
{ $match: { tier: 'platinum' } }, // live Atlas
{ $lookup: {
from: 'orders_archive', // virtual S3-backed collection
localField: '_id',
foreignField: 'customerId',
as: 'purchaseHistory'
}},
{ $project: { name: 1, tier: 1,
totalOrders: { $size: '$purchaseHistory' } } }
])File Format Support in S3
Data Federation reads S3 files in many formats: JSON (one document per line or array), BSON (MongoDB native binary), CSV/TSV (with header row), Avro, ORC, and Parquet (columnar formats widely used in data lakes). For columnar formats, Data Federation can push projection and filter predicates into the file reader for even faster scans.
// In storage config, specify file format per path
{
'dataSources': [{
'storeName': 's3Store',
'path': '/analytics/events/{year string}/{month string}/',
'defaultFormat': '.parquet'
}]
}Cost Model: Query-Based Pricing
Atlas Data Federation charges based on data processed (bytes scanned), not uptime. This makes it cost-effective for infrequent analytical queries over large S3 archives — you pay nothing when no queries run. However, scanning entire unpartitioned S3 datasets can become expensive. Partitioning your S3 data and using projection to reduce scanned bytes are critical for cost control.
Security: Auth and Network
Federated database instances use the same Atlas database users and roles as regular Atlas clusters. You can apply Atlas network peering, private endpoints (AWS PrivateLink), and IP Access Lists to restrict who can connect. The connection to S3 uses IAM roles rather than storing AWS keys directly, following AWS security best practices.
When to Use Atlas Data Federation
Data Federation is a good fit when: 1) You need to run ad-hoc queries across historical S3 archives without loading data into a live cluster. 2) You want to join live transactional data with archived data in a single query. 3) You need a unified analytics interface across multiple Atlas clusters. 4) You want to avoid building and maintaining a separate ETL pipeline for each analytical use case.
Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: Atlas Data Federation creates a virtual MongoDB namespace over S3, Atlas clusters, and other sources, you use the standard aggregation pipeline to query and join data across all sources in one operation, and costs are based on bytes scanned, so partitioning and projection are essential for cost control. Next up we map S3 and Atlas sources to virtual namespaces.
常见问题解答
「什么是 Atlas Data Federation」课时是免费的吗?
是的 — 「什么是 Atlas Data Federation」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 MongoDB Academy 课程的其余内容,请升级到 CoddyKit PRO。 MongoDB Academy 课程共包含 4 节课。
「什么是 Atlas Data Federation」这节课中我会学到什么?
您将描述 Data Federation 的架构、支持的数据源类型,以及统一这些数据源的查询引擎。 你通过在浏览器中直接运行的动手代码来练习 MongoDB Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 MongoDB Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 MongoDB Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「什么是 Atlas Data Federation」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 MongoDB Academy 课中编写并运行代码吗?
能。每节 MongoDB Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 什么是 Atlas Data Federation
- 将 S3 和 Atlas 数据源映射到虚拟命名空间
- 运行跨数据源聚合管道
- 为查询性能划分 S3 数据