桶模式与计算模式
学习者将把时间序列数据分组到桶文档中,以减小索引大小,并预先计算聚合值,避免实时计算带来的高昂成本。
桶模式与计算模式 是 CoddyKit 上的免费 MongoDB Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 MongoDB Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 MongoDB Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Introduction to Schema Design Patterns
Expert MongoDB engineers do not design schemas by instinct — they reach for a library of proven patterns that solve recurring challenges. These patterns encode hard-won lessons about how MongoDB's storage engine, index structure, and aggregation pipeline interact with different document shapes. The Bucket Pattern and Computed Pattern are two of the most impactful for performance-sensitive applications.
The Problem: Too Many Small Documents
Consider an IoT system where each sensor reading is a separate document. A sensor producing one reading per second generates 86,400 documents per day. This means 86,400 index entries per sensor per day, enormous index sizes, and a huge number of operations when querying a day's data. Reading 24 hours of data requires fetching tens of thousands of tiny documents — very inefficient.
// Anti-pattern: one document per reading
{
_id: ObjectId(),
sensorId: 'sensor-42',
timestamp: ISODate('2024-06-01T10:00:01Z'),
temperature: 23.1
}
// 86,400 such documents per sensor per day
// 86,400 index entries per sensor per dayThe Bucket Pattern: Grouping Into Buckets
The Bucket Pattern groups related small documents into a single larger document (a 'bucket'). Instead of one document per reading, one document holds one hour's readings — reducing 3,600 documents to 1, and 3,600 index entries to 1. The bucket document contains an array of measurements plus summary statistics computed at write time. This dramatically reduces index size and improves range query performance.
// Bucket Pattern: one document per sensor per hour
{
_id: ObjectId(),
sensorId: 'sensor-42',
date: ISODate('2024-06-01T10:00:00Z'), // hour bucket boundary
count: 3600,
measurements: [
{ ts: ISODate('2024-06-01T10:00:01Z'), temp: 23.1 },
{ ts: ISODate('2024-06-01T10:00:02Z'), temp: 23.2 },
// ... 3598 more readings ...
],
// Pre-computed summaries
avgTemp: 23.15,
maxTemp: 24.1,
minTemp: 22.8
}Adding to a Bucket With $push and $inc
When a new measurement arrives, find the current bucket document for the sensor and hour, and update it atomically using $push to append to the measurements array and $inc to increment the count. Use upsert: true so MongoDB creates a new bucket document if the current hour's bucket does not exist yet.
const now = new Date()
const hourBucket = new Date(now.getFullYear(), now.getMonth(), now.getDate(), now.getHours())
db.sensorBuckets.updateOne(
{
sensorId: 'sensor-42',
date: hourBucket,
count: { $lt: 3600 } // don't overfill a bucket
},
{
$push: { measurements: { ts: now, temp: 23.5 } },
$inc: { count: 1 },
$min: { minTemp: 23.5 },
$max: { maxTemp: 23.5 }
},
{ upsert: true }
)Querying Across Buckets
To retrieve all readings for a sensor over a time range, query the bucket documents by sensorId and date range, then either use the pre-computed summaries for aggregate results, or $unwind the measurements array for per-reading access. Querying 24 hours of data now reads 24 bucket documents instead of 86,400 individual ones — a 3,600x reduction in documents fetched.
// Get hourly summaries for a day (reads 24 bucket docs)
db.sensorBuckets.find(
{
sensorId: 'sensor-42',
date: {
$gte: ISODate('2024-06-01T00:00:00Z'),
$lt: ISODate('2024-06-02T00:00:00Z')
}
},
{ _id: 0, date: 1, avgTemp: 1, maxTemp: 1, minTemp: 1, count: 1 }
).sort({ date: 1 })The Computed Pattern: Pre-Computing Results
The Computed Pattern pre-calculates expensive aggregations at write time and stores the result directly in the document. Instead of computing an average rating for a product by summing all reviews at read time, compute it whenever a new review is added and store avgRating alongside reviewCount in the product document. Reads become trivially cheap — just return the pre-computed field.
// Product document with pre-computed stats
{
_id: ObjectId(),
name: 'Wireless Headphones',
price: 149.99,
// Pre-computed at write time
reviewCount: 1247,
avgRating: 4.3,
ratingSum: 5362.1 // stored to recompute avg without fetching all reviews
}Updating Computed Fields Incrementally
When a new review arrives, update the computed fields atomically in the same operation rather than recalculating from scratch. Increment reviewCount and ratingSum with $inc, then recompute avgRating. In MongoDB, you can do this with a pipeline-style update (MongoDB 4.2+) that uses $divide to set the average from the updated count and sum.
// Atomic incremental update of computed stats
db.products.updateOne(
{ _id: productId },
[
{
$set: {
reviewCount: { $add: ['$reviewCount', 1] },
ratingSum: { $add: ['$ratingSum', newRating] }
}
},
{
$set: {
avgRating: { $divide: ['$ratingSum', '$reviewCount'] }
}
}
]
)When to Use the Computed Pattern
The Computed Pattern is most valuable when reads vastly outnumber writes and the computation would be expensive at read time. Product rating averages, leaderboard scores, article view counts, revenue totals — all are excellent candidates. The tradeoff is slightly increased write complexity and the need to keep computed values consistent when source data changes. If both reads and writes are frequent, consider background jobs that recompute periodically rather than synchronous updates.
Combining Both Patterns
The Bucket and Computed patterns are frequently used together. A sensor data system might group readings into hourly bucket documents (Bucket Pattern) and maintain a pre-computed daily summary document (Computed Pattern) that stores min/max/avg for the entire day. This tiered approach means dashboards showing daily trends read a single document, while drill-down queries scan only 24 hourly bucket documents.
// Daily summary document (Computed Pattern on top of Bucket Pattern)
{
_id: ObjectId(),
sensorId: 'sensor-42',
date: ISODate('2024-06-01T00:00:00Z'),
totalReadings: 86400,
dailyAvgTemp: 22.8,
dailyMaxTemp: 28.4,
dailyMinTemp: 18.2,
peakHour: 14 // hour with highest average temperature
}Bucket Size Tradeoffs
Bucket documents should not grow without bound — MongoDB has a 16 MB document size limit. For sensors with high data rates, use time-bounded buckets (e.g., one per hour or one per day). The count: { $lt: 3600 } guard in the update filter prevents overfilling. If a bucket reaches capacity, the upsert creates a new bucket automatically. Monitor bucket fill rates in production to tune bucket size for your data rate.
Index Impact of the Bucket Pattern
The primary index on a bucket collection should cover the query pattern: { sensorId: 1, date: 1 }. This compound index means queries filtering by sensor and time range hit the index directly. Compared to indexing the timestamp field of 86,400 individual documents, the bucket collection's index holds only 24 entries per sensor per day — a 3,600x reduction in index size that drastically improves working-set fit in RAM.
// Create supporting index for bucket pattern queries
db.sensorBuckets.createIndex({ sensorId: 1, date: 1 })
// Explain query to verify IXSCAN usage
db.sensorBuckets.find({ sensorId: 'sensor-42', date: { $gte: ISODate('2024-06-01T00:00:00Z') } }).explain('executionStats')Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: the Bucket Pattern groups many small documents into fewer larger ones, dramatically reducing index size and improving range query performance, the Computed Pattern pre-calculates expensive aggregations at write time so reads return pre-stored values instantly, and the two patterns combine effectively for multi-tier time series pipelines. Next up we explore the Extended Reference and Subset Patterns.
常见问题解答
「桶模式与计算模式」课时是免费的吗?
是的 — 「桶模式与计算模式」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 MongoDB Academy 课程的其余内容,请升级到 CoddyKit PRO。 MongoDB Academy 课程共包含 4 节课。
「桶模式与计算模式」这节课中我会学到什么?
学习者将把时间序列数据分组到桶文档中,以减小索引大小,并预先计算聚合值,避免实时计算带来的高昂成本。 你通过在浏览器中直接运行的动手代码来练习 MongoDB Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 MongoDB Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 MongoDB Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「桶模式与计算模式」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 MongoDB Academy 课中编写并运行代码吗?
能。每节 MongoDB Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。