The Bucket and Computed Patterns
Learners will group time series data into bucket documents to reduce index size and pre-compute aggregated values to avoid expensive real-time calculations.
The Bucket and Computed Patterns is a free MongoDB Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the MongoDB Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
Introduction to Schema Design Patterns
Expert MongoDB engineers do not design schemas by instinct — they reach for a library of proven patterns that solve recurring challenges. These patterns encode hard-won lessons about how MongoDB's storage engine, index structure, and aggregation pipeline interact with different document shapes. The Bucket Pattern and Computed Pattern are two of the most impactful for performance-sensitive applications.
The Problem: Too Many Small Documents
Consider an IoT system where each sensor reading is a separate document. A sensor producing one reading per second generates 86,400 documents per day. This means 86,400 index entries per sensor per day, enormous index sizes, and a huge number of operations when querying a day's data. Reading 24 hours of data requires fetching tens of thousands of tiny documents — very inefficient.
// Anti-pattern: one document per reading
{
_id: ObjectId(),
sensorId: 'sensor-42',
timestamp: ISODate('2024-06-01T10:00:01Z'),
temperature: 23.1
}
// 86,400 such documents per sensor per day
// 86,400 index entries per sensor per dayThe Bucket Pattern: Grouping Into Buckets
The Bucket Pattern groups related small documents into a single larger document (a 'bucket'). Instead of one document per reading, one document holds one hour's readings — reducing 3,600 documents to 1, and 3,600 index entries to 1. The bucket document contains an array of measurements plus summary statistics computed at write time. This dramatically reduces index size and improves range query performance.
// Bucket Pattern: one document per sensor per hour
{
_id: ObjectId(),
sensorId: 'sensor-42',
date: ISODate('2024-06-01T10:00:00Z'), // hour bucket boundary
count: 3600,
measurements: [
{ ts: ISODate('2024-06-01T10:00:01Z'), temp: 23.1 },
{ ts: ISODate('2024-06-01T10:00:02Z'), temp: 23.2 },
// ... 3598 more readings ...
],
// Pre-computed summaries
avgTemp: 23.15,
maxTemp: 24.1,
minTemp: 22.8
}Adding to a Bucket With $push and $inc
When a new measurement arrives, find the current bucket document for the sensor and hour, and update it atomically using $push to append to the measurements array and $inc to increment the count. Use upsert: true so MongoDB creates a new bucket document if the current hour's bucket does not exist yet.
const now = new Date()
const hourBucket = new Date(now.getFullYear(), now.getMonth(), now.getDate(), now.getHours())
db.sensorBuckets.updateOne(
{
sensorId: 'sensor-42',
date: hourBucket,
count: { $lt: 3600 } // don't overfill a bucket
},
{
$push: { measurements: { ts: now, temp: 23.5 } },
$inc: { count: 1 },
$min: { minTemp: 23.5 },
$max: { maxTemp: 23.5 }
},
{ upsert: true }
)Querying Across Buckets
To retrieve all readings for a sensor over a time range, query the bucket documents by sensorId and date range, then either use the pre-computed summaries for aggregate results, or $unwind the measurements array for per-reading access. Querying 24 hours of data now reads 24 bucket documents instead of 86,400 individual ones — a 3,600x reduction in documents fetched.
// Get hourly summaries for a day (reads 24 bucket docs)
db.sensorBuckets.find(
{
sensorId: 'sensor-42',
date: {
$gte: ISODate('2024-06-01T00:00:00Z'),
$lt: ISODate('2024-06-02T00:00:00Z')
}
},
{ _id: 0, date: 1, avgTemp: 1, maxTemp: 1, minTemp: 1, count: 1 }
).sort({ date: 1 })The Computed Pattern: Pre-Computing Results
The Computed Pattern pre-calculates expensive aggregations at write time and stores the result directly in the document. Instead of computing an average rating for a product by summing all reviews at read time, compute it whenever a new review is added and store avgRating alongside reviewCount in the product document. Reads become trivially cheap — just return the pre-computed field.
// Product document with pre-computed stats
{
_id: ObjectId(),
name: 'Wireless Headphones',
price: 149.99,
// Pre-computed at write time
reviewCount: 1247,
avgRating: 4.3,
ratingSum: 5362.1 // stored to recompute avg without fetching all reviews
}Updating Computed Fields Incrementally
When a new review arrives, update the computed fields atomically in the same operation rather than recalculating from scratch. Increment reviewCount and ratingSum with $inc, then recompute avgRating. In MongoDB, you can do this with a pipeline-style update (MongoDB 4.2+) that uses $divide to set the average from the updated count and sum.
// Atomic incremental update of computed stats
db.products.updateOne(
{ _id: productId },
[
{
$set: {
reviewCount: { $add: ['$reviewCount', 1] },
ratingSum: { $add: ['$ratingSum', newRating] }
}
},
{
$set: {
avgRating: { $divide: ['$ratingSum', '$reviewCount'] }
}
}
]
)When to Use the Computed Pattern
The Computed Pattern is most valuable when reads vastly outnumber writes and the computation would be expensive at read time. Product rating averages, leaderboard scores, article view counts, revenue totals — all are excellent candidates. The tradeoff is slightly increased write complexity and the need to keep computed values consistent when source data changes. If both reads and writes are frequent, consider background jobs that recompute periodically rather than synchronous updates.
Combining Both Patterns
The Bucket and Computed patterns are frequently used together. A sensor data system might group readings into hourly bucket documents (Bucket Pattern) and maintain a pre-computed daily summary document (Computed Pattern) that stores min/max/avg for the entire day. This tiered approach means dashboards showing daily trends read a single document, while drill-down queries scan only 24 hourly bucket documents.
// Daily summary document (Computed Pattern on top of Bucket Pattern)
{
_id: ObjectId(),
sensorId: 'sensor-42',
date: ISODate('2024-06-01T00:00:00Z'),
totalReadings: 86400,
dailyAvgTemp: 22.8,
dailyMaxTemp: 28.4,
dailyMinTemp: 18.2,
peakHour: 14 // hour with highest average temperature
}Bucket Size Tradeoffs
Bucket documents should not grow without bound — MongoDB has a 16 MB document size limit. For sensors with high data rates, use time-bounded buckets (e.g., one per hour or one per day). The count: { $lt: 3600 } guard in the update filter prevents overfilling. If a bucket reaches capacity, the upsert creates a new bucket automatically. Monitor bucket fill rates in production to tune bucket size for your data rate.
Index Impact of the Bucket Pattern
The primary index on a bucket collection should cover the query pattern: { sensorId: 1, date: 1 }. This compound index means queries filtering by sensor and time range hit the index directly. Compared to indexing the timestamp field of 86,400 individual documents, the bucket collection's index holds only 24 entries per sensor per day — a 3,600x reduction in index size that drastically improves working-set fit in RAM.
// Create supporting index for bucket pattern queries
db.sensorBuckets.createIndex({ sensorId: 1, date: 1 })
// Explain query to verify IXSCAN usage
db.sensorBuckets.find({ sensorId: 'sensor-42', date: { $gte: ISODate('2024-06-01T00:00:00Z') } }).explain('executionStats')Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: the Bucket Pattern groups many small documents into fewer larger ones, dramatically reducing index size and improving range query performance, the Computed Pattern pre-calculates expensive aggregations at write time so reads return pre-stored values instantly, and the two patterns combine effectively for multi-tier time series pipelines. Next up we explore the Extended Reference and Subset Patterns.
Frequently asked questions
Is the “The Bucket and Computed Patterns” lesson free?
Yes — the full text of “The Bucket and Computed Patterns” is free to read here on the web, and the MongoDB Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the MongoDB Academy course, upgrade to CoddyKit PRO.
What will I learn in “The Bucket and Computed Patterns”?
Learners will group time series data into bucket documents to reduce index size and pre-compute aggregated values to avoid expensive real-time calculations. You practise MongoDB Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start MongoDB Academy?
No prior experience is required. MongoDB Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “The Bucket and Computed Patterns” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this MongoDB Academy lesson?
Yes. Every MongoDB Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- The Bucket and Computed Patterns
- The Extended Reference and Subset Patterns
- The Polymorphic and Schema Versioning Patterns
- The Outlier and Tree Structure Patterns