Requirements Analysis and Schema Design
Learners will translate application requirements into a MongoDB schema, choosing embedding vs referencing and applying design patterns appropriately.
Requirements Analysis and Schema Design is a free MongoDB Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the MongoDB Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
The Capstone: Designing a Real Application
In this capstone lesson, you apply all the knowledge from the MongoDB track to design a production-ready application from scratch. We will build a multi-vendor e-commerce platform — a domain rich enough to exercise embedding vs referencing decisions, index planning, aggregation design, and security. The process starts with requirements analysis, which drives every schema decision that follows.
Step 1: Gather Functional Requirements
Begin by listing the application's core entities and operations. For our e-commerce platform: entities — Users, Vendors, Products, Orders, Reviews, Carts; operations — browse products by category, search by keyword, place orders, process payments, track shipping, and write reviews. Each operation maps to one or more MongoDB queries, and those queries drive schema decisions.
// Requirement analysis output (pseudocode spec)
const requirements = {
reads: [
'Get product by slug (very high frequency)',
'List products by category + sort/filter (high frequency)',
'Search products by keyword (high frequency)',
'Get order history for a user (medium frequency)',
'Get order detail (medium frequency)'
],
writes: [
'Create order (medium frequency)',
'Update order status (medium frequency)',
'Add product review (low frequency)',
'Update product inventory (high frequency)'
]
}Step 2: Identify Access Patterns
Access patterns are the specific queries your application will run. Document them precisely before designing the schema — the schema should serve the queries, not the other way around. For each pattern, record: the filter fields, sort fields, projected fields, and estimated frequency. High-frequency patterns drive indexing and embedding decisions. Low-frequency patterns can tolerate joins or aggregation pipeline overhead.
// Access pattern register
const accessPatterns = [
{
name: 'Product page',
filter: { slug: 1 },
projection: 'all except internal fields',
frequency: 'very high',
decision: 'index on slug; embed top 5 reviews'
},
{
name: 'Category listing',
filter: { categoryId: 1, price: 1 },
sort: { price: 1, createdAt: -1 },
frequency: 'high',
decision: 'compound index { categoryId, price, createdAt }'
}
]Designing the Products Collection
The product document is the most frequently read document in the system. Apply the patterns learned: Computed Pattern for pre-computed stats (avgRating, reviewCount); Subset Pattern for embedding only the top 5 reviews; Extended Reference for embedding the vendor's name and logo alongside vendorId. This eliminates joins for 95% of product page renders.
// Product document schema (simplified)
{
_id: ObjectId(),
slug: 'wireless-headphones-pro',
name: 'Wireless Headphones Pro',
categoryId: ObjectId(),
vendor: {
_id: ObjectId(), // reference for updates
name: 'AudioTech Ltd', // Extended Reference
logoUrl: '...' // Extended Reference
},
price: 149.99,
stock: 234,
avgRating: 4.3, // Computed Pattern
reviewCount: 892, // Computed Pattern
topReviews: [ /* 5 most recent */ ], // Subset Pattern
tags: ['audio', 'wireless', 'headphones'],
schema_version: 1
}Designing the Orders Collection
Orders are a classic snapshot document: they capture the state of prices and addresses at purchase time, not the current state. Embed the full shipping address (not a reference to the user's current address), the product snapshot (name, price, image at purchase time), and the vendor name. This ensures orders remain accurate even if prices change or vendors update their profiles.
// Order document schema
{
_id: ObjectId(),
userId: ObjectId(),
status: 'processing', // 'pending', 'processing', 'shipped', 'delivered', 'refunded'
createdAt: new Date(),
shippingAddress: {
name: 'Alice Smith',
street: '42 Elm St',
city: 'Istanbul',
country: 'TR'
},
items: [
{
productId: ObjectId(), // reference for linking
slug: 'wireless-...',
name: 'Wireless Headphones Pro', // snapshot
price: 149.99, // price at purchase time
quantity: 1,
imageUrl: '...'
}
],
subtotal: 149.99,
tax: 27.00,
total: 176.99
}Applying the Schema Design Decision Framework
For each relationship, apply the decision checklist: How often is it accessed together? (embed if always, reference if rarely); How often does the referenced data change? (embed if rarely, reference if frequently); Will the embedded array grow without bound? (reference if yes, embed if bounded); Is the data queried independently? (separate collection if yes, embed if always accessed via parent).
// Decision table for our e-commerce schema
const decisions = [
{ entity: 'Order items', decision: 'embed', reason: 'always fetched with order; price snapshot required' },
{ entity: 'Shipping address', decision: 'embed', reason: 'snapshot at purchase time; changes do not affect order' },
{ entity: 'Product reviews', decision: 'separate + subset', reason: 'grows unbounded; top-5 subset in product doc' },
{ entity: 'Vendor details', decision: 'extended ref', reason: 'name/logo read on every product page; changes rarely' },
{ entity: 'Category tree', decision: 'separate', reason: 'queried independently; used for breadcrumbs' }
]Schema Validation for Critical Collections
Add JSON Schema validators to the products and orders collections to prevent malformed documents from corrupting your data. Enforce required fields (price must be a positive number, status must be one of the valid enum values) and set validationAction: 'error' to reject invalid writes immediately rather than warn.
db.runCommand({
collMod: 'orders',
validator: {
$jsonSchema: {
bsonType: 'object',
required: ['userId', 'status', 'items', 'total', 'createdAt'],
properties: {
status: {
bsonType: 'string',
enum: ['pending', 'processing', 'shipped', 'delivered', 'refunded']
},
total: { bsonType: 'double', minimum: 0 },
items: { bsonType: 'array', minItems: 1 }
}
}
},
validationAction: 'error'
})Planning the Index Set
Define indexes for every high-frequency access pattern. Use the ESR rule (Equality → Sort → Range) for compound indexes. Create indexes only for queries with high frequency — unused indexes waste write performance and RAM. Document each index with its purpose so the team can prune redundant ones as access patterns evolve.
// Products collection indexes
db.products.createIndex({ slug: 1 }, { unique: true }) // product page lookup
db.products.createIndex({ categoryId: 1, price: 1, _id: 1 }) // category listing + keyset pagination
db.products.createIndex({ tags: 1 }) // tag filter
db.products.createIndex({ 'vendor._id': 1 }) // vendor store page
// Orders collection indexes
db.orders.createIndex({ userId: 1, createdAt: -1 }) // user order history
db.orders.createIndex({ status: 1, createdAt: 1 }) // fulfillment queueHandling Concurrent Inventory Updates
A critical challenge in e-commerce is preventing overselling: two users should not both be able to purchase the last item in stock. Use MongoDB's atomic findOneAndUpdate with a stock: { $gt: 0 } guard condition. The update only succeeds if stock is available, and the decrement is atomic — no race condition possible.
// Atomically reserve stock — returns null if out of stock
const product = await db.collection('products').findOneAndUpdate(
{ _id: productId, stock: { $gte: quantity } }, // guard: enough stock
{ $inc: { stock: -quantity } },
{ returnDocument: 'after', projection: { stock: 1, name: 1, price: 1 } }
)
if (!product) {
throw new Error('Insufficient stock')
}
// Proceed to create order with product snapshotAggregation Pipeline for Reporting
Design a sales summary aggregation pipeline that reports revenue by vendor for the last 30 days. This is a classic case where the aggregation pipeline replaces complex application-side computation. The pipeline matches recent orders, unwinds items, groups by vendor, and sorts by total revenue.
// Revenue by vendor, last 30 days
const since = new Date(Date.now() - 30 * 86400000)
db.orders.aggregate([
{ $match: { status: 'delivered', createdAt: { $gte: since } } },
{ $unwind: '$items' },
{
$group: {
_id: '$items.vendorId',
totalRevenue: { $sum: { $multiply: ['$items.price', '$items.quantity'] } },
orderCount: { $addToSet: '$_id' }
}
},
{ $addFields: { orderCount: { $size: '$orderCount' } } },
{ $sort: { totalRevenue: -1 } },
{ $limit: 20 }
])Testing the Schema With explain()
Before going to production, validate every critical query with explain('executionStats'). Confirm that all high-frequency queries show IXSCAN (not COLLSCAN) in the winning plan, docsExamined is close to nReturned, and totalKeysExamined is reasonable. Any query showing COLLSCAN or a high docsExamined/nReturned ratio needs an index.
// Validate the category listing query
const stats = db.products.find(
{ categoryId: ObjectId('...'), price: { $lte: 200 } }
).sort({ price: 1 }).explain('executionStats')
const plan = stats.executionStats
console.log('Stage:', plan.executionStages.inputStage.stage) // should be IXSCAN
console.log('Keys examined:', plan.totalKeysExamined) // should be small
console.log('Docs examined:', plan.totalDocsExamined) // should equal nReturned
console.log('Docs returned:', plan.nReturned)Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: requirements analysis and access pattern documentation come before schema design — the schema serves the queries, not the other way around, a good product document combines Extended Reference, Computed Pattern, and Subset Pattern to eliminate joins on the hot read path, and atomic findOneAndUpdate with guard conditions prevents overselling without requiring transactions. Next up we define the index strategy and validate each index with explain().
Frequently asked questions
Is the “Requirements Analysis and Schema Design” lesson free?
Yes — the full text of “Requirements Analysis and Schema Design” is free to read here on the web, and the MongoDB Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the MongoDB Academy course, upgrade to CoddyKit PRO.
What will I learn in “Requirements Analysis and Schema Design”?
Learners will translate application requirements into a MongoDB schema, choosing embedding vs referencing and applying design patterns appropriately. You practise MongoDB Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start MongoDB Academy?
No prior experience is required. MongoDB Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Requirements Analysis and Schema Design” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this MongoDB Academy lesson?
Yes. Every MongoDB Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Requirements Analysis and Schema Design
- Index Strategy and Query Planner Validation
- Scaling Plan: Replica Set to Sharded Cluster
- Security Hardening and Production Checklist