Sharding Concepts: Chunks, Balancer, and Shard Keys
Learners will define chunks and explain how the balancer moves them across shards to maintain an even distribution of data.
Sharding Concepts: Chunks, Balancer, and Shard Keys is a free MongoDB Academy lesson on CoddyKit — lesson 1 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the MongoDB Academy learning path, one of 4 lessons in the course, and your progress syncs across the web and the CoddyKit app.
What Is Sharding?
Sharding is MongoDB's horizontal scaling strategy. Instead of storing all data on a single replica set, a sharded cluster divides the data across multiple shards, each of which is itself a replica set. This lets you scale storage, throughput, and memory horizontally by adding more shards as your data grows.
Sharded Cluster Components
A sharded cluster has three roles: Shards — replica sets that store the actual data. mongos — a query router that receives client connections and routes operations to the correct shard(s). Config servers — a replica set that stores cluster metadata, including the shard key ranges and chunk locations. Clients always connect to mongos, never directly to shards.
// Connect to the cluster via mongos (same URI format as a standalone)
const client = new MongoClient('mongodb://mongos-host:27017/mydb')The Shard Key: Data Distribution Axis
The shard key is a field (or compound of fields) you choose when sharding a collection. MongoDB uses the shard key value to determine which shard a document belongs to. Every document in a sharded collection must contain the shard key, and the key is immutable — you cannot change it after sharding. Choosing the wrong shard key is the most common MongoDB scaling mistake.
// Enable sharding on a database
sh.enableSharding('mydb')
// Shard a collection by userId
sh.shardCollection('mydb.events', { userId: 1 })What Is a Chunk?
MongoDB divides the shard key range into chunks — contiguous ranges of shard key values. Each chunk lives on exactly one shard. By default, a chunk grows up to 128 MB before MongoDB splits it into two smaller chunks. The shard each chunk lives on is tracked in the config server metadata.
// View the chunks for a collection
use config
db.chunks.find(
{ ns: 'mydb.events' },
{ shard: 1, min: 1, max: 1 }
).limit(10)The Balancer: Evening Out Data
The balancer is a background process that monitors chunk distribution across shards. If one shard has significantly more chunks than others (by default, a difference of 8 or more), the balancer migrates chunks from the busiest shard to less-loaded shards. Migrations happen automatically and try to avoid peak traffic windows.
// Check if the balancer is running
sh.getBalancerState()
// Check balancer status
sh.status()
// Pause the balancer during a maintenance window
sh.stopBalancer()Targeted vs Scatter-Gather Queries
When a query includes the shard key, mongos routes it directly to the one (or few) shards that hold matching documents — a targeted query. When the query does not include the shard key, mongos must send it to all shards and merge results — a scatter-gather query. Targeted queries are dramatically faster. Design your queries around the shard key for best performance.
// Targeted: mongos routes to one shard
db.events.find({ userId: 'u123', date: { $gt: ISODate('2025-01-01') } })
// Scatter-gather: mongos fans out to all shards
db.events.find({ eventType: 'click' })Hot Shards: The Anti-Pattern
A hot shard (or hot spot) occurs when a disproportionate share of reads or writes land on a single shard. Common causes: using a monotonically increasing shard key (like createdAt or ObjectId) so all new inserts always go to the highest chunk, or using a low-cardinality key (like a boolean) that limits parallelism. Hot shards negate the benefits of sharding.
// BAD: ObjectId is monotonically increasing
// All new inserts go to the 'max' chunk on one shard
sh.shardCollection('mydb.events', { _id: 1 }) // AVOID
// BETTER: Use hashed sharding to distribute new inserts
sh.shardCollection('mydb.events', { _id: 'hashed' })Jump Consistent and Jumbo Chunks
A jumbo chunk is a chunk that has grown beyond the maximum size but cannot be split because all of its documents share the same shard key value. Jumbo chunks cannot be migrated by the balancer and create a persistent imbalance. The fix is to choose a shard key with sufficient cardinality so that no single key value maps to more documents than fit in one chunk.
// Identify jumbo chunks (jumbo: true in chunk metadata)
use config
db.chunks.find({ ns: 'mydb.events', jumbo: true }).count()Viewing Shard Distribution With sh.status()
sh.status() gives a complete picture of your sharded cluster: which collections are sharded, how many chunks exist per shard, and the key ranges each shard owns. Use it to quickly spot imbalanced chunk distribution and confirm that the balancer has completed migrations after adding a new shard.
// Full cluster status
sh.status()
// Targeted collection info
db.runCommand({ collStats: 'events' })
// Look at 'sharded', 'count', 'nchunks', 'shards' fieldsAdding a New Shard
Adding a shard to a running cluster is an online operation. The balancer automatically migrates chunks from existing shards to the new one over time. You can pre-split chunks before adding a shard to speed up initial distribution, especially when bulk-loading data. No downtime is required when adding shards.
// Add a new shard (replica set format)
sh.addShard('rs1/new-shard-host:27017')
// Monitor migration progress
sh.status()
db.adminCommand({ balancerCollectionStatus: 'mydb.events' })When to Shard: The Decision Threshold
Sharding adds operational complexity. Before sharding, exhaust vertical scaling and indexing options. Common triggers to shard: data volume > 1–2 TB on a single replica set, write throughput that saturates a single primary, or working set that no longer fits in RAM. Always benchmark with explain() and profiling before deciding to shard.
Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: sharding distributes data across multiple shards (replica sets) using a shard key, chunks are contiguous shard key ranges the balancer migrates to maintain even distribution, and targeted queries (including the shard key) are far more efficient than scatter-gather queries. Next up we explore how to choose the right shard key.
Frequently asked questions
Is the “Sharding Concepts: Chunks, Balancer, and Shard Keys” lesson free?
Yes — the full text of “Sharding Concepts: Chunks, Balancer, and Shard Keys” is free to read here on the web, and the MongoDB Academy course includes 4 lessons in total. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the MongoDB Academy course, upgrade to CoddyKit PRO.
What will I learn in “Sharding Concepts: Chunks, Balancer, and Shard Keys”?
Learners will define chunks and explain how the balancer moves them across shards to maintain an even distribution of data. You practise MongoDB Academy with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.
Do I need any experience to start MongoDB Academy?
No prior experience is required. MongoDB Academy on CoddyKit is structured for beginners through advanced learners; this is — lesson 1 of 4, so you can start here or from the beginning and move at your own pace.
How long does the “Sharding Concepts: Chunks, Balancer, and Shard Keys” lesson take?
Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.
Can I write and run code in this MongoDB Academy lesson?
Yes. Every MongoDB Academy lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.
All lessons in this course
- Sharding Concepts: Chunks, Balancer, and Shard Keys
- Choosing a Shard Key: Cardinality, Frequency, Monotonicity
- Ranged vs Hashed Sharding Strategies
- Zone Sharding: Pinning Data to Regions