分片概念:块、均衡器和分片键
您将定义块,并说明均衡器如何在各分片之间移动它们,以保持数据均匀分布。
分片概念:块、均衡器和分片键 是 CoddyKit 上的免费 MongoDB Academy 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 MongoDB Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 MongoDB Academy 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
What Is Sharding?
Sharding is MongoDB's horizontal scaling strategy. Instead of storing all data on a single replica set, a sharded cluster divides the data across multiple shards, each of which is itself a replica set. This lets you scale storage, throughput, and memory horizontally by adding more shards as your data grows.
Sharded Cluster Components
A sharded cluster has three roles: Shards — replica sets that store the actual data. mongos — a query router that receives client connections and routes operations to the correct shard(s). Config servers — a replica set that stores cluster metadata, including the shard key ranges and chunk locations. Clients always connect to mongos, never directly to shards.
// Connect to the cluster via mongos (same URI format as a standalone)
const client = new MongoClient('mongodb://mongos-host:27017/mydb')The Shard Key: Data Distribution Axis
The shard key is a field (or compound of fields) you choose when sharding a collection. MongoDB uses the shard key value to determine which shard a document belongs to. Every document in a sharded collection must contain the shard key, and the key is immutable — you cannot change it after sharding. Choosing the wrong shard key is the most common MongoDB scaling mistake.
// Enable sharding on a database
sh.enableSharding('mydb')
// Shard a collection by userId
sh.shardCollection('mydb.events', { userId: 1 })What Is a Chunk?
MongoDB divides the shard key range into chunks — contiguous ranges of shard key values. Each chunk lives on exactly one shard. By default, a chunk grows up to 128 MB before MongoDB splits it into two smaller chunks. The shard each chunk lives on is tracked in the config server metadata.
// View the chunks for a collection
use config
db.chunks.find(
{ ns: 'mydb.events' },
{ shard: 1, min: 1, max: 1 }
).limit(10)The Balancer: Evening Out Data
The balancer is a background process that monitors chunk distribution across shards. If one shard has significantly more chunks than others (by default, a difference of 8 or more), the balancer migrates chunks from the busiest shard to less-loaded shards. Migrations happen automatically and try to avoid peak traffic windows.
// Check if the balancer is running
sh.getBalancerState()
// Check balancer status
sh.status()
// Pause the balancer during a maintenance window
sh.stopBalancer()Targeted vs Scatter-Gather Queries
When a query includes the shard key, mongos routes it directly to the one (or few) shards that hold matching documents — a targeted query. When the query does not include the shard key, mongos must send it to all shards and merge results — a scatter-gather query. Targeted queries are dramatically faster. Design your queries around the shard key for best performance.
// Targeted: mongos routes to one shard
db.events.find({ userId: 'u123', date: { $gt: ISODate('2025-01-01') } })
// Scatter-gather: mongos fans out to all shards
db.events.find({ eventType: 'click' })Hot Shards: The Anti-Pattern
A hot shard (or hot spot) occurs when a disproportionate share of reads or writes land on a single shard. Common causes: using a monotonically increasing shard key (like createdAt or ObjectId) so all new inserts always go to the highest chunk, or using a low-cardinality key (like a boolean) that limits parallelism. Hot shards negate the benefits of sharding.
// BAD: ObjectId is monotonically increasing
// All new inserts go to the 'max' chunk on one shard
sh.shardCollection('mydb.events', { _id: 1 }) // AVOID
// BETTER: Use hashed sharding to distribute new inserts
sh.shardCollection('mydb.events', { _id: 'hashed' })Jump Consistent and Jumbo Chunks
A jumbo chunk is a chunk that has grown beyond the maximum size but cannot be split because all of its documents share the same shard key value. Jumbo chunks cannot be migrated by the balancer and create a persistent imbalance. The fix is to choose a shard key with sufficient cardinality so that no single key value maps to more documents than fit in one chunk.
// Identify jumbo chunks (jumbo: true in chunk metadata)
use config
db.chunks.find({ ns: 'mydb.events', jumbo: true }).count()Viewing Shard Distribution With sh.status()
sh.status() gives a complete picture of your sharded cluster: which collections are sharded, how many chunks exist per shard, and the key ranges each shard owns. Use it to quickly spot imbalanced chunk distribution and confirm that the balancer has completed migrations after adding a new shard.
// Full cluster status
sh.status()
// Targeted collection info
db.runCommand({ collStats: 'events' })
// Look at 'sharded', 'count', 'nchunks', 'shards' fieldsAdding a New Shard
Adding a shard to a running cluster is an online operation. The balancer automatically migrates chunks from existing shards to the new one over time. You can pre-split chunks before adding a shard to speed up initial distribution, especially when bulk-loading data. No downtime is required when adding shards.
// Add a new shard (replica set format)
sh.addShard('rs1/new-shard-host:27017')
// Monitor migration progress
sh.status()
db.adminCommand({ balancerCollectionStatus: 'mydb.events' })When to Shard: The Decision Threshold
Sharding adds operational complexity. Before sharding, exhaust vertical scaling and indexing options. Common triggers to shard: data volume > 1–2 TB on a single replica set, write throughput that saturates a single primary, or working set that no longer fits in RAM. Always benchmark with explain() and profiling before deciding to shard.
Quick Check
Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.
Lesson Recap
In this lesson you learned: sharding distributes data across multiple shards (replica sets) using a shard key, chunks are contiguous shard key ranges the balancer migrates to maintain even distribution, and targeted queries (including the shard key) are far more efficient than scatter-gather queries. Next up we explore how to choose the right shard key.
用 AI 导师学习 JavaScript — 免费
在浏览器中编写并运行真实代码,获得全天候 AI 导师的即时帮助,并在网页或应用中继续学习。
- 课程
- 30
- 课程
- 120
常见问题解答
「分片概念:块、均衡器和分片键」课时是免费的吗?
是的 — 「分片概念:块、均衡器和分片键」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 MongoDB Academy 课程的其余内容,请升级到 CoddyKit PRO。 MongoDB Academy 课程共包含 4 节课。
「分片概念:块、均衡器和分片键」这节课中我会学到什么?
您将定义块,并说明均衡器如何在各分片之间移动它们,以保持数据均匀分布。 你通过在浏览器中直接运行的动手代码来练习 MongoDB Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 MongoDB Academy 需要有经验吗?
无需任何先前经验。CoddyKit 上的 MongoDB Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「分片概念:块、均衡器和分片键」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 MongoDB Academy 课中编写并运行代码吗?
能。每节 MongoDB Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 分片概念:块、均衡器和分片键
- 选择分片键:基数、频率和单调性
- 范围分片与哈希分片策略
- 区域分片:将数据固定到区域