0Pricing
MongoDB Academy · 课时

选择分片键:基数、频率和单调性

您将从基数、写入分布和查询定位三个维度评估分片键候选项,并避免热点分片反模式。

选择分片键:基数、频率和单调性 是 CoddyKit 上的免费 MongoDB Academy 课时。 这是第 2 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 MongoDB Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 MongoDB Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

Why Shard Key Choice Is Critical

The shard key is immutable once set and cannot be changed without unsharding and re-sharding the entire collection — an expensive, disruptive operation. Choosing the wrong shard key leads to hot shards, poor query routing, and wasted hardware. You must evaluate candidates against three dimensions: cardinality, frequency, and monotonicity.

Cardinality: How Many Distinct Values?

Cardinality is the number of distinct values the shard key can take. High cardinality (e.g., userId, email, orderId) is good — it gives MongoDB many possible chunk boundaries and lets the balancer distribute data finely. Low cardinality (e.g., status: 'active' | 'inactive', country with 50 values) creates jumbo chunks that cannot be split or migrated.

// HIGH cardinality — good shard key
sh.shardCollection('mydb.users', { userId: 1 })

// LOW cardinality — avoid: only 2 chunk boundaries possible
sh.shardCollection('mydb.users', { status: 1 }) // BAD

Frequency: How Evenly Distributed Are Values?

Frequency measures how many documents share each shard key value. Even high-cardinality keys can be problematic if a small number of values appear in the vast majority of documents. For example, a countryCode field might have 200 distinct values, but 90% of users are from a single country — creating a massive hot chunk that cannot be split.

// Estimate frequency distribution before choosing
db.users.aggregate([
  { $group: { _id: '$countryCode', count: { $sum: 1 } } },
  { $sort: { count: -1 } },
  { $limit: 10 }
])
// If top 1 value has 80%+ of docs, this is a bad shard key

Monotonicity: Are Values Always Increasing?

Monotonicity refers to whether shard key values always increase (or decrease) over time. Fields like createdAt timestamps and ObjectId (_id) are monotonically increasing. This is problematic because all new inserts land on the same 'max' chunk on one shard, creating a write hot spot even when data is evenly distributed historically.

// Monotonic keys cause write hot spots
// All new orders go to the shard with the latest date range
sh.shardCollection('mydb.orders', { createdAt: 1 }) // BAD for high insert rate

// Fix: use hashed sharding to spread monotonic keys
sh.shardCollection('mydb.orders', { createdAt: 'hashed' })

Ideal Shard Key Properties Summary

The ideal shard key has: High cardinality — thousands or millions of distinct values. Low frequency skew — no single value dominates. Non-monotonic distribution — values do not always increase, or you use hashed sharding. Query alignment — matches the filters in your most frequent, latency-sensitive queries for targeted routing.

Compound Shard Keys

A compound shard key combines two fields for better distribution. For example, { tenantId: 1, createdAt: 1 } distributes across tenants (high cardinality) and allows ranged queries within each tenant. The first field determines coarse distribution; the second provides fine-grained splitting. Compound keys can satisfy multi-field query predicates as targeted queries.

// Compound shard key: tenant + date
sh.shardCollection('mydb.events', { tenantId: 1, createdAt: 1 })

// This query is fully targeted (both shard key fields present)
db.events.find({ tenantId: 't123', createdAt: { $gte: ISODate('2025-01-01') } })

Hashed Shard Keys

A hashed shard key applies a hash function to the field value before mapping it to a chunk. This converts monotonic keys (like ObjectId) into randomly distributed hash values, eliminating write hot spots. The trade-off is that range queries on the field become scatter-gather, since hash values are not stored in original order.

// Hashed sharding: uniform write distribution
sh.shardCollection('mydb.events', { _id: 'hashed' })

// Range query on _id is now scatter-gather (con)
// But all inserts are evenly distributed (pro)

Zone Sharding for Geographical Distribution

MongoDB allows zone sharding where you assign shard key ranges to specific shards using tags (zones). This is useful for data residency requirements: European user data can be pinned to EU-region shards, US data to US shards. Zone sharding requires a compound shard key with a region prefix as the first component.

// Tag shards with zones
sh.addShardTag('shard01', 'EU')
sh.addShardTag('shard02', 'US')

// Assign key ranges to zones
sh.addTagRange('mydb.users',
  { region: 'EU', userId: MinKey },
  { region: 'EU', userId: MaxKey },
  'EU'
)

Evaluating Candidates: A Practical Checklist

When evaluating shard key candidates: 1) Run a cardinality check — db.col.distinct('field').length should be in the thousands or more. 2) Check frequency distribution with an aggregation. 3) Determine if the field is monotonic (timestamps, auto-increment). 4) Review your top 5 most frequent queries — does the candidate field appear in their filter?

// Quick cardinality check
db.events.distinct('userId').length   // want > 10,000+

// Frequency check — any value > 1% of docs is a risk
const total = db.events.countDocuments()
db.events.aggregate([
  { $group: { _id: '$userId', n: { $sum: 1 } } },
  { $match: { n: { $gt: total * 0.01 } } }
])

The _id Field as a Hashed Shard Key

A common and safe default for many workloads is to use { _id: 'hashed' }. MongoDB ObjectId values, while monotonic, become evenly distributed after hashing. This gives uniform write distribution out of the box. The main limitation is that any range query on _id becomes scatter-gather — but for most document-level lookup workloads this is acceptable.

// Safe default for write-heavy workloads without range queries
sh.shardCollection('mydb.messages', { _id: 'hashed' })

// Single-document lookup by _id is still targeted
// (hash is deterministic: mongos knows which shard)
db.messages.findOne({ _id: ObjectId('...') })

Shard Key Selection in Atlas

MongoDB Atlas provides a Shard Key Advisor in the Performance Advisor that analyzes your query patterns and recommends shard keys based on actual usage. It can detect monotonic keys, frequency skew, and missing indexes. Using the advisor before sharding is especially helpful for production workloads where query patterns are already established.

Quick Check

Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.

Lesson Recap

In this lesson you learned: a good shard key has high cardinality, low frequency skew, and avoids monotonic values, compound shard keys combine coverage for distribution and query targeting, and hashed sharding neutralizes monotonic key hot spots at the cost of range query efficiency. Next up we compare ranged vs hashed sharding strategies in detail.

常见问题解答

「选择分片键:基数、频率和单调性」课时是免费的吗?

是的 — 「选择分片键:基数、频率和单调性」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 MongoDB Academy 课程的其余内容,请升级到 CoddyKit PRO。 MongoDB Academy 课程共包含 4 节课。

「选择分片键:基数、频率和单调性」这节课中我会学到什么?

您将从基数、写入分布和查询定位三个维度评估分片键候选项,并避免热点分片反模式。 你通过在浏览器中直接运行的动手代码来练习 MongoDB Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 MongoDB Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 MongoDB Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 2 节课,共 4 节。

「选择分片键:基数、频率和单调性」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 MongoDB Academy 课中编写并运行代码吗?

能。每节 MongoDB Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 分片概念:块、均衡器和分片键
  2. 选择分片键:基数、频率和单调性
  3. 范围分片与哈希分片策略
  4. 区域分片:将数据固定到区域
← 返回 MongoDB Academy