MongoDB Academy · 강의

샤드 키 선택하기: 카디널리티, 빈도, 단조성

학습자는 카디널리티, 쓰기 분산, 쿼리 대상 지정이라는 세 가지 기준으로 샤드 키 후보를 평가하고 핫 샤드 안티패턴을 피합니다.

레슨 2/413개 단계

샤드 키 선택하기: 카디널리티, 빈도, 단조성은(는) CoddyKit의 무료 MongoDB Academy 강의입니다. 이것은 4개 중 2번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 MongoDB Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. MongoDB Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Why Shard Key Choice Is Critical

The shard key is immutable once set and cannot be changed without unsharding and re-sharding the entire collection — an expensive, disruptive operation. Choosing the wrong shard key leads to hot shards, poor query routing, and wasted hardware. You must evaluate candidates against three dimensions: cardinality, frequency, and monotonicity.

Cardinality: How Many Distinct Values?

Cardinality is the number of distinct values the shard key can take. High cardinality (e.g., userId, email, orderId) is good — it gives MongoDB many possible chunk boundaries and lets the balancer distribute data finely. Low cardinality (e.g., status: 'active' | 'inactive', country with 50 values) creates jumbo chunks that cannot be split or migrated.

// HIGH cardinality — good shard key
sh.shardCollection('mydb.users', { userId: 1 })

// LOW cardinality — avoid: only 2 chunk boundaries possible
sh.shardCollection('mydb.users', { status: 1 }) // BAD

Frequency: How Evenly Distributed Are Values?

Frequency measures how many documents share each shard key value. Even high-cardinality keys can be problematic if a small number of values appear in the vast majority of documents. For example, a countryCode field might have 200 distinct values, but 90% of users are from a single country — creating a massive hot chunk that cannot be split.

// Estimate frequency distribution before choosing
db.users.aggregate([
  { $group: { _id: '$countryCode', count: { $sum: 1 } } },
  { $sort: { count: -1 } },
  { $limit: 10 }
])
// If top 1 value has 80%+ of docs, this is a bad shard key

Monotonicity: Are Values Always Increasing?

Monotonicity refers to whether shard key values always increase (or decrease) over time. Fields like createdAt timestamps and ObjectId (_id) are monotonically increasing. This is problematic because all new inserts land on the same 'max' chunk on one shard, creating a write hot spot even when data is evenly distributed historically.

// Monotonic keys cause write hot spots
// All new orders go to the shard with the latest date range
sh.shardCollection('mydb.orders', { createdAt: 1 }) // BAD for high insert rate

// Fix: use hashed sharding to spread monotonic keys
sh.shardCollection('mydb.orders', { createdAt: 'hashed' })

Ideal Shard Key Properties Summary

The ideal shard key has: High cardinality — thousands or millions of distinct values. Low frequency skew — no single value dominates. Non-monotonic distribution — values do not always increase, or you use hashed sharding. Query alignment — matches the filters in your most frequent, latency-sensitive queries for targeted routing.

Compound Shard Keys

A compound shard key combines two fields for better distribution. For example, { tenantId: 1, createdAt: 1 } distributes across tenants (high cardinality) and allows ranged queries within each tenant. The first field determines coarse distribution; the second provides fine-grained splitting. Compound keys can satisfy multi-field query predicates as targeted queries.

// Compound shard key: tenant + date
sh.shardCollection('mydb.events', { tenantId: 1, createdAt: 1 })

// This query is fully targeted (both shard key fields present)
db.events.find({ tenantId: 't123', createdAt: { $gte: ISODate('2025-01-01') } })

Hashed Shard Keys

A hashed shard key applies a hash function to the field value before mapping it to a chunk. This converts monotonic keys (like ObjectId) into randomly distributed hash values, eliminating write hot spots. The trade-off is that range queries on the field become scatter-gather, since hash values are not stored in original order.

// Hashed sharding: uniform write distribution
sh.shardCollection('mydb.events', { _id: 'hashed' })

// Range query on _id is now scatter-gather (con)
// But all inserts are evenly distributed (pro)

Zone Sharding for Geographical Distribution

MongoDB allows zone sharding where you assign shard key ranges to specific shards using tags (zones). This is useful for data residency requirements: European user data can be pinned to EU-region shards, US data to US shards. Zone sharding requires a compound shard key with a region prefix as the first component.

// Tag shards with zones
sh.addShardTag('shard01', 'EU')
sh.addShardTag('shard02', 'US')

// Assign key ranges to zones
sh.addTagRange('mydb.users',
  { region: 'EU', userId: MinKey },
  { region: 'EU', userId: MaxKey },
  'EU'
)

Evaluating Candidates: A Practical Checklist

When evaluating shard key candidates: 1) Run a cardinality check — db.col.distinct('field').length should be in the thousands or more. 2) Check frequency distribution with an aggregation. 3) Determine if the field is monotonic (timestamps, auto-increment). 4) Review your top 5 most frequent queries — does the candidate field appear in their filter?

// Quick cardinality check
db.events.distinct('userId').length   // want > 10,000+

// Frequency check — any value > 1% of docs is a risk
const total = db.events.countDocuments()
db.events.aggregate([
  { $group: { _id: '$userId', n: { $sum: 1 } } },
  { $match: { n: { $gt: total * 0.01 } } }
])

The _id Field as a Hashed Shard Key

A common and safe default for many workloads is to use { _id: 'hashed' }. MongoDB ObjectId values, while monotonic, become evenly distributed after hashing. This gives uniform write distribution out of the box. The main limitation is that any range query on _id becomes scatter-gather — but for most document-level lookup workloads this is acceptable.

// Safe default for write-heavy workloads without range queries
sh.shardCollection('mydb.messages', { _id: 'hashed' })

// Single-document lookup by _id is still targeted
// (hash is deterministic: mongos knows which shard)
db.messages.findOne({ _id: ObjectId('...') })

Shard Key Selection in Atlas

MongoDB Atlas provides a Shard Key Advisor in the Performance Advisor that analyzes your query patterns and recommends shard keys based on actual usage. It can detect monotonic keys, frequency skew, and missing indexes. Using the advisor before sharding is especially helpful for production workloads where query patterns are already established.

Quick Check

Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.

Lesson Recap

In this lesson you learned: a good shard key has high cardinality, low frequency skew, and avoids monotonic values, compound shard keys combine coverage for distribution and query targeting, and hashed sharding neutralizes monotonic key hot spots at the cost of range query efficiency. Next up we compare ranged vs hashed sharding strategies in detail.

무료로 시작

AI 튜터와 함께 JavaScript을(를) 배우세요 — 무료

브라우저에서 실제 코드를 작성하고 실행하며, 24/7 AI 튜터로부터 즉각적인 도움을 받고, 웹이나 앱에서 중단한 부분부터 계속 학습하세요.

코스
30
레슨
120

자주 묻는 질문

“샤드 키 선택하기: 카디널리티, 빈도, 단조성” 강의는 무료인가요?

네 — “샤드 키 선택하기: 카디널리티, 빈도, 단조성” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 MongoDB Academy 강의 전체를 잠금 해제할 수 있습니다. MongoDB Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“샤드 키 선택하기: 카디널리티, 빈도, 단조성”에서 뭘 배우나요?

학습자는 카디널리티, 쓰기 분산, 쿼리 대상 지정이라는 세 가지 기준으로 샤드 키 후보를 평가하고 핫 샤드 안티패턴을 피합니다. 브라우저에서 직접 실행하는 실습 코드로 MongoDB Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

MongoDB Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 MongoDB Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 2번째 강의입니다.

“샤드 키 선택하기: 카디널리티, 빈도, 단조성” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 MongoDB Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 MongoDB Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 샤딩 개념: 청크, 밸런서, 샤드 키
  2. 샤드 키 선택하기: 카디널리티, 빈도, 단조성
  3. 범위 기반 샤딩과 해시 기반 샤딩 전략
  4. 영역 샤딩: 지역에 데이터 고정하기
← MongoDB Academy(으)로 돌아가기