0Pricing
MongoDB Academy · 강의

범위 기반 샤딩과 해시 기반 샤딩 전략

학습자는 범위 쿼리에는 범위 기반 샤딩을, 균등한 쓰기 분산에는 해시 기반 샤딩을 사용하도록 컬렉션을 구성하고 두 방식의 절충점을 비교합니다.

범위 기반 샤딩과 해시 기반 샤딩 전략은(는) CoddyKit의 무료 MongoDB Academy 강의입니다. 이것은 4개 중 3번째 강의입니다. 아래에서 전체 강의를 무료로 읽을 수 있으며, 내장 코드 에디터와 24/7 AI 튜터와 함께 브라우저에서 직접 실습할 수 있습니다. 이 강의는 MongoDB Academy 학습 경로의 일부이며, 진행 상황이 웹과 CoddyKit 앱에 동기화됩니다. MongoDB Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

이 강의의 일부는 아직 번역되지 않았으며 영어로 표시됩니다.

Two Sharding Strategies Compared

MongoDB supports two built-in sharding strategies: ranged sharding and hashed sharding. Ranged sharding assigns contiguous ranges of shard key values to specific shards; hashed sharding applies a hash function to the key first and distributes based on the hash. Both have distinct strengths, and the right choice depends on your data access patterns.

Ranged Sharding: How It Works

In ranged sharding, MongoDB divides the shard key's value space into contiguous ranges and assigns each range (chunk) to a shard. For example, users with userId 1–10,000 go to shard A, 10,001–20,000 to shard B, and so on. Documents with nearby shard key values are co-located on the same shard, which is ideal for range queries.

// Enable ranged sharding on a field
sh.shardCollection('mydb.products', { category: 1, price: 1 })

// Range query is now targeted to the shard(s) holding that range
db.products.find({ category: 'electronics', price: { $lt: 100 } })

Ranged Sharding: Strengths

Ranged sharding excels when your application frequently queries ranges of values: date ranges, price ranges, alphabetical name ranges, or paginated results sorted by a numeric ID. Because adjacent values are co-located, range queries become targeted queries that touch only one or a few shards, keeping latency low.

// With ranged sharding on { orderId: 1 }, this is targeted:
db.orders.find({
  orderId: { $gte: 50000, $lte: 60000 }
})
// mongos knows exactly which shard owns this range

Ranged Sharding: Weakness — Hot Spots

The critical weakness of ranged sharding is write hot spots when the shard key is monotonically increasing (timestamps, auto-increment IDs, ObjectId). All new documents cluster at the high end of the range and land on one shard. Until the balancer migrates chunks, this shard absorbs all write traffic while other shards sit idle.

// Problematic: all new events go to the max-range shard
sh.shardCollection('mydb.events', { createdAt: 1 }) // ranged, monotonic = hot spot

// Production symptom: one shard has 90%+ of recent data
// and absorbs all write IOPS

Hashed Sharding: How It Works

In hashed sharding, MongoDB computes a hash of the shard key value and uses the hash to determine the chunk. Documents are distributed based on hash values, which appear random even if the original keys are monotonically increasing. This guarantees a near-uniform initial write distribution across all shards.

// Hashed sharding on _id (neutralizes ObjectId monotonicity)
sh.shardCollection('mydb.events', { _id: 'hashed' })

// Hashed sharding on userId
sh.shardCollection('mydb.sessions', { userId: 'hashed' })

Hashed Sharding: Strengths

Hashed sharding is the best choice when your primary goal is uniform write distribution across shards and you do not need range queries on the shard key. It is ideal for high-insert-rate workloads with monotonic keys (event logs, IoT sensor data, messaging) where every shard should absorb an equal share of write traffic.

// With hashed sharding, inserts are spread uniformly:
// doc1 (hash: 2345...) -> shard A
// doc2 (hash: 8901...) -> shard C
// doc3 (hash: 4567...) -> shard B
// No hot spot regardless of insert order

Hashed Sharding: Weakness — No Range Efficiency

The trade-off of hashed sharding is that range queries on the shard key become scatter-gather. Because adjacent hash values are scattered across shards, a query like { createdAt: { $gte: t1, $lte: t2 } } must fan out to all shards. If range queries are frequent and latency-sensitive, hashed sharding may negate the performance gains from sharding.

// With hashed sharding on createdAt:
// Range query CANNOT be targeted — fans out to all shards
db.events.find({ createdAt: { $gte: ISODate('2025-01-01'), $lte: ISODate('2025-02-01') } })
// Equivalent to a full collection scan across all shards

Choosing Between Ranged and Hashed

Decision guide: Ranged sharding → your shard key has natural distribution (not monotonic) AND your top queries are range queries on that key. Hashed sharding → your shard key is monotonic OR your top queries are point lookups (equality) on a high-cardinality field. When in doubt and inserts are the bottleneck, prefer hashed.

Hybrid: Compound Key With Hashed Component

You can combine both strategies with a compound shard key where the first field is ranged and gives query affinity, while adding a hashed second field spreads the load within each range. Example: { tenantId: 1, _id: 'hashed' } co-locates data by tenant for targeted queries while distributing writes across shards within each tenant.

// Ranged tenantId + hashed _id within tenant
// Writes are distributed; per-tenant queries are targeted
sh.shardCollection('mydb.events', { tenantId: 1, _id: 'hashed' })

// Targeted: all tenantId queries go to the right shard(s)
db.events.find({ tenantId: 'acme', _id: ObjectId('...') })

Checking Which Strategy Is Active

You can inspect a collection's sharding configuration to determine which strategy is in use. The config server metadata stores the shard key and whether it uses 'hashed'. The sh.status() command and db.collection.stats() both expose this information.

// Check sharding info for a collection
use config
db.collections.findOne({ _id: 'mydb.events' })
// { key: { _id: 'hashed' }, unique: false, ... }

// Or via sh.status()
sh.status()

Pre-Splitting Chunks for Bulk Loads

When bulk-loading data into a freshly sharded collection, all chunks initially live on one shard and must be migrated by the balancer — which can be slow. Pre-splitting creates an initial set of empty chunks distributed across all shards before loading. This ensures the balancer's work is minimised and writes are spread from the first insert.

// Pre-split chunks for ranged sharding
// Define desired split points and assign to shards
db.adminCommand({ split: 'mydb.events', middle: { userId: 500000 } })
db.adminCommand({ moveChunk: 'mydb.events',
  find: { userId: 500000 }, to: 'shard02' })

Quick Check

Test your understanding of MongoDB & NoSQL Databases concepts from this lesson.

Lesson Recap

In this lesson you learned: ranged sharding co-locates similar key values for efficient range queries but creates hot spots with monotonic keys, hashed sharding distributes uniformly across shards but makes range queries scatter-gather, and compound keys can combine both benefits. Next up we explore zone sharding for pinning data to specific regions.

자주 묻는 질문

“범위 기반 샤딩과 해시 기반 샤딩 전략” 강의는 무료인가요?

네 — “범위 기반 샤딩과 해시 기반 샤딩 전략” 전체 내용을 이 웹사이트에서 무료로 읽을 수 있습니다. 인터랙티브하게 실습하려면(내장 코드 에디터와 24/7 AI 튜터), CoddyKit PRO로 업그레이드하면 MongoDB Academy 강의 전체를 잠금 해제할 수 있습니다. MongoDB Academy 강의에는 총 4개의 강의가 포함되어 있습니다.

“범위 기반 샤딩과 해시 기반 샤딩 전략”에서 뭘 배우나요?

학습자는 범위 쿼리에는 범위 기반 샤딩을, 균등한 쓰기 분산에는 해시 기반 샤딩을 사용하도록 컬렉션을 구성하고 두 방식의 절충점을 비교합니다. 브라우저에서 직접 실행하는 실습 코드로 MongoDB Academy을(를) 배우며, 24/7 AI 튜터가 강의를 진행하면서 질문에 답변해줍니다.

MongoDB Academy을(를) 시작하는 데 경험이 필요한가요?

사전 경험은 필요하지 않습니다. CoddyKit의 MongoDB Academy은(는) 초급자부터 고급 학습자까지를 위해 구성되어 있으므로, 여기서 시작하거나 처음부터 시작할 수 있으며 자신의 속도대로 진행할 수 있습니다. 이것은 4개 중 3번째 강의입니다.

“범위 기반 샤딩과 해시 기반 샤딩 전략” 강의는 얼마나 걸리나요?

대부분의 CoddyKit 강의는 약 5~10분이 소요됩니다. 각 강의는 간결하고 인터랙티브하여 꾸준한 진행이 가능하며, 웹과 앱에서 중단한 부분부터 바로 시작할 수 있습니다.

이 MongoDB Academy 강의에서 코드를 작성하고 실행할 수 있나요?

네. 모든 MongoDB Academy 강의에는 내장 코드 에디터가 포함되어 있으므로, 브라우저에서 바로 실제 코드를 작성하고 실행한 후 즉시 AI 피드백을 받을 수 있습니다 — 로컬 설정이 필요 없습니다.

이 강의의 모든 강의

  1. 샤딩 개념: 청크, 밸런서, 샤드 키
  2. 샤드 키 선택하기: 카디널리티, 빈도, 단조성
  3. 범위 기반 샤딩과 해시 기반 샤딩 전략
  4. 영역 샤딩: 지역에 데이터 고정하기
← MongoDB Academy(으)로 돌아가기