Reference · Cheat sheet

Sharding

Lesson 0009 distilled — three components, shard key properties, targeted vs scatter-gather. Built to print.

From lesson 0009Context: MongoDB current (7.0/8.0+)

Three components

mongos — query router. Clients connect here. Holds no data. Routes using the chunk map.

Config servers — replica set storing chunk metadata (which chunks live on which shards).

Shards — each is a full replica set holding one horizontal slice of the data.

Shard key — three properties

Cardinality: many distinct values → many chunks → can distribute across many shards. Low cardinality caps chunk count.

Write distribution: are inserts spread evenly? Monotonically increasing key + range sharding = write hotspot.

Query isolation: most queries include the shard key → targeted; missing the shard key → scatter-gather.

Targeted vs scatter-gather

Targeted: filter includes shard key → mongos routes to exactly one shard. Fast.

Scatter-gather: filter missing shard key → mongos broadcasts to all shards, merges results. Latency = slowest shard. Gets worse as you add shards. The COLLSCAN of sharding.

Range vs hash sharding

RangeHash
Write dist.Bad for monotonic keysEven always
Range queriesTargetedScatter-gather
Use whenRange queries dominate, non-monotonic keyWrite throughput, point lookups

Anti-patterns

Low cardinality: boolean/status → max chunks = distinct values → caps shard count.

Monotonic + range: ObjectId / timestamp → all inserts → one shard. Write hotspot.

Unselective compound: 90% data in one value → imbalanced chunks; balancer can't fix it.

Good pattern

{ userId: 1, createdAt: 1 } — userId = cardinality + even distribution; most queries filter by userId (targeted); createdAt = sort ordering within shard.

Or hash on a high-cardinality field when write distribution matters more than range queries.

Warning: shard key is irreversible without resharding (full data copy).