Reference · Cheat sheet

Schema design: embed vs reference

Lesson 0006 distilled — three questions, three tiers, and the anti-patterns. Built to print.

From lesson 0006Context: MongoDB current (7.0/8.0+)

Three questions — ask in order

1. Access pattern: always accessed together? → embed. Sometimes/never? → reference.

2. Cardinality: bounded and small? → embed. Unbounded or large? → reference.

3. Update independence: updated on its own? → reference (embedding rewrites whole parent).

Three tiers

One-to-few (≤~50): embed. Small, bounded, always together. One read, atomic write.

One-to-many (hundreds–thousands): usually reference. Check if bounded + always-together before embedding.

One-to-squillions (millions): reference on the child side — parent _id in each child. Keeps parent document small.

Atomicity is a schema argument

Embedding related data → one document write → atomic with no transaction.

Referencing → two writes → inconsistency possible if process crashes between them → need explicit transaction to fix.

If two pieces of data must always be consistent, embedding lets WiredTiger enforce it for free.

Anti-patterns

Unbounded array: embedding a set that grows forever → hits 16 MB limit. Flip to child-side reference.

Massive document: embedding everything → bloats cache, slow network, expensive writes.

Over-normalization: referencing everything → $lookup chains → destroys read locality.

Subset pattern

Embed the hot slice (top-N most recent/popular) in the parent; full set lives in its own collection.

Hot path = one read (no $lookup). Full list = separate query. Cost: two writes per new item.

Worth it when read:write ratio is high and the hot slice covers >90% of reads.

Relational contrast

RelationalMongoDB
Defaultnormalize + joinembed + co-locate
Join costnear zero (optimizer)always non-zero ($lookup)
Atomicityany transactionsingle document
SchemaDDL-enforcedapplication-level