Reference · Cheat sheet
Lesson 0006 distilled — three questions, three tiers, and the anti-patterns. Built to print.
1. Access pattern: always accessed together? → embed. Sometimes/never? → reference.
2. Cardinality: bounded and small? → embed. Unbounded or large? → reference.
3. Update independence: updated on its own? → reference (embedding rewrites whole parent).
One-to-few (≤~50): embed. Small, bounded, always together. One read, atomic write.
One-to-many (hundreds–thousands): usually reference. Check if bounded + always-together before embedding.
One-to-squillions (millions): reference on the child side — parent _id in each child. Keeps parent document small.
Embedding related data → one document write → atomic with no transaction.
Referencing → two writes → inconsistency possible if process crashes between them → need explicit transaction to fix.
If two pieces of data must always be consistent, embedding lets WiredTiger enforce it for free.
Unbounded array: embedding a set that grows forever → hits 16 MB limit. Flip to child-side reference.
Massive document: embedding everything → bloats cache, slow network, expensive writes.
Over-normalization: referencing everything → $lookup chains → destroys read locality.
Embed the hot slice (top-N most recent/popular) in the parent; full set lives in its own collection.
Hot path = one read (no $lookup). Full list = separate query. Cost: two writes per new item.
Worth it when read:write ratio is high and the hot slice covers >90% of reads.
| Relational | MongoDB | |
|---|---|---|
| Default | normalize + join | embed + co-locate |
| Join cost | near zero (optimizer) | always non-zero ($lookup) |
| Atomicity | any transaction | single document |
| Schema | DDL-enforced | application-level |