Lesson 0008 · Durability

Write concern & read concern

Two orthogonal dials on every write: w controls how many nodes must acknowledge, j controls whether the journal flushed to disk. They don't imply each other — and combining them precisely is how you tune the durability/latency trade-off.

~12 minWarm-up + retrieval quiz + synchronous_commit contrastCheat sheet: here

Warm-up — apply lesson 0007 first (cold, from memory)

The oplog is idempotent. Why does that property matter for crash recovery?

Idempotency means applying f(x) twice = f(x) once. After a crash, the secondary may not know which entries it finished applying. It can safely re-run from its last checkpoint position — re-applying an entry it already applied produces the same document state. No deduplication needed.

A 3-node replica set: to win a primary election, a candidate secondary must satisfy:

Both conditions are required. The most up-to-date oplog ensures no data loss (the new primary has everything the old one committed). Majority votes (2 of 3) prevent split-brain — two nodes can't simultaneously believe they're primary in a 3-node set.

A MongoDB client's default read preference routes reads to:

Default readPreference is "primary" — all reads go to the one primary. You always read the latest committed write. Routing to secondaries is opt-in (secondaryPreferred, secondary) and accepts possible staleness because secondaries apply the oplog asynchronously.

The default write concern (w: 1) means the primary acknowledged your write as soon as it hit its own oplog. But if the primary crashes right after, and the secondaries haven't replicated that oplog entry yet, the next elected primary won't have it — the write is silently rolled back. This isn't a bug; it's a trade-off. Write concern and read concern let you choose exactly where on the durability/latency spectrum you sit.

Vocabulary resolved. Since lesson 0002, warm-ups have drilled j: true and w separately because they're easy to conflate. Here's the definitive split:

j: true — one node's journal question: "Did this node flush the write to its WAL before acking?" About surviving a process crash on one machine.

w: N — the replication question: "How many nodes acknowledged this write?" About surviving a node failure.

They are orthogonal axes. A write can be journaled-but-not-replicated or replicated-but-not-journaled. You set both independently, on every write if you want.

Write concern — the w dial

w specifies how many replica set members must acknowledge the write before the driver returns to the application. — MongoDB Manual: "Write concern describes the level of acknowledgment requested from MongoDB for write operations to a standalone mongod or to replica sets or to sharded clusters."

ValueMeaningRisk
w: 0Fire and forget — no acknowledgment at allNo error reporting; data loss on any failure
w: 1 (default)Primary has written to its oplogRollback if primary crashes before replication
w: 2, w: NPrimary + N−1 secondaries acknowledgedLower rollback risk; higher latency
w: "majority"A majority of voting members acknowledgedMinimal — any elected primary already has the write

w: "majority" is the key value: because elections require a majority, any secondary that wins an election is guaranteed to already have the write. Rollback of a majority-acknowledged write is essentially impossible under normal failure modes. — MongoDB Manual: "If you use writeConcern: majority, MongoDB returns acknowledgment of the write operation after a majority of the voting members have applied the write operation."

The journal flag — the j dial

j is a per-write flag that tells the primary whether to wait for the journal (write-ahead log) to flush to disk before sending acknowledgment. — MongoDB Manual: "The j option requests acknowledgment that the write operation has been written to the on-disk journal."

ValueMeaningRisk
j: false (default)Primary acks without waiting for journal flush; write is in memory (committed to oplog buffer)If the primary process crashes before the next journal write (~100 ms), the write is lost even on that node
j: truePrimary flushes the write to the WAL journal before ackingAdds latency (~journal flush time, sub-millisecond on SSD); write survives any single-process crash

This is the direct analog of the WiredTiger journal from lesson 0002: the journal flushes every 100 ms by default. Setting j: true forces an immediate flush — the same way PostgreSQL's synchronous_commit = local forces a local WAL sync.

The matrix — combining w and j

Pick your position on the durability spectrum

CombinationSurvivesDoesn't surviveUse when
w: 1, j: false (default) Network errors (you got an ack) Primary process crash before journal flush; primary node failure before replication Bulk imports, caches, metrics — data you can reconstruct
w: 1, j: true Primary process crash Primary node failure before replication — write may still roll back Durable on primary, not guaranteed replicated; middle ground
w: "majority", j: false Primary node failure (majority have it) Simultaneous crash of majority members before journal flush — unlikely but possible Read-heavy systems where replicated > journaled
w: "majority", j: true Primary crash + primary node failure; rollback impossible under normal failure Correlated failure of a majority of nodes simultaneously Financial systems, order processing, anything where the write must survive

Read concern — matching durability on the read side

Write concern controls what's durable when you write. Read concern controls what you see when you read — specifically whether you might read data that could later be rolled back. — MongoDB Manual: "The readConcern option allows you to control the consistency and isolation properties of the data read from replica sets and replica set shards."

LevelReadsCan be rolled back?
local (default)Most recent data on the node, regardless of replication stateYes — if the node was primary and loses an election, unacknowledged writes roll back
majorityData acknowledged by a majority of voting membersNo — only data that can't be rolled back
linearizableReflects all majority-acknowledged writes before the read started; reads always from primary with extra verificationNo — strongest guarantee, highest latency

The practical pair: if your write uses w: "majority", your read using readConcern: "majority" sees only data that has met that same bar. Together they give you the MongoDB equivalent of serializable isolation on a distributed system — without a transaction.

The Postgres contrast

MongoDB settingPostgres analogWhat it means
w: 1, j: falsesynchronous_commit = offFast ack; data in memory; may lose recent commits on crash
w: 1, j: truesynchronous_commit = localLocal WAL flush before ack; durable on one node
w: "majority", j: truesynchronous_commit = on (or remote_apply)Wait for remote WAL confirmation; durable on majority; rollback impossible
readConcern: "majority"Reading from primary with synchronous_commit = onOnly see data that can't be lost to a failover

The structural difference: Postgres applies synchronous_commit globally for the server (or per-session); MongoDB sets w and j per write operation or per collection default. Per-operation granularity is more flexible but also more dangerous — it's easy to use the wrong concern on a critical write.

The default write concern w: 1 risks data loss because:

w: 1 means the primary acked as soon as it wrote to its own oplog buffer. If the primary then crashes before any secondary has replicated that entry, the next elected primary won't have the write — rollback. w: "majority" closes this window by requiring acknowledgment from a majority before returning to the client.

j: true on a write means the driver waits for:

j: true is a single-node durability flag — it only affects the primary. It says: "Don't ack until this write is in the on-disk WAL journal." If the primary process crashes after the ack, the write survives (it's on disk). j says nothing about replication — for that you use w.

w: "majority", j: false means the write:

w and j are orthogonal. w: "majority" → a majority of nodes confirmed receipt (oplog-level). j: false → none of them necessarily flushed to their journal yet. The write survives a node failure (majority have it in memory/oplog) but an unlikely simultaneous crash of the majority before their next journal flush could still lose it. In practice, combined hardware failure of a majority is rare enough that w: "majority", j: false is common for performance-sensitive paths.

readConcern: "majority" guarantees you won't read data that:

readConcern: "majority" filters to data that a majority of voting members have durably acknowledged. That's the same bar as w: "majority" on the write side. Any such data cannot be rolled back by an election — an elected primary already has it. So you will never read data that later disappears due to a failover.

The closest Postgres analog to w: "majority", j: true is:

synchronous_commit = on makes Postgres wait for the remote standby to confirm WAL receipt before returning to the client — matching w: "majority" (remote acknowledgment). The j: true part = local WAL flush = synchronous_commit = local. Combined: on is the closest match. remote_apply (stronger — waits for apply, not just receipt) would be even closer if you want the full analog.
Optional — when you have a mongosh instance: insert a document with db.col.insertOne({x: 1}, { writeConcern: { w: "majority", j: true } }) and observe the extra latency vs the default. Then try db.col.find().readConcern("majority") to see majority-read in action. In a 3-node set you can kill the primary immediately after a w: 1 write and watch the new primary roll it back.