Lesson 0007 · Replication
Mongo's high-availability primitive — automatic failover, not a bolt-on. One primary, N secondaries, one oplog. The oplog is what makes it tick: a capped, logical, idempotent operation log that your Postgres and MySQL courses will recognize immediately.
One device, potentially millions of sensor readings. Where does the parent reference go, and why?
You embed an author sub-document inside the blog post document. Updating both the post title and the author's bio in one operation is atomic because:
Why is embedding event-log entries in an array on the parent document an anti-pattern?
In Postgres you configure streaming replication manually and promote standbys by hand (or outsource it to Patroni). In MongoDB, high availability is built into the topology primitive: a replica set. There is no "standalone plus optional replication" mode for production — a replica set is the default deployment, always.
A replica set is a group of mongod instances — typically three — that hold the same data set.
— MongoDB Manual: "A replica set in MongoDB is a group of mongod processes that maintain the same data set. Replica sets provide redundancy and high availability."
Exactly one member is the primary at any moment; the rest are secondaries.
All writes go to the primary; secondaries replicate by tailing the primary's operation log (the oplog).
A third node type is the arbiter: a lightweight process with no data, which exists only to vote in elections. Arbiters let you run a 2-data-node set (cheaper) with a quorum still reachable. They add no redundancy — don't use them if you can afford three full nodes. — MongoDB Manual: "An arbiter does not have a copy of data set and cannot become a primary… Add an arbiter to a replica set to have an odd number of members."
The oplog lives at local.oplog.rs — a capped collection in the local database on every replica set member.
— MongoDB Manual: "The oplog (operations log) is a special capped collection that keeps a rolling record of all operations that modify the data stored in your databases."
Every write the primary executes is recorded here as a logical, idempotent operation.
Secondaries tail this collection and apply each entry in order.
Two properties matter most:
Each oplog entry is a BSON document. The key fields:
The contrast with your existing courses:
| Engine | Replication log | Level | Idempotent? |
|---|---|---|---|
| Postgres streaming | WAL (write-ahead log) | Physical — byte-level page changes | No (page apply is not idempotent) |
| MySQL binlog (row-based) | Binary log | Logical — before/after row images | In practice yes (with row images) |
| MongoDB replica set | oplog (local.oplog.rs) | Logical — full transformation | Yes, by design |
The physical/logical distinction matters for portability and recovery. Postgres WAL is tied to the exact on-disk page format and the major version — you can't stream a Postgres 14 WAL to a Postgres 15 standby. MongoDB's logical oplog is engine-independent: the same oplog entry can be applied regardless of what WiredTiger internal pages look like. — MongoDB Manual: "MongoDB applies database operations on the primary and then records the operations on the primary's oplog. The secondary members then copy and apply these operations in an asynchronous process."
Members send each other heartbeats every two seconds.
If a secondary doesn't hear from the primary within electionTimeoutMillis (default 10 seconds), it concludes the primary is unreachable and calls for an election.
— MongoDB Manual: "When the primary is unavailable, an eligible secondary will hold an election to select itself as the new primary. The first secondary to call the election and receive votes from a majority of the voting members becomes the new primary."
The winner must:
ts) among the candidates.During an election — typically 10–30 seconds — writes fail (no primary to accept them) and reads from primary also fail. This is the availability window your SLA must account for. Applications should use drivers with retryable writes enabled (MongoDB driver default since 4.2) so transient election failures are hidden.
Election vocabulary your Postgres course touches: quorum (here it's majority of voting members, not just nodes); priority (you can set priority 0 on a member to prevent it ever becoming primary — useful for a geographically distant data-center replica you want to keep as read replica only).
By default, all reads go to the primary (readPreference: "primary") — you always read your own writes.
You can change this per operation or per connection:
secondaryPreferred routes to a secondary if one is available, falling back to primary; secondary always routes to a secondary, accepting potential staleness.
— MongoDB Manual: "Read preference describes how MongoDB clients route read operations to the members of a replica set… By default, an application directs its read operations to the primary member in a replica set."
Stale reads are the price: a secondary may be seconds behind the primary's oplog.
Lesson 0008 (write/read concern) covers how to trade consistency and durability precisely.
w: 1, the primary acknowledges as soon as it writes to its own oplog — secondaries may not have applied it yet. With w: "majority", it waits for a majority of voting members to confirm. Combined with j: true (journal flush before ack), you control the full durability/latency trade-off. That's the whole of lesson 0008.
A 3-node replica set can survive how many simultaneous node failures while still electing a new primary?
The oplog is idempotent. What does that guarantee if a secondary crashes mid-apply and replays the same oplog entry on restart?
The MongoDB oplog vs the Postgres WAL — the sharpest distinction:
A secondary's oplog position has fallen behind a point that has been overwritten (the oplog window has passed). What happens next?
By default, where does a MongoDB client send read operations in a replica set?
rs.initiate(). Run rs.status() to see the primary/secondary assignments and oplog lag. Then db.printReplicationInfo() to see the oplog window (how many hours of operations it can hold). Kill the primary process and watch rs.status() as an election fires.