Reference · Glossary

MongoDB glossary

The shared vocabulary of this course. Once a term is defined here, every lesson uses it this way — with the relational contrast attached where it helps.

Context: MongoDB current (7.0/8.0+)All lessons

Document
The basic unit of data: an ordered set of field/value pairs, stored as BSON. A tree, not a flat row — values can be nested documents and arrays. Introduced in 0001. Relational analog: a row (but schema-free and nestable).
BSON
Binary JSON — MongoDB's on-disk and on-the-wire encoding of documents. A binary superset of JSON with extra types (dates, 64-bit ints, binary, ObjectId, Decimal128).
Collection
A grouping of documents. The rough equivalent of a table, but with no enforced schema — documents in one collection can have different shapes.
_id
The mandatory primary-key field, unique within a collection and immutable. Auto-populated with an ObjectId if you don't supply one. MongoDB always keeps a unique index on it. Relational analog: the primary key.
ObjectId
A 12-byte identifier the driver generates for _id when you omit it. Roughly time-ordered, so it sorts by creation time.
Embedding
Storing related data inside a document (nested doc or array), so it's read in one lookup. The default modeling move, driven by "data accessed together is stored together." Introduced in 0001. Relational analog: the opposite of normalizing into separate tables.
Referencing
Linking documents by storing another document's _id and looking it up separately. Used when data is large, unbounded, or accessed independently. Relational analog: a foreign key (but joins are manual / via $lookup).
Single-document atomicity
A write to one document is all-or-nothing, even across many fields and nested arrays — no transaction required. Atomicity is a property of the document boundary. Introduced in 0001.
Distributed (multi-document) transaction
The mechanism for atomic reads/writes spanning more than one document or collection. Costs more than a single-document write; good schema design (embedding) often removes the need. Relational analog: an ordinary BEGIN … COMMIT transaction.
WiredTiger
MongoDB's default storage engine: stores each collection and index as a B-tree, with MVCC, document-level concurrency, a write-ahead journal, and default compression. Introduced in 0002. Relational analog: closest to InnoDB (B+tree + buffer pool + redo log), not the Postgres heap.
MVCC snapshot
The consistent point-in-time view WiredTiger hands an operation at its start, so reads don't block writes. Introduced in 0002. Relational analog: the Postgres/InnoDB snapshot read.
Document-level concurrency
WiredTiger's write granularity: many clients can modify different documents of one collection simultaneously, coordinated optimistically with only intent locks above the document. Relational analog: InnoDB row-level locking.
Checkpoint
A full, consistent snapshot of the data written to disk every 60 seconds; acts as a crash-recovery point. Introduced in 0002.
Journal
WiredTiger's write-ahead log. Flushed ~every 100 ms; on crash, recovery replays journal records written since the last checkpoint. Introduced in 0002. Relational analog: the Postgres WAL / InnoDB redo log.
WiredTiger cache
The engine's internal RAM cache for the working set, default max(50% of (RAM − 1 GB), 256 MB), holding data uncompressed; the OS filesystem cache holds the compressed on-disk blocks. Relational analog: the InnoDB buffer pool.
compact
The command that returns WiredTiger's reused-in-place free space back to the operating system (also achieved by resyncing a replica-set member). Relational analog: VACUUM FULL / OPTIMIZE TABLE.
Collection scan (COLLSCAN)
Reading every document because no useful index exists. Introduced in 0003. Relational analog: a Postgres Seq Scan / MySQL full table scan.
Compound index
An index over several fields, stored sorted by the declared field order. Introduced in 0003.
Index prefix
A beginning subset of a compound index's fields. A query can use the index only via a prefix (it must include the leading field). Relational analog: InnoDB's leftmost-prefix rule.
ESR rule
The field-order guideline for a compound index: Equality, then Sort, then Range. Equality first (most selective, keeps the rest sorted); sort before range (a range breaks sort order for later fields). Variant ERS when the range is very selective. Introduced in 0003.
Multikey index
An index on an array-valued field; MongoDB creates one entry per array element, all pointing to the same document. Automatic. A compound index may hold at most one array field, and a multikey index cannot cover a query. Introduced in 0003. Relational analog: none — scalar columns can't do this.
Covered query
A query answered entirely from an index, without fetching documents, because every field it needs is in the index. Introduced in 0003. Relational analog: Postgres index-only scan / InnoDB covering index.
explain()
The command that shows a query's plan and, in executionStats mode, its measured counts. Verbosities: queryPlanner (default, no run), executionStats (runs it), allPlansExecution (adds losers' trial data). Ignores and doesn't populate the plan cache. Introduced in 0004. Relational analog: EXPLAIN / EXPLAIN ANALYZE (but safe on writes).
IXSCAN / FETCH / SORT
Plan stages: IXSCAN walks index keys; FETCH retrieves the pointed-to document; SORT is an in-memory sort (the index didn't supply order). An IXSCAN with no FETCH is a covered query. Introduced in 0004.
Multi-planner (the race)
MongoDB's classic plan selection: candidate plans are run side by side for a short trial and the one producing the most results for the least work wins. 8.3+ adds a cost-based ranker backup. Introduced in 0004. Relational analog: none — Postgres/MySQL estimate cost from statistics instead of racing.
examined-vs-returned
The efficiency tell: totalKeysExamined/totalDocsExamined compared to nReturned. Near 1:1 is efficient; a big gap means much work per row. Introduced in 0004. Relational analog: Postgres estimated-vs-actual rows.
Plan cache
Stores the winning plan keyed by query shape and reuses it for matching queries. Cleared by any DDL (creating/dropping/hiding an index), LRU eviction, and restart. Introduced in 0004.
Aggregation pipeline
MongoDB's language for computing over documents: an ordered list of stages, each transforming a stream of documents and feeding the next. Read-only unless it ends in $out/$merge. Introduced in 0005. Relational analog: a SELECT … WHERE … GROUP BY … ORDER BY, written as stages.
Stage
One step in a pipeline. Core stages: $match (WHERE), $group (GROUP BY), $project (SELECT list), $sort (ORDER BY), $limit/$skip, $unwind (≈ UNNEST), $lookup (join). Introduced in 0005.
$lookup
The stage that performs a left outer join to another collection in the same database, adding an array field of matched foreign documents. Overuse signals a schema that should embed instead. Introduced in 0005. Relational analog: LEFT OUTER JOIN — but native and cheap there, avoided here.
$unwind
The stage that explodes an array field into one document per element, so later stages can group or match on individual elements. Introduced in 0005. Relational analog: ≈ UNNEST — no clean equivalent; arrays are Mongo-specific.
Predicate pushdown (aggregation)
The optimizer moving $match earlier — ahead of $sort (less to sort) and ahead of a projection (so the first stage can use an index). Also $sort+$limit coalescing into a top-N sort. Introduced in 0005.
One-to-few / one-to-many / one-to-squillions
A three-tier cardinality model for schema design. One-to-few (≤~50): embed. One-to-many (hundreds–thousands): usually reference. One-to-squillions (millions): reference on the child side — store the parent's _id in each child document to avoid growing an unbounded array on the parent. Introduced in 0006.
Unbounded array (anti-pattern)
Embedding a set of sub-documents that grows without limit in an array on the parent document. Causes monotonic document growth, eventually hitting the 16 MB limit, and bloats every parent read with data rarely needed. Fix: reference on the child side, or use the subset pattern. Introduced in 0006.
Subset pattern
Embedding only the hot slice of a large set (e.g., the 5 most recent reviews) in the parent document, while the full set lives in its own collection. Hot-path reads are served from the embed (no $lookup); full-list reads query the separate collection. Cost: two writes per new item (both the child collection and the embedded array on the parent). Introduced in 0006.
Replica set
A group of mongod instances (typically 3) that hold the same data. Exactly one is the primary; the rest are secondaries. All writes go to the primary; secondaries replicate by tailing the oplog. Includes built-in automatic failover via elections. Introduced in 0007. Relational analog: Postgres streaming replication + Patroni, but built-in rather than bolted on.
oplog (local.oplog.rs)
A capped collection on every replica set member that records every write operation as a logical, idempotent transformation. Secondaries tail the primary's oplog and apply entries in order. When the oplog rolls past a secondary's last-applied position, the secondary needs a full resync. Introduced in 0007. Relational analog: Postgres WAL (physical, not idempotent) / MySQL binlog (logical, less so).
Election
The process triggered when a secondary detects the primary is unreachable (after electionTimeoutMillis, default 10 s). A candidate secondary receives votes from a majority of voting members; the most up-to-date one wins and becomes the new primary. Writes fail during the election window (~10–30 s). Introduced in 0007.
Read preference
Where the driver routes read operations: primary (default — strong consistency), secondaryPreferred (secondary if available, else primary), secondary (always secondary, accepts staleness). Introduced in 0007.
Arbiter
A lightweight replica set member that votes in elections but holds no data and can never become primary. Used to give a 2-data-node set an odd number of voting members. Adds no data redundancy. Introduced in 0007.
Write concern (w)
How many replica set members must acknowledge a write before the driver returns to the application. w: 1 (default) = primary only; w: "majority" = a majority of voting members, eliminating rollback risk. Controls replication durability. Introduced in 0008. Relational analog: Postgres synchronous_commit (off / local / on).
Journal flag (j)
A per-write boolean controlling whether the primary waits for the journal (write-ahead log) to flush to disk before acking. j: false (default) = in-memory ack; j: true = on-disk journal flush. Controls single-node disk durability. Orthogonal to w. Introduced in 0008.
Read concern
Controls the consistency of data returned by a read. local (default) = most recent on the node, may include data that could roll back; majority = only data acknowledged by a majority (cannot roll back); linearizable = strongest, reflects all majority-acked writes before the read started. Introduced in 0008. Relational analog: Postgres isolation level / reading from a synchronized standby.
Sharding
Horizontal partitioning of a collection across multiple replica sets (shards), each holding a slice of the data keyed by the shard key. Enables storage and write throughput beyond single-machine capacity. The cluster is accessed through mongos routers. Introduced in 0009. Relational analog: none in-core for Postgres/MySQL — requires CitusDB, Vitess, or similar external tools.
Shard key
The field (or compound fields) used to assign documents to chunks and shards. A good shard key has high cardinality, even write distribution, and query isolation (most queries include it). An irreversible choice without resharding. Introduced in 0009.
Chunk
A contiguous range of shard key values that lives on exactly one shard. Chunks split when they exceed the configured size (default 128 MB) and are migrated by the balancer to maintain even distribution. Introduced in 0009.
mongos
The query router for a sharded cluster. Clients connect to mongos (not to shards). It reads the chunk map from config servers and routes each query to the correct shard(s). Holds no data. Introduced in 0009.
Targeted query / scatter-gather
A targeted query includes the shard key — mongos routes it to exactly one shard. A scatter-gather query is missing the shard key — mongos broadcasts to every shard and merges results. Scatter-gather latency = slowest shard and grows with shard count. The sharding analog of COLLSCAN vs IXSCAN. Introduced in 0009.
Multi-document transaction
A session-scoped, all-or-nothing operation that can span multiple documents and collections. Uses snapshot isolation — the transaction reads from a consistent snapshot taken at its start; writes are buffered until commitTransaction(). Available on replica sets (4.0+) and sharded clusters (4.2+). The escape hatch from single-document atomicity, not the default. Introduced in 0010.
Snapshot isolation (transaction-level)
Each transaction reads from a consistent point-in-time snapshot. Concurrent writes made after the snapshot are invisible for the transaction's lifetime. At commit, writes apply atomically. Write conflicts abort the later transaction. The same MVCC mechanism WiredTiger uses for single reads (0002), extended to span the entire transaction lifetime. Introduced in 0010.
Write conflict
Occurs when two transactions concurrently attempt to write the same document. The second transaction to reach the conflict point must abort and retry. MongoDB uses optimistic concurrency — no pre-locking; conflicts detected at write time. Introduced in 0010.