Reference · Glossary
Glossary
Canonical definitions for this course. When a term is defined here, lessons use it exactly this way.
Living documentGrows with each lesson
- query context
- Evaluation mode that asks “how well does this document match?” and produces a relevance
_score. Clauses in must and should run here.
— intro'd in lesson 0001
- filter context
- Evaluation mode that asks “does this document match, yes or no?” — no score is computed. Cacheable. Clauses in
filter and must_not run here.
— intro'd in lesson 0001
- _score
- A non-negative float ranking how well a document matched the query. Higher = more relevant. Only produced in query context; default scoring is BM25.
— intro'd in lesson 0001
- bool query
- The primary compound query. Combines clauses under four keys:
must, should, filter, must_not.
— intro'd in lesson 0001
- minimum_should_match
- How many
should clauses a document must satisfy. Defaults to 1 when the bool has no must/filter, and to 0 when it does.
— intro'd in lesson 0001
- node query cache (filter cache)
- A per-node LRU cache holding bitsets of which documents match reused filter-context clauses. Why filters are cheap on repeat queries.
— intro'd in lesson 0001
- constant_score
- A query that wraps a filter and assigns every matching document the same fixed score (default 1.0), skipping relevance computation entirely.
— intro'd in lesson 0001
- BM25
- Best Match 25 — the default similarity (since ES 5.0), replacing classic TF/IDF. Scores each query term as
IDF · [saturated TF, length-normalized], summed over terms.
— intro'd in lesson 0001, derived in 0002
- IDF (inverse document frequency)
- Weights a term by rarity:
ln(1 + (N − n + 0.5)/(n + 0.5)), where N = total docs and n = docs containing the term. Rarer term → higher weight.
— intro'd in lesson 0002
- TF saturation
- BM25's property that repeated occurrences of a term add ever less to the score, approaching a ceiling of
IDF · (k1+1). The fix for raw TF/IDF's unbounded growth.
— intro'd in lesson 0002
- k1
- BM25 knob for TF saturation rate. Default 1.2.
0 ignores term frequency; higher values saturate more slowly.
— intro'd in lesson 0002
- b
- BM25 knob for field-length normalization, range 0–1, default 0.75.
0 ignores length; 1 applies the full long-field penalty via |D|/avgdl.
— intro'd in lesson 0002
- dfs_query_then_fetch
- A search type that gathers global term statistics across shards before scoring, so per-shard IDF differences don't skew results. Useful for small or test indices.
— intro'd in lesson 0002
- analysis
- The process of turning a text value into a stream of terms via an analyzer. Runs at index time (on field values) and search time (on the query string). Mismatched runs cause silent zero-hit failures.
— intro'd in lesson 0003
- analyzer
- The pipeline that performs analysis, always in three stages: char filters → tokenizer → token filters. The default
standard analyzer = standard tokenizer + lowercase filter.
— intro'd in lesson 0003
- term (token)
- The atomic unit stored in and searched against the inverted index — the output of analysis. BM25 scores per term. What counts as a term is decided by the analyzer, not the query.
— intro'd in lesson 0003
- tokenizer
- The single mandatory analyzer stage that splits a character stream into tokens (e.g.
standard breaks on word boundaries). Preceded by char filters, followed by token filters.
— intro'd in lesson 0003
- token filter
- An analyzer stage that adds, removes, or rewrites tokens after the tokenizer — e.g.
lowercase, stop, stemmer, synonym. Zero or more per analyzer.
— intro'd in lesson 0003
- character filter
- An analyzer stage that rewrites the raw character stream before tokenizing — e.g.
html_strip, character mapping. Zero or more per analyzer.
— intro'd in lesson 0003
- text field
- A mapping type that is analyzed: its value becomes many terms, powering full-text
match queries. Not suitable for sorting or aggregations.
— intro'd in lesson 0003
- keyword field
- A mapping type that is not analyzed: the whole value is stored verbatim as one exact term. Used for exact match, sorting, and aggregations. Dynamic strings map to both a
text field and a .keyword sub-field.
— intro'd in lesson 0003
- _analyze API
- Endpoint that returns the token stream an analyzer (or a mapped field) produces for a given text — the ground truth of what lands in the index. The primary tool for diagnosing the term-mismatch trap.
— intro'd in lesson 0003
- inverted index
- The core search structure: maps each term → the documents that contain it (the inverse of a document→words forward index). One per field. Made of a term dictionary plus postings lists.
— intro'd in lesson 0004
- term dictionary
- The sorted list of every unique term in a field's inverted index. Sorting enables binary-search term lookup and contiguous prefix/range scans (and is why leading wildcards are slow).
— intro'd in lesson 0004
- postings list
- For one term, the sorted list of doc IDs that contain it (optionally with term frequencies, positions, and offsets). Its length is the term's document frequency. Sorted doc IDs make multi-term AND a one-pass merge.
— intro'd in lesson 0004
- document frequency (df)
- The number of documents containing a term = the length of its postings list. This is IDF's
n from lesson 0002; rarer term → shorter postings → higher IDF.
— intro'd in lesson 0004
- _termvectors API
- Endpoint returning the per-term statistics stored in the inverted index for a document —
term_freq, and with term_statistics, doc_freq. Lets you read BM25's raw inputs directly.
— intro'd in lesson 0004
- segment
- A self-contained inverted index and the unit a shard is built from. Immutable once written — never edited in place; changes come as new segments. A search runs against every segment in the shard.
— intro'd in lesson 0005
- shard (Lucene index)
- A Lucene index = a collection of segments plus a commit point (a file listing the live segments). The unit Elasticsearch distributes and replicates.
— intro'd in lesson 0005
- refresh
- The operation (default every 1s) that writes the in-memory buffer to a new searchable segment in the filesystem cache. Until it runs, newly indexed docs are not searchable — hence “near real-time.” Tunable via
refresh_interval.
— intro'd in lesson 0005
- near real-time search
- ES's property that document changes become visible to search within about one second (after the next refresh), not instantly.
— intro'd in lesson 0005
- .del file (soft delete)
- A per-segment file marking documents as deleted. Because segments are immutable, a delete only marks the doc (filtered from results); an update is a delete-mark plus a re-index. The bytes are removed only at merge.
— intro'd in lesson 0005
- merge
- The automatic background process that combines smaller segments into larger ones. Deleted/old-version docs are not copied across, so merge is when disk is reclaimed. Throttled to protect search performance.
— intro'd in lesson 0005
- force merge
- A manual API (
_forcemerge) that merges a shard down to a target segment count. Appropriate for read-only indices; an anti-pattern on actively-written ones (creates oversized segments the normal policy won't re-merge).
— intro'd in lesson 0005
- occurrence type
- A
bool clause key defining how a clause participates: must (match, scored), should (optional, scored, summed), filter (match, unscored, cached), must_not (exclude, unscored, cached).
— intro'd in lesson 0006
- more-matches-is-better
- The
bool scoring rule: the BM25 _score of every matching must and should clause is added together to form the final score.
— intro'd in lesson 0006
- aggregation
- A summary computed over the documents matched by the query — counts, statistics, and groupings. Runs in the same request as the search; set top-level
size: 0 to return only aggregations.
— intro'd in lesson 0007
- metric aggregation
- An aggregation that calculates a value from field values across the docs — e.g.
avg, sum, min/max, stats, cardinality.
— intro'd in lesson 0007
- bucket aggregation
- An aggregation that groups documents into buckets (each a collection of docs meeting a criterion) by field value, range, or interval. Each bucket carries a
doc_count. Examples: terms, range, date_histogram.
— intro'd in lesson 0007
- pipeline aggregation
- An aggregation whose input is the output of other aggregations rather than documents — e.g.
avg_bucket, derivative.
— intro'd in lesson 0007
- sub-aggregation (nesting)
- An aggregation placed inside a bucket aggregation's own
aggs, computing a metric or further buckets per bucket. The mechanism behind drill-down analysis.
— intro'd in lesson 0007
- terms aggregation
- A bucket aggregation building one bucket per unique value (default top 10 by
doc_count). Run it on a keyword field, not text. Counts are approximate across shards — see doc_count_error_upper_bound / sum_other_doc_count.
— intro'd in lesson 0007
- primary shard
- An original shard of an index that owns a subset of its documents. The number of primaries is fixed at index creation and determines how the data is partitioned.
— intro'd in lesson 0008
- replica shard
- A copy of a primary shard, for redundancy and read throughput. A replica is never allocated on the same node as its primary — so on a single node, replicas stay
UNASSIGNED.
— intro'd in lesson 0008
- over-sharding
- Having too many small shards. Because each shard carries fixed heap/thread/cluster-state overhead and a search uses one thread per shard, over-sharding degrades search performance and can destabilize the cluster. Target 10–50 GB per shard.
— intro'd in lesson 0008
- search thread pool
- The bounded pool of threads a node uses to run shard-level searches — one thread per shard per query. Fanning a query across very many shards can exhaust it, the sharp edge of over-sharding.
— intro'd in lesson 0008
← lesson 0001 · lesson 0008 →