Reference · Glossary

Glossary

Canonical definitions for this course. When a term is defined here, lessons use it exactly this way.

Living documentGrows with each lesson

query context
Evaluation mode that asks “how well does this document match?” and produces a relevance _score. Clauses in must and should run here. — intro'd in lesson 0001
filter context
Evaluation mode that asks “does this document match, yes or no?” — no score is computed. Cacheable. Clauses in filter and must_not run here. — intro'd in lesson 0001
_score
A non-negative float ranking how well a document matched the query. Higher = more relevant. Only produced in query context; default scoring is BM25. — intro'd in lesson 0001
bool query
The primary compound query. Combines clauses under four keys: must, should, filter, must_not. — intro'd in lesson 0001
minimum_should_match
How many should clauses a document must satisfy. Defaults to 1 when the bool has no must/filter, and to 0 when it does. — intro'd in lesson 0001
node query cache (filter cache)
A per-node LRU cache holding bitsets of which documents match reused filter-context clauses. Why filters are cheap on repeat queries. — intro'd in lesson 0001
constant_score
A query that wraps a filter and assigns every matching document the same fixed score (default 1.0), skipping relevance computation entirely. — intro'd in lesson 0001
BM25
Best Match 25 — the default similarity (since ES 5.0), replacing classic TF/IDF. Scores each query term as IDF · [saturated TF, length-normalized], summed over terms. — intro'd in lesson 0001, derived in 0002
IDF (inverse document frequency)
Weights a term by rarity: ln(1 + (N − n + 0.5)/(n + 0.5)), where N = total docs and n = docs containing the term. Rarer term → higher weight. — intro'd in lesson 0002
TF saturation
BM25's property that repeated occurrences of a term add ever less to the score, approaching a ceiling of IDF · (k1+1). The fix for raw TF/IDF's unbounded growth. — intro'd in lesson 0002
k1
BM25 knob for TF saturation rate. Default 1.2. 0 ignores term frequency; higher values saturate more slowly. — intro'd in lesson 0002
b
BM25 knob for field-length normalization, range 0–1, default 0.75. 0 ignores length; 1 applies the full long-field penalty via |D|/avgdl. — intro'd in lesson 0002
dfs_query_then_fetch
A search type that gathers global term statistics across shards before scoring, so per-shard IDF differences don't skew results. Useful for small or test indices. — intro'd in lesson 0002
analysis
The process of turning a text value into a stream of terms via an analyzer. Runs at index time (on field values) and search time (on the query string). Mismatched runs cause silent zero-hit failures. — intro'd in lesson 0003
analyzer
The pipeline that performs analysis, always in three stages: char filters → tokenizer → token filters. The default standard analyzer = standard tokenizer + lowercase filter. — intro'd in lesson 0003
term (token)
The atomic unit stored in and searched against the inverted index — the output of analysis. BM25 scores per term. What counts as a term is decided by the analyzer, not the query. — intro'd in lesson 0003
tokenizer
The single mandatory analyzer stage that splits a character stream into tokens (e.g. standard breaks on word boundaries). Preceded by char filters, followed by token filters. — intro'd in lesson 0003
token filter
An analyzer stage that adds, removes, or rewrites tokens after the tokenizer — e.g. lowercase, stop, stemmer, synonym. Zero or more per analyzer. — intro'd in lesson 0003
character filter
An analyzer stage that rewrites the raw character stream before tokenizing — e.g. html_strip, character mapping. Zero or more per analyzer. — intro'd in lesson 0003
text field
A mapping type that is analyzed: its value becomes many terms, powering full-text match queries. Not suitable for sorting or aggregations. — intro'd in lesson 0003
keyword field
A mapping type that is not analyzed: the whole value is stored verbatim as one exact term. Used for exact match, sorting, and aggregations. Dynamic strings map to both a text field and a .keyword sub-field. — intro'd in lesson 0003
_analyze API
Endpoint that returns the token stream an analyzer (or a mapped field) produces for a given text — the ground truth of what lands in the index. The primary tool for diagnosing the term-mismatch trap. — intro'd in lesson 0003
inverted index
The core search structure: maps each term → the documents that contain it (the inverse of a document→words forward index). One per field. Made of a term dictionary plus postings lists. — intro'd in lesson 0004
term dictionary
The sorted list of every unique term in a field's inverted index. Sorting enables binary-search term lookup and contiguous prefix/range scans (and is why leading wildcards are slow). — intro'd in lesson 0004
postings list
For one term, the sorted list of doc IDs that contain it (optionally with term frequencies, positions, and offsets). Its length is the term's document frequency. Sorted doc IDs make multi-term AND a one-pass merge. — intro'd in lesson 0004
document frequency (df)
The number of documents containing a term = the length of its postings list. This is IDF's n from lesson 0002; rarer term → shorter postings → higher IDF. — intro'd in lesson 0004
_termvectors API
Endpoint returning the per-term statistics stored in the inverted index for a document — term_freq, and with term_statistics, doc_freq. Lets you read BM25's raw inputs directly. — intro'd in lesson 0004
segment
A self-contained inverted index and the unit a shard is built from. Immutable once written — never edited in place; changes come as new segments. A search runs against every segment in the shard. — intro'd in lesson 0005
shard (Lucene index)
A Lucene index = a collection of segments plus a commit point (a file listing the live segments). The unit Elasticsearch distributes and replicates. — intro'd in lesson 0005
refresh
The operation (default every 1s) that writes the in-memory buffer to a new searchable segment in the filesystem cache. Until it runs, newly indexed docs are not searchable — hence “near real-time.” Tunable via refresh_interval. — intro'd in lesson 0005
near real-time search
ES's property that document changes become visible to search within about one second (after the next refresh), not instantly. — intro'd in lesson 0005
.del file (soft delete)
A per-segment file marking documents as deleted. Because segments are immutable, a delete only marks the doc (filtered from results); an update is a delete-mark plus a re-index. The bytes are removed only at merge. — intro'd in lesson 0005
merge
The automatic background process that combines smaller segments into larger ones. Deleted/old-version docs are not copied across, so merge is when disk is reclaimed. Throttled to protect search performance. — intro'd in lesson 0005
force merge
A manual API (_forcemerge) that merges a shard down to a target segment count. Appropriate for read-only indices; an anti-pattern on actively-written ones (creates oversized segments the normal policy won't re-merge). — intro'd in lesson 0005
occurrence type
A bool clause key defining how a clause participates: must (match, scored), should (optional, scored, summed), filter (match, unscored, cached), must_not (exclude, unscored, cached). — intro'd in lesson 0006
more-matches-is-better
The bool scoring rule: the BM25 _score of every matching must and should clause is added together to form the final score. — intro'd in lesson 0006
aggregation
A summary computed over the documents matched by the query — counts, statistics, and groupings. Runs in the same request as the search; set top-level size: 0 to return only aggregations. — intro'd in lesson 0007
metric aggregation
An aggregation that calculates a value from field values across the docs — e.g. avg, sum, min/max, stats, cardinality. — intro'd in lesson 0007
bucket aggregation
An aggregation that groups documents into buckets (each a collection of docs meeting a criterion) by field value, range, or interval. Each bucket carries a doc_count. Examples: terms, range, date_histogram. — intro'd in lesson 0007
pipeline aggregation
An aggregation whose input is the output of other aggregations rather than documents — e.g. avg_bucket, derivative. — intro'd in lesson 0007
sub-aggregation (nesting)
An aggregation placed inside a bucket aggregation's own aggs, computing a metric or further buckets per bucket. The mechanism behind drill-down analysis. — intro'd in lesson 0007
terms aggregation
A bucket aggregation building one bucket per unique value (default top 10 by doc_count). Run it on a keyword field, not text. Counts are approximate across shards — see doc_count_error_upper_bound / sum_other_doc_count. — intro'd in lesson 0007
primary shard
An original shard of an index that owns a subset of its documents. The number of primaries is fixed at index creation and determines how the data is partitioned. — intro'd in lesson 0008
replica shard
A copy of a primary shard, for redundancy and read throughput. A replica is never allocated on the same node as its primary — so on a single node, replicas stay UNASSIGNED. — intro'd in lesson 0008
over-sharding
Having too many small shards. Because each shard carries fixed heap/thread/cluster-state overhead and a search uses one thread per shard, over-sharding degrades search performance and can destabilize the cluster. Target 10–50 GB per shard. — intro'd in lesson 0008
search thread pool
The bounded pool of threads a node uses to run shard-level searches — one thread per shard per query. Fanning a query across very many shards can exhaust it, the sharp edge of over-sharding. — intro'd in lesson 0008

← lesson 0001 · lesson 0008 →