The Index File Format
The experimental grep-index crate gives ripgrep a persistent on-disk index so repeated searches can identify relevant candidates without starting from an ordinary full traversal each time. The command-line feature is explicitly experimental and is grouped under indexing flags.
The format shown here is deliberately small at its public boundary: an index is a directory containing index.db, opened through an embedded database handle. The excerpts expose how the path is found and opened, plus how regexes become literal and n-gram queries; they do not expose the database’s table schema.
Sources: crates/index/src/index.rs:12-15, crates/index/src/index.rs:52-67, crates/index/src/index.rs:69-91
Core concepts
Index directory
An index directory is the filesystem location that contains the database file named index.db.
Sources: crates/index/src/index.rs:22-25
Index
Index is the in-memory wrapper that pairs the database path with a read-only or read-write database handle.
Sources: crates/index/src/index.rs:12-15, crates/index/src/index.rs:175-178
IndexDiscovery
IndexDiscovery decides where to find or create an index, using the current directory, an environment-variable name, and a directory name.
Sources: crates/index/src/index.rs:104-109, crates/index/src/index.rs:112-119
GramQuery
GramQuery is the logical query language produced from regex literals: a literal, an And of queries, or an Or of queries.
Sources: crates/index/src/literal.rs:16-20
Analysis
Analysis tracks the extracted query together with exact literals, possible prefixes, and possible suffixes.
Sources: crates/index/src/literal.rs:320-326
N-gram
An n-gram is a fixed-size byte window extracted from a literal, such as the three-byte windows of foobar.
Sources: crates/index/src/literal.rs:822-830, crates/index/src/literal.rs:870-890
How the on-disk index is located and opened
The crate treats a valid index as a directory whose index.db child is a file; this is the test used by Index::exists. The IndexBuilder carries two independent choices: whether to create the index and whether to open it for writing.
IndexBuilder::build branches on the creation flag. Creation calls build_create; otherwise, opening calls build_open.
pub fn build(&self, path: impl AsRef<Path>) -> anyhow::Result<Index> {
let path = path.as_ref();
if self.create {
self.build_create(path)
} else {
self.build_open(path)
}
}The important point is that the directory operation and database operation are separate: creation first creates the directory if needed, then joins INDEX_FILE_NAME and creates the database at that path. Opening similarly joins INDEX_FILE_NAME, then chooses Database::open or ReadOnlyDatabase::open according to the write flag.
Handle preserves that access mode after opening. Read transactions are available from either variant, but write access fails when the handle is read-only.
A caller normally starts discovery with IndexDiscovery::new, whose defaults are an unset working directory, the environment variable RIPGREP_INDEX_PATH, and the .ripgrep directory name. discover first uses the configured or process current directory, then gives the environment-variable path priority when it is non-empty; otherwise it walks ancestors and accepts the first .ripgrep location that either is being created or passes Index::exists.
The discovery path is therefore a search policy, not part of the database contents. It lets the same index.db format be reached through an explicit environment path or through a .ripgrep directory in the current directory or an ancestor.
The following architecture shows the boundary between discovery policy, index construction, and the opened database handle.
Evidence
- index-discoverycrates/index/src/index.rs:104
- index-discoverycrates/index/src/index.rs:121
- index-buildercrates/index/src/index.rs:33
- index-buildercrates/index/src/index.rs:43
- index-wrappercrates/index/src/index.rs:12
- database-handlecrates/index/src/index.rs:175
- database-handlecrates/index/src/index.rs:181
- index-dbcrates/index/src/index.rs:9
- index-dbcrates/index/src/index.rs:52
- index-dbcrates/index/src/index.rs:69
The call order is compact: discovery selects a path, the builder chooses creation or opening, and the resulting Index retains the selected handle.
Evidence
Sources: crates/index/src/index.rs:22-25, crates/index/src/index.rs:33-36, crates/index/src/index.rs:43-50, crates/index/src/index.rs:175-178, crates/index/src/index.rs:181-189, crates/index/src/index.rs:191-198, crates/index/src/index.rs:112-119, crates/index/src/index.rs:121-146
How regexes become index queries
Indexed searching needs a cheaper prefilter than running the complete regex against every file. The index crate parses a regex into HIR and walks its expression kinds in b, producing an Analysis; build_analysis then finalizes that analysis, and GramQueryBuilder::build returns its query.
The analysis distinguishes certainty from approximation. A literal becomes Analysis::exact_one; a bounded character class becomes an exact LiteralSet; an unbounded repetition becomes anything; and a repetition with a positive minimum makes an otherwise exact analysis inexact.
Analysis therefore has three literal collections with different roles:
| Field | Plain meaning |
|---|---|
query | The logical GramQuery used for candidate filtering. |
exact | Literals known to occur exactly in the analyzed expression. |
prefix | Prefix fragments retained after exactness is lost. |
suffix | Suffix fragments retained after exactness is lost. |
These fields are the state that Analysis combines and later simplifies into an n-gram query.
Concatenation combines adjacent literal HIR nodes before analysis, so a run of literal pieces is treated as one literal where possible. When concatenation contains inexact pieces, the analysis crosses suffixes with prefixes and intersects the component queries; this preserves only n-grams that remain necessary across the concatenation.
Alternation unions analyses. If one branch is exact and another is not, the exact branch is saved into the query before the analysis becomes inexact, while prefixes and suffixes are merged. This is why the output can express alternatives without pretending that every branch supplies one common literal.
The builder’s defaults are a three-byte ngram_size, a maximum literal-analysis length of 250, and a maximum character-class size of 10. These limits matter because very large classes return anything, and long concatenated analyses are replaced with anything once the configured length limit is exceeded.
The data movement from regex structure to candidate query is summarized below.
Evidence
- regex-hircrates/index/src/literal.rs:677
- analysiscrates/index/src/literal.rs:320
- analysiscrates/index/src/literal.rs:671
- literal-setscrates/index/src/literal.rs:505
- literal-setscrates/index/src/literal.rs:551
- gram-querycrates/index/src/literal.rs:16
- gram-querycrates/index/src/literal.rs:667
- ngramscrates/index/src/literal.rs:822
- ngramscrates/index/src/literal.rs:256
and_ngrams skips a literal set whose shortest literal is smaller than the configured n-gram size. Otherwise it extracts windows from each sufficiently long literal, unions the resulting alternatives, and intersects them into the query. The underlying ngrams iterator uses overlapping windows, which is why foobar yields foo, oob, oba, and bar at size three.
Sources: crates/index/src/literal.rs:677-748, crates/index/src/literal.rs:671-675, crates/index/src/literal.rs:667-669, crates/index/src/literal.rs:320-326, crates/index/src/literal.rs:795-816, crates/index/src/literal.rs:406-443, crates/index/src/literal.rs:386-404, crates/index/src/literal.rs:644-646, crates/index/src/literal.rs:256-271, crates/index/src/literal.rs:822-830, crates/index/src/literal.rs:870-890
How indexed search differs from ordinary regex search
An ordinary ripgrep search uses a Searcher to read bytes, applies a Matcher, and reports results to a Sink. The matcher interface can support substring or full regex implementations, while the searcher remains responsible for consuming source bytes.
The index path adds a candidate-selection stage based on the extracted GramQuery: literal analysis turns a full regex into necessary literal or n-gram constraints before the ordinary matching machinery is relevant. The shown material establishes that query construction and the ordinary searcher are separate layers; it does not show the database table lookup or the exact candidate-record encoding.
At the CLI level, indexed candidates are filtered by explicit globs, file types, hidden-file and depth settings, and maximum file size. Ignore files are not reapplied during indexed search, and queries that require transformed contents or every file fall back to ordinary search. If no index is found, ripgrep also performs an ordinary search.
That distinction explains the crate’s purpose: the index avoids rediscovering all possible files for eligible repeated queries, while the final regex search still needs the normal matcher/searcher behavior for correctness. The available excerpts show the query-side optimization clearly, but not the complete indexed candidate execution path.
Sources: crates/searcher/src/lib.rs:7-15, crates/matcher/src/lib.rs:3-12, crates/index/src/literal.rs:667-669, crates/index/src/literal.rs:256-271, crates/core/flags/defs.rs:3497-3502, crates/core/flags/defs.rs:3493-3497
How it connects
The command-line layer exposes index reads and writes only when the unstable-index feature is enabled; index_write creates discovery and index_read discovers an existing index. Read Index Integration and Feature Gating for how this feature enters a search.
The index crate sits beside ripgrep’s reusable crates, while the grep facade re-exports the matcher, printer, regex, and searcher crates. Read Repository and Crate Map for that workspace boundary.
Literal extraction is related to the default Rust regex implementation, which also derives optimized regexes from configured HIR literals. Read The Default Rust Regex Matcher for the normal matcher-side optimization.
Once candidates are selected, ordinary byte reading and match reporting remain the responsibility of the searcher and its sink. Read The Searcher Core for that execution path.
Sources: crates/core/flags/hiargs.rs:935-953, crates/grep/src/lib.rs:15-21, crates/regex/src/config.rs:146-156, crates/searcher/src/lib.rs:7-15
Key takeaways
Indexrepresents an index directory plus a read-only or read-write database handle.- The on-disk filename is index.db; the shown excerpts do not define its internal table schema.
IndexDiscoveryprefersRIPGREP_INDEX_PATH, then searches.ripgrepdirectories through ancestors.- Regex HIR is reduced into
Analysis, literal sets, n-grams, and aGramQuery. - Indexed search narrows eligible candidates first, while ordinary search applies the matcher through the searcher; unsupported indexed queries fall back.
Sources: crates/index/src/index.rs:12-15, crates/index/src/index.rs:175-178, crates/index/src/index.rs:52-67, crates/index/src/index.rs:112-119, crates/index/src/index.rs:121-146, crates/index/src/literal.rs:671-675, crates/index/src/literal.rs:677-748, crates/index/src/literal.rs:256-271, crates/core/flags/defs.rs:3497-3502, crates/searcher/src/lib.rs:7-15