BurntSushi/ripgrepUnlicense3fce3b5Report / request removal

The Searcher Core

Searcher is the searcher crate’s coordinator: it validates configuration, reads or receives bytes, chooses a search strategy, and passes matches and context to a Sink.

It exists to keep input handling separate from matching and output. The same core can search a file, an arbitrary reader, or an in-memory slice while preserving line positions, byte offsets, binary-data state, and context behavior.

Sources: crates/searcher/src/searcher/mod.rs:597-625, crates/searcher/src/searcher/mod.rs:727-765, crates/searcher/src/searcher/mod.rs:643-657, crates/searcher/src/searcher/mod.rs:769-795

Core concepts

Searcher

A Searcher owns search configuration, decoding state, a line buffer, and a dedicated buffer for complete multi-line haystacks.

Sources: crates/searcher/src/searcher/mod.rs:597-625

Sink

A Sink is the callback interface that receives matches, context lines, context breaks, binary-data notifications, and search lifecycle events.

Sources: crates/searcher/src/sink.rs:102-223

Line-oriented search processes complete line ranges from either a rolling reader buffer or a slice, then asks the core to match each line.

Sources: crates/searcher/src/searcher/core.rs:330-383, crates/searcher/src/searcher/glue.rs:38-51

Multi-line search treats the entire haystack as one in-memory slice because matches may cross line boundaries.

Sources: crates/searcher/src/searcher/mod.rs:597-625, crates/searcher/src/searcher/glue.rs:166-206

Memory-map choice

MmapChoice controls whether file searching may use a memory map; its default is disabled, while auto enables it.

Sources: crates/searcher/src/searcher/mmap.rs:22-24, crates/searcher/src/searcher/mmap.rs:49-51

How a file search chooses its input strategy

The file path is opened first, then search_file_maybe_path tries the memory-map path before choosing a complete-file or incremental-reader strategy.

Input strategy selection — How does a file search choose between mapped, complete, and incremental input?

Evidence

If MmapChoice::open is disabled, running on macOS, or unable to create a map, it returns no map and the caller falls back; a successful map is passed to search_slice.

The fallback depends on multi_line_with_matcher: a compatible multi-line search fills the entire file into multi_line_buffer and runs MultiLine; otherwise the file is handled through search_reader.

For a reader, search_reader builds a decoder and then chooses between MultiLine and ReadByLine.

The decisive condition is not simply the builder’s multi_line flag. multi_line_with_matcher returns false when multi-line is disabled, when the matcher’s line terminator agrees with the configured terminator, or when the matcher reports that the terminator cannot occur in a match.

A heap limit of zero permits only the memory-map strategy; configuration checking rejects the search when memory maps are disabled or unavailable.

Sources: crates/searcher/src/searcher/mod.rs:643-657, crates/searcher/src/searcher/mod.rs:678-714, crates/searcher/src/searcher/mmap.rs:65-115, crates/searcher/src/searcher/mod.rs:678-712, crates/searcher/src/searcher/mod.rs:896-915, crates/searcher/src/searcher/mod.rs:727-765, crates/searcher/src/searcher/mod.rs:151-185, crates/searcher/src/searcher/mod.rs:805-821

How bytes become matches and context

Incremental searching uses LineBufferReader over a reusable LineBuffer. The reader clears the buffer at construction, fills it from the input, exposes complete searchable bytes, and consumes bytes after the core advances.

The buffer keeps separate positions for the searchable end and the end of a possible partial line, so a read can retain an incomplete line until more bytes arrive. The fill loop reads more whenever no line terminator is available, rolls unconsumed bytes to the front, and can either quit or convert when binary data is detected.

ReadByLine::run begins the sink, repeatedly fills the reader, matches the current buffer, consumes remaining bytes when matching stops, and finishes with byte and binary offsets.

Incremental reader call order — How do bytes move from a reader to a sink?

Evidence

For a slice, SliceByLine::run checks binary data in an initial range, repeatedly calls line matching, computes a byte count, and finishes. Both slice and reader paths use Core::match_by_line, which selects a fast line search when eligible and otherwise uses the slow line search.

The fast path finds candidate lines, handles optional before and after context, advances to the line end, and calls sink_matched; it can switch to the slow path when configuration requires it. The slow path iterates with LineStep, strips the line terminator before matching, updates the position, and routes successful, after-context, or passthrough lines separately.

A match is sent as SinkMatch, which carries its bytes, absolute byte offset, optional line number, the containing buffer, and its byte range within that buffer. A sink can inspect the raw bytes, iterate its lines, and read offsets and line numbers through the match accessors.

Sources: crates/searcher/src/line_buffer.rs:219-225, crates/searcher/src/line_buffer.rs:256-258, crates/searcher/src/line_buffer.rs:274-276, crates/searcher/src/line_buffer.rs:294-323, crates/searcher/src/line_buffer.rs:406-472, crates/searcher/src/line_buffer.rs:480-493, crates/searcher/src/searcher/glue.rs:38-51, crates/searcher/src/searcher/glue.rs:117-131, crates/searcher/src/searcher/core.rs:170-183, crates/searcher/src/searcher/core.rs:385-427, crates/searcher/src/searcher/core.rs:330-383, crates/searcher/src/sink.rs:366-373, crates/searcher/src/sink.rs:379-381, crates/searcher/src/sink.rs:392-394, crates/searcher/src/sink.rs:401-403, crates/searcher/src/sink.rs:410-412

How context and multi-line boundaries are tracked

Before-context calculation starts at the last visited line, walks backward by the configured number of lines, then emits each selected line through the sink. After-context starts at the last visited line and emits lines until after_context_left reaches zero.

The core tracks last_line_visited, last_line_counted, absolute_byte_offset, and the current line number; count_lines updates line numbers only for bytes not already counted. When a rolling buffer is reused, roll retains enough preceding context and adjusts the absolute offset before matching resumes.

A context break is emitted only when context reporting is enabled, a previous line was sunk, and the next group is separated by a gap.

Sink::matched receives every reported match. Without multi-line search, the reported match spans exactly one non-empty line; with multi-line search, it may span multiple lines and contain multiple matches. Sink::context receives one context line at a time, while Sink::context_break marks a gap between non-contiguous context groups.

Search lifecycle — What states does a search pass through before it completes?

Evidence

Multi-line mode changes the boundary model: the complete haystack stays available, MultiLine::sink locates the whole line range for a match, merges adjacent or overlapping line ranges, and delays emission until the group is complete. This prevents adjacent matches on the same lines from causing the same line to be sunk more than once.

Sources: crates/searcher/src/searcher/core.rs:240-274, crates/searcher/src/searcher/core.rs:276-309, crates/searcher/src/searcher/core.rs:185-213, crates/searcher/src/searcher/core.rs:661-671, crates/searcher/src/searcher/core.rs:646-659, crates/searcher/src/sink.rs:102-223, crates/searcher/src/searcher/glue.rs:166-206, crates/searcher/src/searcher/glue.rs:208-258

How it connects

The search worker supplies a matcher, a Searcher, and a printer-facing sink; for JSON output, it creates a sink and calls searcher.search_path. The printer’s standard implementation consumes SinkMatch values to build its output representation.

For the broader execution path, continue with Search Worker and I/O. For matcher behavior, see The Matcher Trait and Internal Iteration. For output handling, see Output Printers.

Sources: crates/core/search.rs:403-409, crates/printer/src/standard.rs:900-914

Key takeaways

  • File searches try MmapChoice::open first, then choose complete-file multi-line search or incremental reader search.
  • Sink is the callback boundary for matches, context, binary notifications, and lifecycle events.
  • Incremental search retains partial lines and required context while advancing absolute offsets.
  • Multi-line search keeps the entire haystack in memory and groups adjacent match lines before sinking them.

Sources: crates/searcher/src/searcher/mod.rs:678-714, crates/searcher/src/sink.rs:102-223, crates/searcher/src/line_buffer.rs:294-323, crates/searcher/src/searcher/core.rs:185-213, crates/searcher/src/searcher/mod.rs:597-625, crates/searcher/src/searcher/glue.rs:208-258

Want this for your repos?

Try Angada AI Wiki