Testing, Benchmarking, and Fuzzing
This page covers three verification tools around ripgrep: the Rust integration-test harness, the Python benchsuite runner, and the fuzz_glob fuzz target.
They exist to check observable command behavior, compare search performance across tools and workloads, and exercise glob matching with generated inputs.
Sources: tests/macros.rs:2-15, fuzz/Cargo.toml:21-24, tests/macros.rs:18-38, benchsuite/benchsuite:19-23, fuzz/Cargo.toml:10-12
Core concepts
Integration test case
An integration test case is a generated Rust test that runs once with the normal configuration and, when enabled, once with PCRE2. The rgtest macro creates the test, calls setup, and repeats the test through setup_pcre2 when the pcre2 feature is active.
Sources: tests/macros.rs:2-15
Temporary fixture
A temporary fixture is a per-test directory that supplies files, the working directory, and the selected regex-engine mode. Dir stores the executable root, test directory, and pcre2 flag.
Sources: tests/util.rs:58-67
Command wrapper
A command wrapper is the object that configures and evaluates a child rg process. TestCommand stores both the fixture and the underlying Command, while command sets the working directory, removes RIPGREP_CONFIG_PATH, fixes the path separator, and optionally adds --pcre2.
Sources: tests/util.rs:252-257, tests/util.rs:167-176
Performance corpus
A performance corpus is a named dataset used to compare command-line search tools under a defined workload. benchsuite distinguishes large-file and many-small-file corpora because their performance characteristics and relevance strategies differ.
Sources: benchsuite/benchsuite:19-23
Glob fuzz target
A glob fuzz target is a cargo-fuzz binary dedicated to exercising the globset crate with arbitrary inputs. The fuzz package declares libfuzzer-sys, enables globset’s arbitrary feature, and names the target fuzz_glob.
Sources: fuzz/Cargo.toml:7-12, fuzz/Cargo.toml:21-24
How the integration harness runs rg
The harness starts by creating a unique fixture directory and then building a command wrapper for that directory. setup performs those two operations directly. Dir::new obtains a counter value, derives the test executable root, creates a path below the ripgrep-tests directory, removes an existing directory if necessary, and retries directory creation before returning the fixture.
The command wrapper resolves the rg binary through CARGO_BIN_EXE_rg when available, otherwise falling back to a path relative to the test executable; when CROSS_RUNNER selects a cross runner, it places the binary behind that runner. It then launches from the fixture directory, disables the ambient RIPGREP_CONFIG_PATH, forces / as the path separator, and adds --pcre2 for PCRE2 fixtures.
The test body can populate the fixture with text or bytes before invoking the command. create delegates to create_bytes, while try_create_bytes joins the file name to the fixture directory, creates the file, writes all bytes, and flushes it.
The normal output path turns process output into text only after requiring success. stdout calls output, output calls raw_output and expect_success, and raw_output removes lines beginning with <jemalloc>: from stderr before returning the result.
Evidence
- test-macrotests/macros.rs:2
- fixturetests/util.rs:22
- fixturetests/util.rs:73
- test-commandtests/util.rs:167
- rg-processtests/util.rs:179
- output-checktests/util.rs:410
The sequence shows the harness boundary: test generation creates the fixture and command, the command resolves the binary, and successful output is checked before a test consumes it.
Assertions cover both successful and unsuccessful process behavior. assert_err rejects a successful status, assert_exit_code compares the process exit code, and assert_non_empty_stderr requires failure plus non-empty stderr; each diagnostic includes command, working directory, directory listing, status, stdout, and stderr.
For exact textual comparisons, eqnice and eqnice_repr compare expected and actual values and panic with labeled expected-versus-got output when they differ.
| Helper | What it verifies |
|---|---|
assert_err | The command did not report success. |
assert_exit_code | The command returned the expected numeric code. |
assert_non_empty_stderr | The command failed and produced stderr. |
eqnice | Human-readable outputs are equal. |
eqnice_repr | Debug representations are equal. |
These helpers make failures actionable by preserving process context and showing the differing output or status.
Sources: tests/util.rs:22-26, tests/util.rs:73-87, tests/util.rs:14, tests/util.rs:179-193, tests/util.rs:492-505, tests/util.rs:167-176, tests/util.rs:102-104, tests/util.rs:134-143, tests/util.rs:305-308, tests/util.rs:331-334, tests/util.rs:338-342, tests/util.rs:523-528, tests/macros.rs:2-15, tests/util.rs:410-439, tests/util.rs:345-364, tests/util.rs:368-386, tests/util.rs:389-408, tests/macros.rs:18-38, tests/macros.rs:41-61
How benchsuite structures performance comparisons
The benchsuite file is an executable Python 3 script whose stated role is comparing command-line search tools. Its corpus configuration includes English and Russian subtitle files, a Linux source-tree clone, and explicit URLs for the subtitle data.
The script accounts for locale because grep’s Unicode behavior affects performance. GREP_ASCII sets LC_ALL to C, while GREP_UNICODE sets it to en_US.UTF-8; the script also defines a constrained SIFT command for code-search comparisons.
Each benchmark requires its corpus, chooses a working directory, defines a pattern, and builds a Benchmark containing named commands. The default Linux literal comparison uses PM_RESUME and contrasts rg, ag, git grep, ugrep, and recursive grep, while the fairer literal comparison adds line numbers and includes an rg (mmap) variant.
The benchmark suite varies the search question rather than measuring only one command: it includes default literal search, fair literal search, case-insensitive literals, literals inside regular expressions, word matching, and Unicode-category matching.
To reproduce a comparison represented by this script, use the benchsuite/benchsuite runner with the corresponding corpus available and inspect the benchmark definition for its working directory, pattern, flags, environment, and competing commands. The shown excerpt does not include the script’s command-line entry point or result-reporting code, so those details cannot be specified here.
Sources: benchsuite/benchsuite:1-5, benchsuite/benchsuite:25-36, benchsuite/benchsuite:37-51, benchsuite/benchsuite:53-79, benchsuite/benchsuite:82-110, benchsuite/benchsuite:113-143, benchsuite/benchsuite:146-197, benchsuite/benchsuite:200-241
What the glob fuzz target exercises
The fuzz package is isolated from the main workspace through its own [workspace] declaration, is marked with cargo-fuzz metadata, and keeps debug information in release builds. Its only shown dependency under test is the repository’s globset crate with the arbitrary feature enabled.
The target is named fuzz_glob and is located at fuzz_targets/fuzz_glob.rs; the manifest disables ordinary tests and documentation generation for that binary. Therefore, the grounded local entry point is the fuzz_glob cargo-fuzz target, and the behavior it is intended to stress is globset handling of arbitrary generated data. The target source and its exact local invocation are not included in this pack, so no stronger claim about specific assertions, mutations, or command syntax is warranted.
Sources: fuzz/Cargo.toml:7-19, fuzz/Cargo.toml:10-12, fuzz/Cargo.toml:21-24, fuzz/Cargo.toml:7-12
How it connects
The integration harness exercises the built rg binary rather than calling the internal search crates directly, making it a boundary check for the command-line product. For CLI behavior and flag resolution, continue to CLI Entry and Flag Parsing.
The benchmark runner compares the command-line tool with other search programs and explicitly varies options such as line numbering, case sensitivity, Unicode patterns, and memory mapping. The search behavior behind those commands is described in Search Worker and I/O and The Searcher Core.
The fuzz target is connected directly to the reusable glob-matching crate through a path dependency and its arbitrary feature. For the matching implementation and its role in filtering, continue to The Globset Matching Engine and Gitignore and Override Matching.
Sources: tests/util.rs:167-176, tests/util.rs:179-193, benchsuite/benchsuite:82-110, benchsuite/benchsuite:113-143, benchsuite/benchsuite:53-79, fuzz/Cargo.toml:10-12
Key takeaways
rgtestcreates normal and optional PCRE2 integration-test runs.Dircreates isolated fixtures, andTestCommandlaunchesrgfrom them.- Output and exit-status helpers turn process behavior into readable test failures.
benchsuitecompares multiple tools across corpus types, locales, patterns, and flags.fuzz_globis the configured globset fuzzing entry point, but its target implementation and exact invocation are not shown.
Sources: tests/macros.rs:2-15, tests/util.rs:58-67, tests/util.rs:167-176, tests/util.rs:331-334, tests/util.rs:345-364, tests/util.rs:410-439, benchsuite/benchsuite:19-23, benchsuite/benchsuite:82-110, fuzz/Cargo.toml:21-24