24 min read
AI assisted

Can tgrep, CodeGraph, Mantic, or ripwire Replace ripgrep?

Four code search tools measured against ripgrep on eight repos

Point an agent at a codebase and it mostly calls rg. A few tools now offer to make that faster with an index, or to answer the question a different way entirely. I measured four of them against ripgrep on real repositories, and what decides each tool's success turns out to be a different kind of variable.

  • tgrep's speedup has nothing to do with repository size. kubernetes alone contains both 45.13x and 0.33x
  • CodeGraph is binary. All five repositories above 64% language coverage scored full marks; the two at 24% scored under half. There is no middle
  • Mantic doesn't read file contents. That's the design, not a marketing simplification
  • ripwire goes from 1 to 10 rank-1 answers on phrasing alone — the same 15 tasks, asked in natural language versus by symbol name

The tools are tgrep (trigram-indexed regex grep), CodeGraph (symbol and call-graph index), Mantic (path-based intent search), and ripwire (a symbol graph with no daemon). ripgrep is the baseline.

What was measured

Two machines measured independently without knowing about each other, then the results were merged. Three public repositories were added afterward to fill the scale gap.

Repositories Files
Run A 3 personal repos (Go+TS, SQL-heavy C, Python) 330 – 973
Run B an internal plugin repo, the Linux kernel 4,076 / 95,916
Run C vllm, litellm, kubernetes 6,963 / 10,628 / 31,261
Run D the same three, with ripwire added 6,963 / 10,628 / 31,261

Eight repositories, 48 regex cells, 40 retrieval tasks. A task gives the kind of question an agent would ask and checks whether the tool reaches the correct file; the answers were established by hand with ripgrep and by reading the code.

Before trusting any timing, the result sets were compared. Across the first 30 cells, matched line counts were identical.

The four answer different questions

Lining them up by speed is a category error to begin with. Mantic's own README says "Quick Text Searches → Use ripgrep."

tgrep CodeGraph Mantic ripwire
Question "which lines match this regex" "what breaks if I change this symbol" "which file should I open first" "what context does this job need"
Indexes trigrams of file contents file symbols and call relationships file paths and metadata file symbols and bodies (cached)
Returns matched lines (same as rg) symbols, call paths, blast radius ranked file paths ranked symbols and line spans
Resident process tgrep serve codegraph serve --mcp none none
Relation to ripgrep replacement complement complement complement

A trigram index slices file contents into three-character grams ahead of time; at query time it pulls literals out of the pattern and narrows the candidate set to files holding those grams. Everything that makes tgrep fast — and everything that makes it slow — comes from that.

tgrep — the query decides, not the repository

Six query stages of decreasing selectivity, run in the same order against each repository. Higher rows have fewer matches and more distinctive literals.

Query stage internal 4k vllm 7k litellm 10k kubernetes 31k linux 96k median
1 rare literal 12.56x 12.56x 13.10x 45.13x 206.43x 13.10x
2 -l file list 11.73x 9.81x 3.42x 4.62x 62.28x 9.81x
3 common literal 9.54x 2.37x 3.18x 1.55x 10.18x 3.18x
4 case-insensitive 4.25x 1.03x 0.66x 1.21x 8.28x 1.21x
5 anchored regex 7.67x 0.45x 0.51x 0.65x 1.83x 0.65x
6 no extractable literal 0.98x 0.25x 0.25x 0.33x 0.98x 0.33x

Read down a column and four of the five decrease monotonically. Knowing the query stage is enough to know the ordering of the speedup. The exception is the internal repo, where anchored regex sits above case-insensitive.

Read across and there's no rule. Ordered by repository size, the medians go 8.61 · 1.70 · 1.92 · 1.38 · 9.23x — up and down. kubernetes holds 45.13x and 0.33x at once. Size is not the variable.

Two mechanisms explain it. The index narrows the candidate set to save scanning, but with no literal to extract it can't narrow at all (stage 6), and when matches are plentiful the delivery cost survives the narrowing (stages 3–5). When both apply, the result isn't a tie — it's a 4x loss.

In the first two runs tgrep's worst case was 0.84x, essentially a tie. Adding the large repositories produced 0.25x. The difference is match count: [A-Z]{20,} returned 2,143 lines on Linux, while [A-Z]{8,} here returns 19,057 to 81,579.

"Up to 52x" did not reproduce

Within the Linux kernel alone, swapping only the query mix moves the aggregate multiple this much.

Query mix Aggregate
selective only (rare literal + -l) 90.9x
all six 3.4x
adversarial only (240K matches + no literal) 1.28x

And the 4,076-file repository and 95,916-file Linux produced aggregates of 3.42x and 3.36x — nearly identical. A 20x difference in file count with the same query mix gives the same multiple.

Break-even, and the index that isn't saved

The build cost has to be repaid in queries.

Repository Files Index build Saved per query Break-even
linux 95,916 11.8 s 1,643 ms 7.2 queries
internal plugin repo 4,076 0.84 s 66 ms 12.7 queries
personal repo (Go+TS) 873 140 ms 7.6 ms 18 queries
personal repo (Python) 973 280 ms 8.1 ms 35 queries

The smaller the repository, the farther the break-even — because ripgrep is already fast there.

The awkward case is having no index at all. Measured on Linux:

State Latency
server running 0.01 s
server stopped, index on disk 0.02 s
no index 16.3 s
ripgrep (reference) 1.75 s

With no index it's 9.3x ripgrep, and that 16-second build is not written to disk. The second run took 15.9 seconds and again produced no index directory. The same non-persistence reproduced on a small repository, where no .tgrep appeared. For an agent that touches a new repository every session, this is the default path.

It isn't a drop-in

In 46 of 48 cells ripgrep and tgrep returned the same answer. kubernetes broke the streak.

Cell ripgrep tgrep Δ
case-insensitive namespace 76,458 lines 75,757 lines 701
no literal [A-Z]{8,} 81,579 lines 81,371 lines 208

One file caused both: a 3.3 MB generated protobuf. tgrep classified it as binary by content inspection at index time and excluded it; ripgrep returns 1,156 matching lines from it. This is a difference in binary-file classification, not in search correctness. If you're looking for source, excluding it is the better behavior — but it does mean tgrep is not a semantic drop-in.

CodeGraph — language coverage is the whole story

Repository Primary language Text files Indexed Coverage Tasks
litellm Python + TSX 10,031 8,489 85% 5/5
personal repo Python 973 827 85% 5/5
personal repo Go + TS 873 745 85% 5/5
vllm Python 6,577 5,474 83% 5/5
kubernetes Go + YAML 30,933 19,834 64% 5/5
personal repo mostly SQL 330 80 24% 3/5
internal plugin repo mostly Markdown 4,076 992 24% 3/10

There is no middle. The five above 64% are perfect; the two at 24% score under half. Quality doesn't degrade gradually — the primary language is either on the supported list or it isn't. Markdown and SQL are not.

When it is on the list, the tool is dominant. It placed all 15 tasks across vllm, litellm, and kubernetes at rank 1, with a median output of 484 tokens. It's the only one of the four that pinpoints definition sites.

Symbol CodeGraph Actual
NewCloudNodeController node_controller.go:117 :117
NewHorizontalController horizontal.go:138 :138
LeaderElector leaderelection.go:188 match

This explains why CodeGraph managed only 16 of 25 tasks in the first two runs. Everything it missed was in an unsupported language — 2,262 Markdown files and 64 SQL files.

Indexing cost, though, diverges with scale.

Repository tgrep CodeGraph
vllm 6,963 0.86 s · 70 MB 17.5 s · 377 MB
litellm 10,628 1.86 s · 136 MB 27.3 s · 633 MB
kubernetes 31,261 2.85 s · 220 MB 60.4 s · 992 MB

On kubernetes it takes 21x longer and produces an index 4.5x larger, while covering 63% of the tree. On the Linux kernel (95,916 files), codegraph init had still not finished after 20 minutes. That's a single observation, and interference with the background run couldn't be ruled out.

Mantic — only where filenames give away the answer

Mantic doesn't read file contents. The README's "infers intent from file structure and metadata rather than brute-force reading content" is literal. PG_FUNCTION_ARGS appears in three C files, and Mantic returned nothing for it — while returning sql correctly, because those characters appear in the path.

Repository Answer reached Character
vllm · litellm · kubernetes 15/15 router.py, budget_manager.py, kubelet.go — the name is the answer
personal repo (Go+TS) 4/5 descriptive filenames
internal plugin repo 5/10 (+1 wrong) mostly Markdown
personal repo (SQL-heavy C) 1/5 filenames say less about function
personal repo (Python) 0/5 nested git repository

"Reached" means the answer file was somewhere in the results, not that it was ranked first. On the 15 large-repo tasks it reached the answer 15 times but ranked it first 5 times, and within the top five 10 times. That is a different axis from CodeGraph, which placed all 15 at rank 1.

That last row is a separate failure. The repository contained a nested repo with its own .git and no .gitmodules.

Tool Paths returned Under the nested repo
ripgrep · tgrep 391 372
CodeGraph 188 181
Mantic 27 0

No query returns any of them. Of 1,046 tracked files, 1,007 — 96% — are invisible. git ls-files reports a nested repository as a single gitlink entry, and Mantic doesn't descend into it. Pointing --path at it directly works, so only the default traversal fails, and it fails silently.

Cost also inverts with scale.

Small repos Large repos
Files returned 2 – 6 exactly 100, always
Median output tokens 12 1,038
15-task output total — 72,549 B

kubelet, scheduler, leader — every query returns exactly 100. It's a hard cap. The tool that was cheapest at 253 bytes on small repositories spends 2.7x CodeGraph's output (26,660 B) on large ones. The answer is inside those 100, but at a median rank of 5 you read from the top.

They go stale differently

For an indexed tool the real risk isn't speed. Four scenarios on a fixture repository. STALE-MISS means failing to see newly written code; STALE-PATH means returning a path that no longer exists.

Tool File modified File added File deleted File moved Watcher
ripgrep fresh fresh fresh fresh not needed
ripwire fresh fresh fresh fresh not needed
tgrep serve fresh fresh fresh fresh required
CodeGraph (daemon) fresh fresh fresh fresh required
tgrep (CLI alone) STALE-MISS STALE-MISS fresh STALE-MISS —
CodeGraph (no daemon) STALE-MISS STALE-MISS STALE-PATH STALE-PATH —
Mantic (git repo) fresh fresh STALE-PATH STALE-PATH —

ripwire gets its own section below. The other four first.

tgrep fails safely. It misses new code but never invents a path, because it narrows candidates with the index and then actually reads the files at query time. No file, no match.

CodeGraph fails dangerously. Called without a daemon it hands back deleted files and pre-move paths as answers. And codegraph status won't tell you the index is stale — across every run, not one staleness warning appeared. For an agent that's worse than an empty result: an empty result sends it looking elsewhere, a wrong one sends it reading the wrong file.

Both tools are only complete while a watcher process is attached, and that watcher isn't in the CLI. tgrep's lives in tgrep serve, CodeGraph's inside codegraph serve --mcp. Propagation took 283ms–7,105ms for tgrep and 0.7–2.0s for CodeGraph. Agents usually attach over MCP, so watched mode is the default there — but calling the CLI from a script or a hook is a different world.

The two runs disagreed completely about Mantic, so a three-file minimal reproduction settled it.

Condition Result after deleting beta
git repository alpha, beta, gamma — a ghost
git repository, after git add -A alpha, gamma — correct
not a git repository alpha, gamma — correct

Inside a git repository it uses git's file list and never checks whether the file is on disk. Additions show up immediately; deletions linger until staged. Test only additions after an edit and it looks always-correct.

A high rank and a correct answer are different things

On a task asking where the claude-opus-5 model database lives, both Mantic and CodeGraph promoted a plausibly named file. Checking it, claude-opus-5 appeared zero times; the real answer was a same-named file on a different path. Without a procedure that verifies answers by content, it would have been scored correct.

Operational risk — the daemon doesn't die

After deleting the fixture directory outright and recreating it, the CodeGraph daemon stayed alive for over three minutes, still answering from the old database of a directory that no longer existed. Every query against the recreated repository was wrong.

codegraph daemon is an interactive picker — you select an entry and press enter. There's no --stop-all, so a script can't stop it; you kill the pid yourself. The tool that manages staleness becomes a source of it.

Installation residue cost a round too. In all three repositories .codegraph was a broken symlink from an earlier install. codegraph status says only "Not initialized, run codegraph init," and init then dies with ENOENT ... mkdir. It's an unexplained loop until you delete the dangling link by hand.

The same shape showed up in a real deployment

code-yeongyu/oh-my-openagent shipped CodeGraph as a bundled MCP server across three harnesses, then removed it entirely on 2026-09-02 — PR #7644, 274 files changed, rg -il codegraph going from 241 files to none. Listing what it deleted from the OpenCode plugin, the PR body includes:

the built-in codegraph MCP, the codegraph-bootstrap hook, the config schema block, the tool-result init guidance, and the process-sweep family for its orphaned daemons.

The infrastructure built to reap orphaned daemons went out with the tool. Before that came a zombie-sweep evidence directory (2026-07-27) and a Windows timeout repair (2026-08-16).

The reason for removal appears nowhere. Reading the PR body in full, it says what was deleted, why the config-loader change shipped alongside it, and what QA was run. The rationale for the decision isn't there. So that operational history is circumstance, not cause — and reading past that distinction is reading past the evidence.

The loader change that shipped with it isn't CodeGraph's fault either. OmoConfigLayerSchema was .strict(), so one unrecognized key dropped an entire config layer; removing the codegraph key would have silently reset the settings of every user whose config still carried that block. That's a schema policy problem on the bundling side.

Version matters here. Those incidents are CodeGraph 1.5.0; these measurements are 1.6.0. What changed in between wasn't established, and the 1 GiB orphan storm did not reproduce.

ripwire — the only graph tool that stays fresh unwatched

That was the daemon story, and redhat-et/ripwire v0.5.0 aims squarely at it. A single C++23 binary with no runtime dependencies, leading with No API key. No embeddings. No index server. No daemon. The last sentence targets the axis of the two preceding sections, so it went onto the same three repositories. The release tarball was checked against its SHA-256, and only the binary was used — not the install script or the bundled agent hooks.

The no-daemon claim holds. It passed all four scenarios in both --no-cache and --cache (incremental) modes. The incremental mode was correct on deletes and moves too — exactly the two cells where CodeGraph without a daemon and Mantic hand back paths that no longer exist. Apart from ripgrep, it's the only tool that filled that table with no watcher.

The price is memory and latency

The README reports indexing its own repository in 0.25 s and 6.6 MB. Real repositories give different numbers.

Repository Phase Wall Peak memory Cache
vllm 6,963 cold map (no cache) 3.66 s 628 MB —
lean cache build 4.16 s 649 MB 58 MB
rich cache fill (first --for) 30.67 s 3,471 MB 280 MB
warm query 1.81 s 2,955 MB 280 MB
litellm 10,628 rich cache fill 25.10 s 3,530 MB 319 MB
warm query 2.60 s 3,123 MB 319 MB
kubernetes 31,261 rich cache fill 8.03 s 1,002 MB 115 MB
warm query 2.49 s 1,012 MB 115 MB

It peaks at 3.5 GB on a 6,963-file repository — 3.3x CodeGraph's 1.06 GB init peak. Nothing broke on a 24 GB machine, but there isn't much headroom either. kubernetes has 4.5x more files than vllm and its rich phase is cheaper; Go files being numerous but individually small would explain it, though that wasn't verified, so it stays an observation.

"No index server. No daemon." does not mean no index. There is a 115–319 MB cache file, and filling it the first time takes 8–31 seconds. What's absent is a resident process, not an index. That distinction is the point: with no resident process there's nothing to go stale and nothing to leave a zombie behind.

Accuracy depends on how you phrase it

The same 15 tasks, asked two ways: in natural language (the wording given to Mantic) and by symbol name (the wording given to CodeGraph). ripwire reports its own routing — routed: name-exact BM25 — query names a symbol — so the two paths really are different.

Tool Rank 1 Top 5 Reached Median output tokens Median latency
CodeGraph 15/15 15/15 15/15 484 248 ms
tgrep 13/15 14/15 15/15 11 18 ms
ripgrep 12/15 14/15 15/15 11 122 ms
ripwire (symbol name) 10/15 13/15 13/15 264 2,372 ms
Mantic 5/15 9/15 15/15 1,038 440 ms
ripwire (natural language) 1/15 5/15 10/15 2,166 2,398 ms

Phrasing alone moves rank-1 answers from 1 to 10. Natural language sends --for down a subtoken+body BM25 path that spreads wide; a symbol name takes the name-exact path and narrows. Where Mantic is built to receive natural language, ripwire's --for is strongest when you already know the name.

Even in symbol mode it falls short of CodeGraph. Output is cheaper at 264 tokens against 484, but latency is 2,372 ms — 9.6x CodeGraph and 132x tgrep.

The two misses differ in kind. litellm keeps a Rust subproject beside the Python body; asked for Router, ripwire anchors on the Rust Router and puts six .rs files above the correct litellm/router.py. That isn't a coverage failure — a separate check returned another symbol from that same file at rank 1, so it was indexed. It couldn't pick which language the question meant. The other miss asked for proxy_server, which is a filename rather than a symbol, an unfavorable question for a symbol-graph tool. That one reads as a task design problem and wasn't counted against ripwire.

Markdown arrives through a different verb

CodeGraph's governing variable was language coverage, and not seeing Markdown left it at 0/5 on the md tasks. ripwire lists Markdown as supported, and that turns out to be true — but not through --for. On kubernetes (6,391 YAML files, 591 Markdown):

  • --for="contributing guidelines code review" → everything returned is .go. No documents
  • --recall="Contributor License Agreement" → 16 of 323 document files matched, returning CONTRIBUTING.md by section (lines="7-12,13-20")
  • the same query to CodeGraph → No results found

--for is the code lens and --recall is the document lens. Markdown headings become section symbols exactly as documented. CodeGraph has no counterpart at all, so in repositories whose product is documentation, ripwire fills a real gap — provided the agent knows which verb to call.

Verdict

If staleness is your main worry, ripwire is currently the only answer. Nothing else returns symbol- and call-graph-level answers while staying fresh with no watcher. CodeGraph's surviving daemon, Mantic's ghost paths, and tgrep's CLI false negatives are all absent here.

On every other axis an existing tool is better.

Axis Best ripwire
Rank-1 answers CodeGraph 15/15 10/15 (symbol name)
Query latency tgrep 18 ms 2,372 ms
Peak memory tgrep 121–159 MB 1.0–3.5 GB
Output tokens tgrep · rg 11 264

This is v0.5.0, released a day before the measurement. It's pre-1.0, and this run covers three repositories and 15 tasks.

Query breadth dominates tokens, not the tool

Measured on one task — understanding an emoji-blocking code path.

Method Bytes
Mantic --files 253
CodeGraph context --no-code 1,134
Mantic default JSON 2,712
CodeGraph context 4,064
reading the target file in full 5,920
ripgrep -l candidate list (154 files) 10,015
CodeGraph explore 20,609

Two things invert. An unranked -l list costs more than reading the answer file in full. On the worst task ripgrep spent 152,439 bytes listing 1,709 files; the answer file is 5,920 bytes. And CodeGraph explore is 3.5x the cost of reading the file, so the "fewer tokens" claim depends on which subcommand you use.

The improvement here needs no change of tool. Stop throwing a broad rg -l and reading the resulting pile. Narrow the pattern, or cap it with -m.

Checking the published claims

Claim Verdict
tgrep "up to 52x" Not reproduced. The multiple is a function of query selectivity, ranging 0.25x to 206x
Mantic "2-10x slower than ripgrep" Partly wrong. 5x slower on small repos, 4x faster on Linux
CodeGraph "fewer tokens" Inverts by subcommand. context --no-code 1,134 B, explore 20,609 B
CodeGraph "auto syncs on code changes" Conditionally true. The watcher lives in the MCP server, not the bare CLI
Mantic "100% multi-repo accuracy" Disproved on nested repos. It returned none of 1,007 files
ripwire "No index server. No daemon." True, but not the same as no index. A 115–319 MB cache, 8–31 s to fill

So what do you use

ripgrep, by default. Across 40 tasks only ripgrep and tgrep never missed the answer, and only ripgrep is still correct immediately after an edit. With no index and no daemon, there's no state to go stale or die.

Each tool needs a different question asked of it.

Tool Governing variable What to ask
tgrep query type what kind of queries do I actually throw
CodeGraph language coverage is my repository's primary language on the supported list
Mantic how descriptive paths are can you answer from filenames alone in this repo
ripwire how the question is phrased do I know the symbol name, or am I asking in prose

tgrep earns adoption when four conditions overlap: a large repository, sessions long enough to clear the break-even, queries dominated by rare literals or -l, and tgrep serve running continuously. That combination produced the 206x, and it's the same profile under which GitHub Copilot CLI adopted tgrep. None of the eight repositories here satisfied all four. If you already run it, there's reason to keep it — going stale, it misses code rather than inventing paths.

CodeGraph starts with "Files by Language" in codegraph status. If your primary extension isn't there, nothing else about the tool matters. If it is, the tool is strong: impact, callers, callees, and affected have no counterpart in the grep family. Attach it over MCP so the daemon runs, and if you call it from a script or a hook, run codegraph index first every time. Bundling it into a shipped product is a different bet.

ripwire covers the places a daemon can't go — CI, hooks, one-shot scripts, or anywhere you've decided not to pay the cost of managing a daemon. Where you can run and manage one, CodeGraph is more accurate and 9x faster. If you do run ripwire, phrase queries as symbol names and reach for --recall when looking for documentation. Check that you can afford 3.5 GB first.

Mantic isn't worth adopting, and the reason is a paradox. If filenames alone can answer the question, a person can answer it too and the tool is barely needed; if they can't, the tool gets it wrong. The 253-byte --files output is worth remembering, and it's usually rank 1 when right — but that doesn't outweigh silent blindness to nested repositories and confidently wrong answers.

What wasn't established

  • Cold cache. purge needs sudo, so both machines measured warm. Real first queries are slower
  • Per-MCP-session measurement. The central claim of both indexed tools — fewer tool calls — was approximated with CLI invocations instead
  • What agents actually query. The token accounting conclusion depends on it, and whether narrow patterns or broad keywords are closer to real use wasn't settled
  • Why CodeGraph's Linux index didn't finish. One observation; outside evidence raises the prior but doesn't prove cause
  • What changed between 1.5.0 and 1.6.0 regarding orphaned daemons. No reproduction attempt was designed for it
  • How ripwire behaves after 1.0. v0.5.0 was measured on three repositories and 15 tasks
  • Rank-1 counts for tools with no relevance ranking. Running the same 15 tasks twice gave ripgrep 12 and then 14 rank-1 answers (lite-02, k8s-04); tgrep gave 13 both times. For the grep family rank is traversal order, so don't read those counts precisely