How the CodeMesh benchmark works

The short answer

CodeMesh's published benchmark asked the same 98 software-engineering questions about the same public codebase, twice: once to an agent with only filesystem and grep access, once to an agent with only the CodeMesh code-answer tool attached. Every token is measured from the run artifact, every question is published, and the 11 the system got wrong are kept in the results rather than removed.

The two arms, side by side

Both answered the same questions about the same checkout with the same model. One had filesystem and grep access and no MCP tools; the other had only the CodeMesh tool. Everything measured below is the difference between these two columns.

trace_query.log · codemesh
WITHOUT CODEMESH
01
Search repo
02
Open file
03
Read thousands of tokens
04
Follow imports
05
Open related files
06
Search callers
07
Read callers
08
Read configuration
09
Reason about answer
WITH CODEMESH
01
Search structural graph
02
Identify implementation
03
Identify callers
04
Retrieve precise context
05
Send relevant context to model
06
Reason about answer
steps
9→
6
tokens
100%→
5%

Repository

The benchmark runs against openclaw, a real, publicly available open-source codebase. We deliberately chose a public repository rather than our own product so that the questions, the code they refer to, and the ground-truth answers can all be checked by someone with no relationship to CodeMesh. We haven't published a pinned commit hash or mirror URL here. If you want to verify a specific answer against the exact code, ask us and we'll point you at the checkout the run used.

Questions

The benchmark consists of 98 questions spread across four difficulty tiers:

  • Medium: 32 questions. Direct lookups: a single constant, timeout or file.
  • Hard: 40 questions. Require connecting two or three related pieces of code.
  • Extreme: 17 questions. Multi-hop reasoning across several files or subsystems.
  • Expert: 9 questions. Require explaining why the code is built the way it is, not just what a value is.

Questions were written before either arm was run, against the real code, and none were excluded afterward for any reason, including the ones CodeMesh answered incorrectly. All 98 questions that were written appear in the published results.

Models

Both arms of the benchmark run through the Claude Code CLI, and both arms use the same underlying model for the run. The only difference between the two arms is which tools the agent is allowed to use to find information in the repository, not the model answering the question.

Baseline workflow

Arm A: reading the files. A fresh session gets ordinary filesystem and grep access and no MCP tools at all. It answers the way most coding agents answer today: search for a likely file, open it, read it, follow an import or a caller, open another file, read more, and eventually assemble an answer from whatever it found along the way.

CodeMesh workflow

Arm B: CodeMesh. The same session, asked the same question, with exactly one tool attached: the CodeMesh code-answer MCP tool. Retrieval and synthesis against the parsed code graph happen server-side; only the resulting answer crosses back into the model's context, instead of the raw files it took to construct it.

Token measurement

Every token count in this benchmark is a real, measured usage number pulled from the run artifact for that question, not an estimate derived from file sizes or a token-per-character heuristic. That includes input tokens, output tokens, cache-read tokens, cache-creation tokens, and any tokens consumed by helper-model calls made along the way. The headline reduction figure is computed from total tokens consumed across all 98 questions, not an average of per-question ratios.

Cost measurement

The published cost figures are real measured API spend for this run, not a projection from a rate card. Because model API pricing can change over time, the run date matters for interpreting the dollar figures: this run was measured on 2026-08-13, and the cost numbers reflect the pricing in effect on that date.

Failures

11 of the 98 questions (11.2%) were answered incorrectly by the CodeMesh arm. Those 11 results are kept in the published data, not excluded or reworked after the fact. A benchmark that quietly drops its failures isn't measuring anything.

Outliers

No individual question was excluded as an outlier in either direction, not the ones with unusually large token savings, and not the ones with the weakest results. 92.9% of questions individually clear an 85% token-reduction floor; the remainder are shown alongside the rest rather than trimmed out to smooth the distribution.

Limitations

Read this before citing the numbers
This benchmark reflects one measured run, against one public repository, using one model, on one date. It is not a claim that every repository, every model, every coding agent or every task will see the same reduction. Individual questions in this same run range from under 80% to over 99% reduction. The published 95.1% is a total across 98 questions, not a floor guaranteed for any single task. Results on your repository, with your model and your usage pattern, will vary.

Reproducibility

Because openclaw is public, this benchmark can in principle be independently re-run by anyone with access to the same checkout, the same questions, and API access to run both arms. We have not packaged this into a one-click reproduction script: reproducing it today means running the same two workflows described above yourself. If that changes, this page will be updated.

Data

The complete 98-row result set, every question, both token counts, the reduction percentage, the difficulty tier and the verdict, is published in full on the benchmark page. Nothing summarized here replaces that table; it is the same data these numbers are computed from.

See the results this methodology produced

Every question, every token count, every verdict, published in full.

See Every TestExplore the Data