How the CodeMesh benchmark works
The short answer
CodeMesh's published benchmark asked the same 98 software-engineering questions about the same public codebase, twice: once to an agent with only filesystem and grep access, once to an agent with only the CodeMesh code-answer tool attached. Every token is measured from the run artifact, every question is published, and the 11 the system got wrong are kept in the results rather than removed.
The two arms, side by side
Both answered the same questions about the same checkout with the same model. One had filesystem and grep access and no MCP tools; the other had only the CodeMesh tool. Everything measured below is the difference between these two columns.
Repository
The benchmark runs against openclaw, a real, publicly available open-source codebase. We deliberately chose a public repository rather than our own product so that the questions, the code they refer to, and the ground-truth answers can all be checked by someone with no relationship to CodeMesh. We haven't published a pinned commit hash or mirror URL here. If you want to verify a specific answer against the exact code, ask us and we'll point you at the checkout the run used.
Questions
The benchmark consists of 98 questions spread across four difficulty tiers:
- Medium: 32 questions. Direct lookups: a single constant, timeout or file.
- Hard: 40 questions. Require connecting two or three related pieces of code.
- Extreme: 17 questions. Multi-hop reasoning across several files or subsystems.
- Expert: 9 questions. Require explaining why the code is built the way it is, not just what a value is.
Questions were written before either arm was run, against the real code, and none were excluded afterward for any reason, including the ones CodeMesh answered incorrectly. All 98 questions that were written appear in the published results.
Models
Both arms of the benchmark run through the Claude Code CLI, and both arms use the same underlying model for the run. The only difference between the two arms is which tools the agent is allowed to use to find information in the repository, not the model answering the question.
Baseline workflow
Arm A: reading the files. A fresh session gets ordinary filesystem and grep access and no MCP tools at all. It answers the way most coding agents answer today: search for a likely file, open it, read it, follow an import or a caller, open another file, read more, and eventually assemble an answer from whatever it found along the way.
CodeMesh workflow
Arm B: CodeMesh. The same session, asked the same question, with exactly one tool attached: the CodeMesh code-answer MCP tool. Retrieval and synthesis against the parsed code graph happen server-side; only the resulting answer crosses back into the model's context, instead of the raw files it took to construct it.
Token measurement
Every token count in this benchmark is a real, measured usage number pulled from the run artifact for that question, not an estimate derived from file sizes or a token-per-character heuristic. That includes input tokens, output tokens, cache-read tokens, cache-creation tokens, and any tokens consumed by helper-model calls made along the way. The headline reduction figure is computed from total tokens consumed across all 98 questions, not an average of per-question ratios.
Cost measurement
The published cost figures are real measured API spend for this run, not a projection from a rate card. Because model API pricing can change over time, the run date matters for interpreting the dollar figures: this run was measured on 2026-08-13, and the cost numbers reflect the pricing in effect on that date.
Failures
11 of the 98 questions (11.2%) were answered incorrectly by the CodeMesh arm. Those 11 results are kept in the published data, not excluded or reworked after the fact. A benchmark that quietly drops its failures isn't measuring anything.
Outliers
No individual question was excluded as an outlier in either direction, not the ones with unusually large token savings, and not the ones with the weakest results. 92.9% of questions individually clear an 85% token-reduction floor; the remainder are shown alongside the rest rather than trimmed out to smooth the distribution.
Limitations
Reproducibility
Because openclaw is public, this benchmark can in principle be independently re-run by anyone with access to the same checkout, the same questions, and API access to run both arms. We have not packaged this into a one-click reproduction script: reproducing it today means running the same two workflows described above yourself. If that changes, this page will be updated.
Data
The complete 98-row result set, every question, both token counts, the reduction percentage, the difficulty tier and the verdict, is published in full on the benchmark page. Nothing summarized here replaces that table; it is the same data these numbers are computed from.
See the results this methodology produced
Every question, every token count, every verdict, published in full.