CodeMesh Proof & Benchmarks

The short answer

Marketing claims are only useful when they can be inspected. Every reduction figure CodeMesh publishes comes from one measured benchmark run: 98 questions, one public repository, real token counts and real API spend, and every one of those 98 results, including the ones CodeMesh got wrong, is published in full.

What the benchmark measured

Both arms answered the same questions about the same checkout. One had filesystem and grep access, the other had only the CodeMesh tool. This is the difference that produced the numbers below.

trace_query.log · codemesh
WITHOUT CODEMESH
01
Search repo
02
Open file
03
Read thousands of tokens
04
Follow imports
05
Open related files
06
Search callers
07
Read callers
08
Read configuration
09
Reason about answer
WITH CODEMESH
01
Search structural graph
02
Identify implementation
03
Identify callers
04
Retrieve precise context
05
Send relevant context to model
06
Reason about answer
steps
9→
6
tokens
100%→
5%
25.3M → 1.23M
Tokens, 98 questions
95.1%
Fewer tokens
$23.67 → $1.65
Measured API spend
93.1%
Lower measured cost

Results are from the published CodeMesh benchmark (98 questions, one public repository, one model). Actual results vary by repository, model, coding agent, task and usage pattern. Read the methodology →

Token reduction, by question
98 questions · nothing excluded
≥98% (50x+)
4 questions
95–98% (20–50x)
29 questions
90–95% (10–20x)
45 questions
85–90% (6.7–10x)
13 questions
below 85%
7 questions

It is easy for any product in this space to claim large token or cost savings. It is harder to show the underlying data well enough that a skeptical reader can check it themselves. This section exists for that reader: the full benchmark result set, a detailed account of how the benchmark was run, and an interactive way to explore every individual question.

Nothing here is a projection or a rate-card estimate. The token counts are measured from run artifacts, and the dollar figures are real measured API spend for the same 98 questions, asked twice: once with an agent reading files directly, once with an agent using only CodeMesh.

Inspect Every Test

Read the questions, the tokens, and the results CodeMesh didn't get right.

Inspect Every TestRead the Methodology