CodeMesh Proof & Benchmarks
The short answer
Marketing claims are only useful when they can be inspected. Every reduction figure CodeMesh publishes comes from one measured benchmark run: 98 questions, one public repository, real token counts and real API spend, and every one of those 98 results, including the ones CodeMesh got wrong, is published in full.
What the benchmark measured
Both arms answered the same questions about the same checkout. One had filesystem and grep access, the other had only the CodeMesh tool. This is the difference that produced the numbers below.
Results are from the published CodeMesh benchmark (98 questions, one public repository, one model). Actual results vary by repository, model, coding agent, task and usage pattern. Read the methodology →
It is easy for any product in this space to claim large token or cost savings. It is harder to show the underlying data well enough that a skeptical reader can check it themselves. This section exists for that reader: the full benchmark result set, a detailed account of how the benchmark was run, and an interactive way to explore every individual question.
Nothing here is a projection or a rate-card estimate. The token counts are measured from run artifacts, and the dollar figures are real measured API spend for the same 98 questions, asked twice: once with an agent reading files directly, once with an agent using only CodeMesh.
The Full Benchmark
All 98 questions, sortable and unfiltered, including the results CodeMesh got wrong.
Methodology
How the repository, questions, models and measurements were chosen, and what this benchmark does not claim.
Benchmark Explorer
Filter by difficulty tier or area, sort by tokens or reduction, and inspect any individual question.
Raw Results
Every one of the 98 rows behind these numbers is shown in full on the benchmark page. Nothing is held back.
Inspect Every Test
Read the questions, the tokens, and the results CodeMesh didn't get right.