CodeMesh Research
The short answer
CodeMesh Research is where we publish original, reproducible, transparently-documented research into AI coding economics and context engineering, not a content-marketing blog. It starts from one real result: a published 98-question benchmark measuring actual token and API cost differences between an agent reading a repository directly and the same agent using CodeMesh's structural retrieval.
See the full benchmark results → Read the methodology →
This hub is organized by research category so the roadmap is visible even where the work isn't done yet. Every card below is either a link to a real, published result, or a plainly-labeled topic we haven't completed research on. Nothing here is presented as finished when it isn't.
AI Coding Economics
Original research into what AI-assisted coding actually costs, and where that cost goes.
State of AI Coding Costs
An industry-wide look at what teams are actually spending on AI coding agents.
Tokens per Software Engineering Task
How token consumption scales with task type: bug fix, feature, refactor, review.
Cost per Successful Coding Task
Cost measured against successful outcomes, not just raw token counts.
Repository Size vs Token Consumption
How token consumption scales as a repository grows in size and complexity.
None of the four studies above have been run yet. The published benchmark already contains real, tier-by-tier token and cost data relevant to this category; see the full benchmark results for measured numbers by difficulty tier rather than a projection.
Context Engineering
Research into how much context an agent needs, and how supplying it well or badly changes outcomes.
Context Efficiency Study
How much of the context supplied to a coding agent is actually used in its final answer.
Structural vs Semantic Retrieval
A conceptual comparison of when an explicit relationship query wins and when semantic search wins. This is analysis and worked examples, not an empirical study with measured results.
Context Quantity vs Answer Quality
Whether supplying more context reliably improves answer quality, or where it stops helping.
Context Pollution Study
How irrelevant or excess context in the window affects a coding agent's accuracy.
Repository Intelligence
Benchmarks planned against different kinds of repository structure, not yet run.
Monorepo Retrieval Benchmark
Retrieval accuracy and token cost specifically in large monorepos with many packages.
Large-Codebase Retrieval Benchmark
How retrieval quality and cost hold up as a single codebase grows very large.
Cross-Language Repository Benchmark
Retrieval behavior in repositories that mix multiple programming languages.
Coding Agent Benchmarks
This is the one category with a completed, published entry, not a planned topic. It measures how a coding agent performs, in tokens, cost and accuracy, with and without CodeMesh.
- Published: 2026-08-13
- Sample size: 98 questions, across four difficulty tiers
- Repository: openclaw, a real, public codebase, chosen so results can be independently checked
- Model: the same underlying model in both arms, run through the Claude Code CLI; only the retrieval tool available to the agent changes between arms
- Limitations: one measured run, one repository, one model, one date; results vary by repository, model, coding agent and task; see the methodology writeup for the full set of caveats
The Full Benchmark
All 98 questions, sortable and unfiltered, including the results CodeMesh got wrong.
Methodology
How the repository, questions, models and measurements were chosen, and what this benchmark does not claim.
Read the one study that's real today
Every number in it is measured from a run artifact, not projected from a rate card.