CodeMesh Research

The short answer

CodeMesh Research is where we publish original, reproducible, transparently-documented research into AI coding economics and context engineering, not a content-marketing blog. It starts from one real result: a published 98-question benchmark measuring actual token and API cost differences between an agent reading a repository directly and the same agent using CodeMesh's structural retrieval.

25.3M → 1.23M
Tokens, 98 questions
95.1%
Fewer tokens
$23.67 → $1.65
Measured API spend
93.1%
Lower measured cost

See the full benchmark results → Read the methodology →

This hub is organized by research category so the roadmap is visible even where the work isn't done yet. Every card below is either a link to a real, published result, or a plainly-labeled topic we haven't completed research on. Nothing here is presented as finished when it isn't.

AI Coding Economics

Original research into what AI-assisted coding actually costs, and where that cost goes.

Planned, not yet published

State of AI Coding Costs

An industry-wide look at what teams are actually spending on AI coding agents.

Planned, not yet published

Tokens per Software Engineering Task

How token consumption scales with task type: bug fix, feature, refactor, review.

Planned, not yet published

Cost per Successful Coding Task

Cost measured against successful outcomes, not just raw token counts.

Planned, not yet published

Repository Size vs Token Consumption

How token consumption scales as a repository grows in size and complexity.

None of the four studies above have been run yet. The published benchmark already contains real, tier-by-tier token and cost data relevant to this category; see the full benchmark results for measured numbers by difficulty tier rather than a projection.

Context Engineering

Research into how much context an agent needs, and how supplying it well or badly changes outcomes.

Planned, not yet published

Context Efficiency Study

How much of the context supplied to a coding agent is actually used in its final answer.

Published, comparison article

Structural vs Semantic Retrieval

A conceptual comparison of when an explicit relationship query wins and when semantic search wins. This is analysis and worked examples, not an empirical study with measured results.

Planned, not yet published

Context Quantity vs Answer Quality

Whether supplying more context reliably improves answer quality, or where it stops helping.

Planned, not yet published

Context Pollution Study

How irrelevant or excess context in the window affects a coding agent's accuracy.

Repository Intelligence

Benchmarks planned against different kinds of repository structure, not yet run.

Planned, not yet published

Monorepo Retrieval Benchmark

Retrieval accuracy and token cost specifically in large monorepos with many packages.

Planned, not yet published

Large-Codebase Retrieval Benchmark

How retrieval quality and cost hold up as a single codebase grows very large.

Planned, not yet published

Cross-Language Repository Benchmark

Retrieval behavior in repositories that mix multiple programming languages.

Coding Agent Benchmarks

This is the one category with a completed, published entry, not a planned topic. It measures how a coding agent performs, in tokens, cost and accuracy, with and without CodeMesh.

  • Published: 2026-08-13
  • Sample size: 98 questions, across four difficulty tiers
  • Repository: openclaw, a real, public codebase, chosen so results can be independently checked
  • Model: the same underlying model in both arms, run through the Claude Code CLI; only the retrieval tool available to the agent changes between arms
  • Limitations: one measured run, one repository, one model, one date; results vary by repository, model, coding agent and task; see the methodology writeup for the full set of caveats
Where this hub stands today
CodeMesh Research currently consists of one published benchmark and one published comparison article. Every other topic on this page is a planned area of research, explicitly marked as such, not a finished study with results held back. This hub will grow only as we complete and publish new reproducible research, not by adding topics before the work behind them exists.

Read the one study that's real today

Every number in it is measured from a run artifact, not projected from a rate card.

See Every TestRead the Methodology