The Complete Guide to AI Coding Cost Optimization

The short answer

AI coding cost is bigger than a model's per-token price: it includes subscriptions, metered API usage, supporting infrastructure and developer time. Within the usage-based slice, the biggest controllable lever isn't which model you pick: it's how much unnecessary context an agent has to search, read and re-read before it can do anything useful. This guide covers both: where the costs come from, and the concrete levers, prompt optimization, context optimization, caching, model routing and measurement, for controlling them, regardless of which coding agent or model you use.

Work it out as you read

The guide gives you the levers. This gives you the order to pull them in, on your own usage rather than on a benchmark.

50
$100
30%
MONTHLY SPEND
$5,000
ANNUAL SPEND
$60,000
ELIGIBLE SPEND (CONTEXT/RETRIEVAL)$1,500/mo
POTENTIAL SAVINGS
$1,427
PER MONTH
$17,118
PER YEAR

Estimates are illustrative and show potential savings on eligible context/retrieval usage only, not your entire AI bill. Actual savings depend on model pricing, coding-agent behavior, repository size, task type and the proportion of usage attributable to code retrieval and context loading.

1. What AI coding costs consist of

Ask most engineering leaders what AI coding costs their team, and they'll quote a number from one invoice: the per-seat subscription for a coding assistant, or last month's API bill. Both are real costs, but they're each only one piece of a larger picture that spans four distinct categories, usually tracked by different people, on different schedules, using different units.

Subscription or seat costs are flat, predictable fees, per developer, per month or year, for a coding assistant or IDE plugin. They're easy to budget because they don't vary with usage, which is also their limitation: a seat costs the same whether the developer uses it twice a day or fifty times. API or usage-based costs are the opposite: metered by tokens or requests actually consumed, whether through a direct model API key, a usage-based agent platform, or overage charges once a subscription's included usage runs out. This is the category that scales unpredictably with how a codebase and a team's AI usage grow, and the one most of this guide is about.

Infrastructure covers everything standing behind those two: compute for any self-hosted or fine-tuned models, hosting and storage for vector databases or embedding pipelines, and the cost of running agents inside CI or other automated pipelines. It's usually invisible on an AI tooling invoice because it shows up on a cloud bill instead, under a different cost center. And developer time, while it never appears as a dollar figure on any AI-related bill, is frequently the largest true cost of all: time spent waiting on slow completions, reviewing AI-generated changes that need correction, or re-prompting an agent that misunderstood the task the first time.

Cost typeHow it's typically billedWho usually owns the budgetExample
Subscription / seatFlat fee per user, per month or yearProcurement / engineering opsA per-seat coding-assistant license
API / usage-basedMetered by tokens or requests consumedEngineering or cloud/infra budgetDirect model API calls from an agent
InfrastructureCompute, storage and hosting for supporting systemsPlatform / infra teamVector databases, embedding pipelines, CI-invoked agents
Developer timeNot billed directly; shows up as hours, not dollarsNo single line item: everyone's budgetTime spent waiting on, reviewing and correcting AI output

A cost-optimization conversation that only looks at the API bill is looking at one of four buckets. The rest of this guide focuses mostly on the usage-based bucket, because it's the one that scales with codebase size and agent behavior, but a complete accounting of AI coding cost has to include all four.

2. How token pricing works

Model providers charge based on tokens processed, not per request and not per completed task. A token is a chunk of text: in English, roughly three-quarters of a word on average, though the exact split varies by model and by language, since tokenization is model-specific. Every call to a model, every question an agent asks, every file it reads, every response it generates, consumes some number of tokens, and that consumption is what gets billed.

Rates vary by model tier: larger, more capable "frontier" models cost meaningfully more per token than smaller, faster models built for simpler tasks, and providers regularly ship new model generations with their own rate cards. Rates also vary between input and output tokens (chapter 3). What matters conceptually, independent of whatever the specific numbers are on any given day, is the basic shape of the bill:

Cost per call  =  (input_tokens  × input_rate)
               +  (output_tokens × output_rate)

Total cost     =  sum of cost_per_call, across every call in a session

That shape has one important consequence: total cost is a function of two independent things, how many tokens you send and generate, and which model's rate you're paying. You can control cost by changing either one, and the two levers don't substitute for each other. A cheaper model still costs more than it needs to if it's fed ten times the context it actually needed; a precisely-scoped context still costs more than it needs to if it's always sent to the most expensive model available, even for trivial edits.

Why no prices here
This guide deliberately avoids quoting specific per-token prices. Provider rate cards change every time a new model ships, often every few months, so any dollar figure printed on this page would be stale within a short time. Check your provider's current pricing page for today's numbers; the structural points made here hold regardless of what those numbers are.

3. Input tokens vs. output tokens

Input tokens are everything sent to the model before it responds: the prompt, any code or files included as context, and the running conversation history. Output tokens are what the model generates in response: an explanation, a diff, a new function. Providers typically price these differently, and output tokens are usually the more expensive of the two, per token.

There are two structural reasons for that gap. First, generating output is autoregressive: the model produces one token, feeds it back in, produces the next, and repeats, which is inherently more expensive to compute than processing a batch of input tokens in parallel. Second, providers increasingly offer discounted rates for repeated or cached input (chapter 10), which lowers the effective cost of input further without touching output pricing at all, widening the gap between what an input token and an output token cost in practice.

For coding tasks specifically, this creates a counterintuitive result: input tokens are cheaper per unit, but they usually dominate the bill by volume. A coding agent's output, a diff, a patch, a short explanation, is typically small. Its input, the files it read, the context it assembled, the conversation history it carried forward, is typically large, especially once repository discovery (chapter 4) is included. That means the dollars spent on a typical coding session are driven far more by how much (and how relevant) the input context is than by how verbose the model's answer turns out to be. It also means that improving retrieval quality, sending the right input, not just less of it, is one of the highest-leverage places to intervene on cost, because it acts on the larger and more controllable side of the ledger.

4. Repository discovery

Before an AI coding agent can generate anything useful, it usually has to figure out where in a codebase the relevant code lives. In practice that means running searches, listing directories, opening candidate files, reading them, often in full, even when only a small part is relevant, following imports to related files, and sometimes repeating the whole process because the first guess didn't pan out. None of that is generation. All of it is input tokens.

flow
1Developer question
2Search the repository (grep / file search)
3Open candidate files
4Read full file contents
5Follow imports to related files
6Search again: first guess was incomplete
7Locate the actually-relevant code
8Only now: reasoning and generation begins

On a small, familiar codebase, this loop is short and cheap. On a large or unfamiliar one, it can run for many rounds, and every round re-reads material, re-sends conversation history, and adds tokens before the model produces a single line of the actual change. Discovery is a structural property of how most coding agents operate today: a loop of search, read, and search again, not a one-off inefficiency in any particular prompt. And because most billing views report "input tokens" as a single aggregate number, discovery's share of that number is usually invisible unless you go looking for it.

5. Code context

"Code context" is the information an AI coding agent draws on to complete a task: the functions and classes it's editing or calling, their direct dependencies, the callers that could break if it changes a signature, relevant tests, and enough of the surrounding codebase to match existing conventions. Ideally, an agent receives a bounded, task-specific slice of the repository, not the whole thing, and not too little of it either. See what code context actually covers for a fuller breakdown of the textual, structural, relational and operational information a repository can offer an agent.

What an agent typically receives, absent deliberate retrieval, is whatever its default discovery process happened to turn up: entire files opened because a search matched a substring somewhere inside them, tangential files pulled in because a keyword search couldn't distinguish relevance from coincidence, or, at the other extreme, as much of the repository as fits in the context window, included "just in case" it turns out to matter.

Both directions have a cost. Too little context makes an agent guess, and guessing tends to produce wrong output that needs another round to correct, which costs more tokens overall than getting it right the first time would have. Too much context means paying, in tokens, to have the model read past material that was never relevant, and, as the next chapter covers, can measurably reduce how well the model uses the part of the context that actually mattered. Getting this right is fundamentally a retrieval-quality problem, not a token-counting problem.

6. Context windows

Definition
A context window is the maximum number of tokens a model can take into account at once: its input, its own output, and in some architectures a hidden reasoning allowance, all counted against the same ceiling. Once a conversation's cumulative content exceeds that ceiling, the oldest or lowest-priority material has to be dropped, summarized, or truncated to make room.

Context windows have grown substantially across model generations, to the point that fitting an entire small-to-medium repository into one call is often technically possible. That has led to a natural assumption: a bigger window is strictly better, because there's more room. That assumption doesn't hold in practice.

Bigger isn't automatically better: recall degrades with long, noisy contexts.

Independent of whether a context window has run out of room, models tend to recall and reason over information less reliably the longer and noisier the surrounding context becomes, particularly material buried in the middle of a long input, away from the beginning or the end. This is a distinct failure mode from simply running out of space. Running out of room is a hard, visible ceiling: the call fails, or the framework has to summarize and drop something. Degraded recall is soft and mostly invisible: nothing errors out, the model just silently misses or under-weights something relevant that was technically present in the context, and the response quietly gets worse and more likely to need a follow-up correction. A large window stuffed with mostly irrelevant material can underperform a small, tightly curated one, even though the large one had "more room to spare."

7. Context-window waste

Context-window waste is irrelevant, duplicated, or purely discovery-oriented content occupying space in the context window that could otherwise hold material the model actually needs, or that simply costs tokens without ever being used in the response. It's a different diagnosis from context-window size: a huge window with plenty of headroom can still be mostly waste, and a small one can be used efficiently.

Common sources include: the same file re-read multiple times over the course of a conversation because the agent lost track of having already read it; verbose tool output that carries formatting or metadata far beyond what the task needs; whole files included when only one function inside them is relevant; and stale history from earlier, abandoned attempts at the same task that never gets cleared out. None of this is dramatic on its own: it's the accumulation across a long agentic session that turns it into a meaningful share of total spend. See the dedicated breakdown of context-window waste for a closer look at how it accumulates and how to audit for it.

8. Prompt optimization

Prompt optimization is about improving the instructions themselves: being specific about the desired behavior, format and constraints up front; giving acceptance criteria so the agent knows what "done" looks like; providing an example of the intended style when one exists; and structuring a multi-step ask clearly enough that it doesn't need to be untangled through back-and-forth. Its cost benefit comes from two places: fewer round trips (each of which re-sends context and adds its own overhead), and less generation wasted on a misunderstood request that has to be redone.

It's worth being precise about how this differs from context optimization, covered next. Prompt optimization is about what you ask and how you phrase the ask. Context optimization is about what background material accompanies that ask. The two are independent: a perfectly worded, unambiguous prompt attached to sloppy or irrelevant context still wastes tokens on material the model didn't need. A precisely curated context attached to a vague or underspecified prompt still risks producing the wrong thing, which then needs another round to fix. They're complementary levers against cost, not substitutes for one another, and a team that has optimized one without the other has only captured half the available saving.

9. Context optimization

Definition
Context optimization is the process of minimizing the amount of information supplied to an AI model while preserving the information required to complete a task correctly. For coding agents, this means providing relevant functions, classes, dependencies and code relationships instead of unnecessarily loading large portions of a repository into the context window.

Context optimization is the category most of this guide's middle chapters belong to: repository discovery, code context, context windows and context-window waste are all facets of the same underlying question: is the model getting the information it needs, without also getting a large amount it doesn't? Unlike prompt optimization (what you ask) or caching and model routing (covered next, which act on how a call is billed rather than on what's in it), context optimization acts directly on the content of the input itself. See what context optimization means for coding agents for the full treatment, including how it compares to each of the adjacent techniques in this guide.

CodeMesh is a context optimization layer for AI coding agents. It helps agents understand a codebase without repeatedly searching, opening and reading large amounts of source code, allowing them to work with more precise repository context and dramatically fewer unnecessary tokens. That's one concrete way of applying context optimization specifically to the retrieval side of a coding agent's workflow: parsing a repository's structure once and answering an agent's context needs directly from that understanding, instead of letting every session rediscover the codebase from scratch through search and re-reading. See what CodeMesh is for the full picture of how that works.

10. Caching

Many providers support some form of prompt or context caching: a mechanism for marking a stable portion of the input, a system prompt, a tool schema, a large file that doesn't change between calls, so that a later call reusing that same prefix is billed at a fraction of the normal input rate on a cache hit, with a smaller separate cost for writing the cache the first time. This helps precisely the kind of workload an agentic coding session generates: the same system prompt, tool definitions, and often the same file context get resent, verbatim, call after call within one conversation.

Technical note
Two related things both get called "caching." Provider-side prompt caching lets you mark a stable prefix so a repeat call reuses it at a discounted input rate: this is the one that shows up as a lower number on your bill. Separately, some agent frameworks cache their own retrieval or embedding results client-side to avoid recomputing them. Both reduce repeated-context cost, but only the provider-side version changes what you're actually billed per token.

Caching has real limits worth being clear-eyed about. It helps most when context is stable and reused across many calls within a fairly short window, since cache entries typically expire. And it's a mitigation for the cost of resending the same material, not a way to make that material smaller or more relevant in the first place. A cached hundred-thousand-token context is still a hundred-thousand-token context; caching just makes resending it cheaper, it doesn't make it more useful. That's why caching combines so well with context optimization rather than substituting for it: the smaller and more stable the genuinely relevant context is, the more complete, and more valuable, each cache hit becomes.

11. Model routing

Model routing means sending different tasks to different models based on how demanding they are: cheaper, faster models for classification, simple edits, formatting and boilerplate, and reserving frontier or flagship reasoning models for genuinely hard problems like ambiguous requirements, complex refactors, or multi-file architectural changes. Done well, it means a team isn't paying frontier-model rates for work a much cheaper model could have handled just as correctly.

Model routing is conceptually independent of context optimization, and the two don't overlap in what they act on. Routing chooses which model answers a call; context optimization decides what that call contains, regardless of which model receives it. A team can route intelligently and still hand every model a badly-curated, oversized context; equally, a team can curate context precisely and still default to calling the most expensive model out of habit for tasks that never needed it. Controlling cost across the full stack means addressing both: neither one is sufficient by itself.

Scope note
CodeMesh does not perform model routing: it does not choose which model handles a request. It is a context layer that sits underneath whatever model or coding agent a team has already chosen, including a setup that already does its own model routing. The two are complementary, not the same capability.

12. Structural code retrieval

Definition
Structural code retrieval finds relevant source code using explicit software relationships rather than relying only on textual or semantic similarity. It can retrieve information based on relationships such as function calls, imports, dependencies, inheritance and containment, helping an AI coding agent identify relevant context more directly.

Rather than guessing relevance from text similarity, structural retrieval answers questions like "what calls this function," "what does this class inherit from," or "what breaks if this signature changes" directly, by walking a repository's actual dependency structure. See the dedicated page on structural code retrieval for how that compares in practice to text and vector-based search.

CodeMesh's approach (chapter 9) is one implementation of this idea: it parses a repository into a structural representation ahead of time, so that when an agent needs to know a function's callers, a class's dependents, or a module's imports, that answer can be looked up directly instead of reconstructed through search and re-reading every time it's needed.

13. Semantic retrieval

Semantic (or vector) retrieval represents chunks of code and text as embeddings, and finds relevant material by similarity between an embedding of the query and embeddings of the corpus. It's well suited to fuzzy, natural-language questions where the right code doesn't share vocabulary with the question, finding "where we handle payment retries" even if the word "retry" never appears in the relevant code, and to broad exploration across an unfamiliar codebase, including non-code material like documentation and comments.

Its limitation is that similarity isn't the same thing as relevance or completeness. Two pieces of code can be tightly coupled: a configuration constant and the fifteen call sites that read it, while reading nothing alike, in which case a vector search can miss the connection entirely. Conversely, code that reads similarly may have no real relationship at all. Vector search returns what's similar; it doesn't guarantee it has returned what's actually required for a change to be correct.

Semantic (vector) retrievalStructural retrieval
Matches byMeaning / similarity of textExplicit relationships: calls, imports, inheritance
Best forFuzzy, natural-language questions; unfamiliar territoryPrecise dependency, caller and impact questions
Can missCode that's related but doesn't read similarlyCode that's conceptually similar but not directly connected
NeedsAn embedding index, kept in sync with the repositoryA parsed structural representation of the repository

See structural retrieval vs. vector search for a fuller side-by-side. In practice, the two are often complementary rather than competing: semantic search to find the likely area of a large, unfamiliar repository, then structural traversal from there to pull in the exact dependencies once the right neighborhood is found.

14. Code graphs

Definition
A code knowledge graph represents entities in a software repository and the relationships between them. Nodes may represent files, functions, classes or modules, while relationships may represent calls, imports, inheritance, containment or dependencies.

A code graph is what makes structural retrieval (chapter 12) possible in practice: instead of re-deriving relationships by parsing and cross-referencing files on the fly, an agent, or the system supporting it, can traverse relationships that were already resolved when the repository was parsed. That's the difference between answering "what calls this function" by searching text for the function's name across every file, versus following an edge that already encodes every caller. See what a code knowledge graph is for how it compares to a plain AST, a dependency graph, or a call graph.

15. Retrieval-augmented generation (RAG)

Retrieval-augmented generation is the general pattern of retrieving relevant material from an external corpus and inserting it into a prompt before generation, so a model can answer using material beyond whatever it already knows or beyond what fits by default in its context. RAG as a pattern is agnostic about how retrieval happens: the retrieval step could be keyword search, semantic/vector search, or structural traversal of a graph. In practice, "RAG" for code has usually meant the semantic version: chunk source into embeddings and retrieve by similarity (chapter 13).

Structural code retrieval relates to RAG without being identical to it. Both are strategies for avoiding "stuff the whole repository into context and hope": retrieving a bounded, relevant slice instead of everything. Where they differ is the signal used to decide what counts as relevant: typical RAG implementations select by textual or semantic similarity; structural retrieval selects by explicit code relationships. The two can also be combined, as noted in chapter 13: using semantic search to narrow down a broad or unfamiliar codebase, then structural traversal to pull in the exact, correct dependencies once the relevant area has been found. See code graph vs. RAG for a closer comparison.

16. MCP (Model Context Protocol)

The Model Context Protocol is an open standard for how AI agents connect to external tools and context sources. Instead of every coding agent needing a bespoke integration for every filesystem, database, search index, ticket tracker or code graph it might want to query, a tool can expose itself once, over MCP, and any MCP-compatible agent can reach it using the same interface.

Technical note
Technically, MCP defines a client-server protocol, typically running over stdio or HTTP, through which a coding agent (the MCP client) can list and call "tools" and read "resources" exposed by an MCP server. A code-graph query engine, a filesystem, a ticket tracker or an internal API can each expose themselves as MCP servers using the same interface, so an agent that speaks MCP can reach any of them without a custom integration built per tool.

MCP is worth understanding in a cost-optimization guide because it's become the practical, standard way agents reach context providers of any kind, but it's important to be precise about what it does and doesn't guarantee. MCP standardizes how an agent reaches a context source; it says nothing about what that source returns. A well-built MCP server that performs precise, structural retrieval will reduce the tokens an agent spends on context. A poorly built one that dumps raw file contents over the same protocol will not: MCP is the transport layer, not a retrieval-quality guarantee. See how CodeMesh integrates over MCP for the specific case.

17. Measuring developer AI expenditure

The most natural starting metric for AI coding spend is dollars spent per developer per month, across subscriptions and metered API usage combined. It's a reasonable baseline for three reasons: it's usually already visible on a bill (seats and API keys are commonly tied to individuals or teams), it scales in a way that's roughly comparable across a team of similar size, and tracking it monthly reveals a trend: whether spend per developer is growing faster than headcount, usage, or any measure of output.

Its limitation is that it says nothing about whether the spend produced anything. A developer spending very little because they rarely use AI tooling and a developer spending a great deal because they use it constantly and productively both just look like different numbers on this metric: it can't distinguish light usage from wasteful usage, or heavy usage from productive usage. That's the gap the next chapter's metric is meant to close.

18. Cost per completed coding task

A more meaningful unit than raw spend, or raw tokens, or cost per API call, is cost per successfully completed unit of work: a merged pull request, a passing test suite, a resolved ticket. The distinction matters because cost per call, on its own, doesn't penalize a wrong answer at all; it only penalizes it indirectly, through however many additional follow-up calls are needed to correct it, and only after the fact.

Cost per completed task  =  sum of (tokens × rate) across every round
                            until the task is verified done
                          ÷  number of tasks actually completed

This reframes efficiency correctly: an approach that's cheaper per individual call but needs twice as many rounds to reach a correct result can easily end up more expensive per completed task than an approach that costs more per call but gets it right the first time. Counting only the numerator, tokens or dollars per call, makes the first approach look cheaper right up until you divide by whether it actually finished the job.

19. Team-wide AI economics

Per-developer costs multiply across a team, but not in a simple, linear way. Usage varies widely between individual developers on the same team: heavy users and light users on identical tooling can differ by an order of magnitude in monthly spend, and different teams within the same organization may standardize on different tools or models entirely. Shared infrastructure costs (a shared vector index, a shared code-graph service) also amortize differently across a team than per-seat subscription costs do, which means a purely per-developer view can understate or overstate the real total depending on how much of the stack is shared versus individually metered.

A team-wide view should track total spend, cost per active AI-using developer, not just cost divided by headcount, since adoption is rarely 100%, and how that figure moves as the team or its codebases grow. That last point matters more than it first appears: because repository discovery (chapter 4) tends to consume more tokens per task as a codebase gets larger, older, or more tangled, team-wide AI economics don't stay flat even at constant headcount: a team working in a codebase that's growing in size and complexity should expect its per-developer AI spend to drift upward over time unless something is done specifically about context and discovery efficiency. See how this plays out for engineering teams in more detail.

20. The CodeMesh benchmark

To make the retrieval side of this guide concrete rather than purely conceptual, CodeMesh publishes a benchmark of 98 developer questions against openclaw, a public open-source codebase, answered two ways by the same coding agent: once with ordinary filesystem search and file reading, and once with only a structural, context-optimized retrieval tool available instead. Because the repository is public, the questions, the ground truth, and the token counts are all independently checkable.

25.3M → 1.23M
Tokens, 98 questions
95.1%
Fewer tokens
$23.67 → $1.65
Measured API spend
93.1%
Lower measured cost

The result is a direct illustration of the ideas in chapters 4 through 14: most of the difference between the two conditions came from how much of the 25.3M baseline tokens were spent on repository discovery, searching, opening and reading files, rather than on reasoning about the actual question.

Results are from the published CodeMesh benchmark (98 questions, one public repository, one model). Actual results vary by repository, model, coding agent, task and usage pattern. Read the methodology →

21. Optimization checklist

Pulling the levers from this guide together into something actionable, regardless of which model, coding agent, or MCP tools a team already uses:

  • Prompt: Give explicit acceptance criteria and format up front, so fewer clarification rounds are needed to reach a usable answer.
  • Context: Prefer structural or scoped retrieval so an agent gets the relevant slice of a repository, not whole files pulled in by a keyword match.
  • Context: Periodically audit long agent conversations for duplicated file reads and stale history from abandoned attempts: context-window waste accumulates quietly.
  • Context: Don't assume a bigger context window is automatically safer: a smaller, curated context can outperform a larger, noisy one.
  • Caching: Keep system prompts, tool schemas and frequently reused file context stable across calls so cache hits stay high within a session.
  • Model routing: Reserve frontier or flagship models for genuinely hard reasoning tasks; route boilerplate, formatting and simple edits to smaller, cheaper models.
  • Measurement: Track dollars spent per developer per month as a baseline, but don't stop there.
  • Measurement: Track cost per completed task, not just cost per call or per token, since it's the number that reflects whether spend bought working software.
  • Measurement: Review the team-wide trend monthly rather than treating a single snapshot as a steady state, especially as codebases grow.

22. Calculator

Reading about how these costs stack up is one thing; seeing what they add up to for a specific team is another. The cost calculator applies the token-pricing and team-economics concepts from this guide, discovery overhead, seat count, usage patterns, to a team's own numbers, to give a rough estimate of where spend is going and what a reduction in unnecessary context consumption would be worth.

See what this would save you

Use the calculator to turn the concepts in this guide into a rough number for your own team.

Executive summary

AI coding cost is bigger than the number on an API invoice: it's made up of subscription seats, metered API usage, supporting infrastructure, and developer time spent waiting on, reviewing and correcting AI output, and most conversations about "AI cost" only look at the first two. Within the usage-based slice, providers charge per token, with output priced higher than input and rates that vary by model tier, but for coding agents, token volume is usually dominated by input: the code, files and history an agent reads before it generates anything. A large share of that input is repository discovery, search, open, read, re-read, that happens before real reasoning starts, and a large share of it is waste: irrelevant, duplicated, or discovery-only content sitting in the context window.

Prompt optimization, context optimization, caching and model routing are four distinct, complementary levers against that spend; none substitutes for the others. What ties them together is measurement: tracking raw spend per developer is a reasonable starting point, but cost per completed task, not per call, and not per raw token, is the number that actually reflects whether spend is buying working software. Structural retrieval, code graphs, semantic search and MCP are, in different ways, tools for getting an agent the right context instead of simply more of it, which, for most teams, is still the highest-leverage place they haven't looked yet.

Put context optimization into practice

CodeMesh applies structural context optimization to AI coding agents: see how it measured up in a published, reproducible benchmark.

Start Saving TokensSee the Benchmark