Why do AI coding agents use so many tokens?

The short answer

AI coding agents can consume large numbers of tokens because solving a software-development task often requires more than generating code. The agent may first search the repository, open files, read source code, follow dependencies, inspect callers, review configuration and collect enough context to reason about the task. Repeated repository discovery can therefore consume substantial context before useful generation begins.

Fourteen files, one question

The bars are files in a repository and the lit ones are the files an agent opens to answer one question. Discovery is the cost of finding which bars to light, and it is paid again on the next question.

Where the tokens actually go

It's tempting to assume that AI coding costs are mostly the cost of the model "writing code." In practice, generation is often only the last step in a longer chain. Before an agent produces a diff, it may need to figure out which files are relevant, open them, read the code inside them, follow imports and calls to related code, check configuration, and assemble all of that into context the model can reason over.

The cost isn't only generation. It's discovery.

Each step below can cost tokens on its own, and on an unfamiliar or large repository, the discovery steps can add up to more context than the reasoning and generation that follow them.

flow
1PROMPT
2Repository search
3File reads
4Related files
5Dependencies
6Context assembly
7Reasoning
8Code generation

Two kinds of token consumption

Not all token consumption is the same problem. It helps to separate tokens spent doing the actual work from tokens spent figuring out what the actual work requires.

Useful model consumption

  • Reasoning about the task
  • Generating code
  • Debugging an issue
  • Planning a multi-step change
  • Explaining a decision or a piece of code

Potentially reducible consumption

  • Repeated discovery of the same parts of a repository
  • Repeated file reads across turns or sessions
  • Irrelevant files pulled into context
  • Duplicate context sent more than once
  • Overly broad repository scans
  • Unnecessary supporting code that doesn't affect the task

CodeMesh focuses primarily on the second category. It doesn't change how a model reasons, debugs or generates code: it targets the repeated, avoidable discovery work that can happen before any of that begins.

Common questions

It needs enough context to reason correctly: function signatures, related code, dependencies and conventions, before it can safely generate or modify code. Without that context, generated code is more likely to be wrong or inconsistent with the rest of the repository.
No. A larger context window can hold more content, but every token placed into it is still processed and billed. A bigger window doesn't make discovery more targeted: it just raises the ceiling on how much unnecessary content can be loaded before the limit is hit.
On most current model pricing, yes: input tokens are typically priced lower per token than output tokens, and cached input tokens are cheaper still. Exact ratios vary by provider and model, so this isn't a fixed rule across every API.
A larger or less familiar repository gives an agent a bigger space to search before it locates relevant code. That can mean more files opened, more dependencies followed and more re-reading, all before useful reasoning starts.
Caching can reduce the cost of re-sending identical, previously-seen content, which helps with repeated requests. It doesn't help an agent discover new or different context for a new question, so it reduces repetition rather than discovery itself.
CodeMesh parses a repository once and maintains a structural understanding of it, so an agent can retrieve relevant functions, classes and relationships directly instead of repeatedly searching, opening and re-reading files from scratch.

Give your agents the context they need

Start Saving TokensSee the Benchmark