Why do AI coding agents use so many tokens?
The short answer
AI coding agents can consume large numbers of tokens because solving a software-development task often requires more than generating code. The agent may first search the repository, open files, read source code, follow dependencies, inspect callers, review configuration and collect enough context to reason about the task. Repeated repository discovery can therefore consume substantial context before useful generation begins.
Fourteen files, one question
The bars are files in a repository and the lit ones are the files an agent opens to answer one question. Discovery is the cost of finding which bars to light, and it is paid again on the next question.
Where the tokens actually go
It's tempting to assume that AI coding costs are mostly the cost of the model "writing code." In practice, generation is often only the last step in a longer chain. Before an agent produces a diff, it may need to figure out which files are relevant, open them, read the code inside them, follow imports and calls to related code, check configuration, and assemble all of that into context the model can reason over.
The cost isn't only generation. It's discovery.
Each step below can cost tokens on its own, and on an unfamiliar or large repository, the discovery steps can add up to more context than the reasoning and generation that follow them.
Two kinds of token consumption
Not all token consumption is the same problem. It helps to separate tokens spent doing the actual work from tokens spent figuring out what the actual work requires.
Useful model consumption
- Reasoning about the task
- Generating code
- Debugging an issue
- Planning a multi-step change
- Explaining a decision or a piece of code
Potentially reducible consumption
- Repeated discovery of the same parts of a repository
- Repeated file reads across turns or sessions
- Irrelevant files pulled into context
- Duplicate context sent more than once
- Overly broad repository scans
- Unnecessary supporting code that doesn't affect the task
CodeMesh focuses primarily on the second category. It doesn't change how a model reasons, debugs or generates code: it targets the repeated, avoidable discovery work that can happen before any of that begins.