What is structural code retrieval?

The short answer

Structural code retrieval finds relevant source code using explicit software relationships rather than relying only on textual or semantic similarity. It can retrieve information based on relationships such as function calls, imports, dependencies, inheritance and containment, helping an AI coding agent identify relevant context more directly.

The relationships retrieval walks

Structural retrieval is only as good as the relationships it has to follow. These are the ones CodeMesh stores: containment from repository to file, definition from file to function and class, and the calls and imports that join them.

RepositoryCodeFileCodeFileCodeFunctionCodeFunctionCodeClassCodeImportCodeChunkCodeChunkCodeChunk

Software isn't only text.

Software is a network of relationships.

Two functions can share almost no vocabulary and still be tightly coupled: one calls the other directly. Two functions can use very similar wording and have nothing to do with each other. Retrieval methods that only compare text or meaning can miss the first case and mistake the second. Structural retrieval works from the relationships themselves.

A worked example

Take the question: "What depends on validateToken()?"

Semantic retrieval
A semantic or text-similarity search looks for code whose wording resembles "validateToken": variable names, comments or strings that are textually or semantically close to the query. It may surface functions that mention tokens or validation in general, whether or not they actually call validateToken(), and it may miss a caller that uses different vocabulary entirely.
Structural retrieval
Structural retrieval instead looks up the explicit call relationships already resolved for validateToken() and returns exactly what invokes it, regardless of naming or wording.
relationship graph
login()
       \
refreshSession() ---> validateToken()
       /
authenticateAPI()

Here, login(), refreshSession() and authenticateAPI() all call validateToken() directly: a dependency relationship a call graph represents explicitly, rather than one that has to be inferred from text.

The building blocks of structural retrieval

Structural retrieval is built on a handful of established concepts for representing code as something other than plain text.

AST

An abstract syntax tree is a parsed representation of source code's grammar (functions, statements, expressions) as a tree rather than as raw characters. It's typically the first structural representation a parser produces.

Symbols

Symbols are the named entities in a codebase (functions, classes, variables, types) resolved to a specific definition rather than treated as arbitrary text tokens.

Call graphs

A call graph represents which functions call which other functions, letting a system answer questions like "what calls this?" or "what does this call?" directly.

Dependency graphs

A dependency graph represents which modules, packages or files depend on which others, at a coarser granularity than individual function calls.

Import relationships

Import relationships capture which files or modules bring in code from which other files or modules, the concrete links a dependency graph is often built from.

Inheritance

Inheritance relationships capture which classes extend or implement which others, letting a system trace behavior defined in a parent class to the subclasses that use it.

A related but distinct concept is data flow: how a value moves and transforms as it passes between variables and function calls. Data-flow analysis is a separate, more involved category of structural analysis on its own.

See structural retrieval on a real repository

CodeMesh's published benchmark measures what structural retrieval saves in tokens and cost.

Start Saving TokensSee the Benchmark