Run the same 98 software engineering questions against the same codebase using the same model, once with a <file-reading-agent> and once using only <CodeMesh>.
| ID↑ | Tier | Question | Reading files | CodeMesh | Reduction |
|---|---|---|---|---|---|
| Q001 | medium | Which file defines openclaw's gateway wire protocol version constant, and what is the current PROTOCOL_VERSION value?gateway-protocol | 142,511 | 9,481 | 93.3% |
| Q002 | hard | openclaw's gateway protocol version file declares four minimum/current version constants. Name all four and give each value, and explain what the asymmetry between the client minimum and the node/probe minimum means in practice.gateway-protocol | 119,695 | 10,130 | 91.5% |
| Q003 | hard | When the openclaw gateway is still starting its sidecars and cannot serve a websocket connection yet, which WebSocket close code does it use, what retry-after hint does it advertise, and what are the reason strings? Name the file.gateway-protocol | 524,184 | 14,742 | 97.2% |
| Q004 | medium | What is the worker heartbeat interval used by openclaw's worker admission protocol, and what is WORKER_RPC_SET_VERSION set to?worker-protocol | 119,709 | 9,623 | 92.0% |
| Q006 | hard | What are the transcript-commit batch limits in openclaw's worker admission protocol: max messages per batch, max content parts, and max JSON nesting depth?worker-protocol | 136,668 | 9,632 | 93.0% |
| Q007 | medium | In openclaw's worker inference protocol schema, what is the maximum number of context messages and the maximum output tokens allowed?worker-protocol | 176,774 | 9,728 | 94.5% |
| Q008 | medium | What is the maximum number of chat history entries openclaw's gateway protocol schema allows, and which file defines that cap?gateway-protocol | 234,344 | 9,620 | 95.9% |
| Q009 | medium | How many session targets can a single sessions.patchMany request address in openclaw, and what constant name carries that limit?gateway-protocol | 267,645 | 9,635 | 96.4% |
| Q010 | medium | What is the maximum allowed length of a chat send session key in openclaw's gateway protocol primitives, and how many viewer-presence keys can a session track?gateway-protocol | 158,645 | 14,309 | 91.0% |
| Q011 | medium | What are openclaw's default gateway request timeout and default pre-auth handshake timeout, and which package defines them?gateway-client | 157,226 | 9,548 | 93.9% |
| Q012 | medium | In openclaw's gateway client timeouts module, what is MIN_CONNECT_CHALLENGE_TIMEOUT_MS and what is MAX_SAFE_TIMEOUT_DELAY_MS?gateway-client | 76,978 | 9,733 | 87.4% |
| Q013 | hard | What are the documented default values for openclaw's gateway auth rate limiter config (max failed attempts, sliding window, lockout duration, max tracked identities)?gateway-auth | 410,958 | 9,749 | 97.6% |
| Q014 | hard | How does openclaw's gateway auth rate limiter treat loopback addresses (127.0.0.1 / ::1) by default, and what still happens to failed auth attempts from loopback?gateway-auth | 252,492 | 9,981 | 96.0% |
| Q015 | hard | Name the rate-limit scope constants exported by openclaw's gateway auth rate limiter and say what a scope is for.gateway-auth | 130,778 | 15,065 | 88.5% |
| Q016 | expert | Why does openclaw give the pre-auth device bootstrap-token verify path its own dedicated rate-limit scope instead of reusing the shared-secret scope? Cite the mechanism named in the code.gateway-auth | 155,835 | 26,786 | 82.8% |
| Q017 | medium | What is openclaw's gateway control-plane rate limit (requests per window) and window length?gateway | 158,656 | 9,599 | 93.9% |
| Q018 | medium | What is openclaw's default channel connect grace period, and which module defines it?gateway | 237,262 | 9,461 | 96.0% |
| Q019 | hard | In openclaw's gateway channel health policy, when a channel snapshot reports lifecycle "starting" or "recovering", under what condition is it judged healthy, what reason string is returned, and what happens if the start timestamp is missing or the window has passed?gateway | 127,855 | 20,641 | 83.9% |
| Q020 | medium | What HTTP path prefixes does openclaw's gateway use for board views and for the standalone MCP app?gateway | 448,401 | 9,586 | 97.9% |
| Q021 | medium | What is openclaw's default maximum number of ingress delivery attempts, and what is the default maximum backoff delay between ingress retries?channels-ingress | 151,949 | 9,560 | 93.7% |
| Q022 | hard | What two conditions must both hold before openclaw dead-letters a retryable ingress event, and what happens to an event that has exceeded the attempt limit but not the other condition?channels-ingress | 120,034 | 10,246 | 91.5% |
| Q023 | hard | In openclaw's ingress failure disposition logic, what are the two possible disposition kinds, and what reason string is recorded when the retry limit is exceeded?channels-ingress | 145,360 | 9,581 | 93.4% |
| Q024 | hard | When openclaw computes the remaining backoff delay for a released ingress event, what backoff factor and jitter does it use, and how is the attempt count bounded before the backoff is computed?channels-ingress | 467,252 | 19,391 | 95.8% |
| Q025 | hard | Which openclaw workspace package implements computeBackoff, what is its exact formula, and what does src/infra/backoff.ts contribute?packages/retry | 160,754 | 9,961 | 93.8% |
| Q026 | hard | In openclaw's retry package, what does computeBackoffSchedule return for attempt <= 0, and what happens when the attempt exceeds the provided schedule length?packages/retry | 139,204 | 9,715 | 93.0% |
| Q027 | hard | How does openclaw's sleepWithAbort guard against invalid or oversized delays, and what error does it raise when the abort signal fires?packages/retry | 167,941 | 9,970 | 94.1% |
| Q028 | medium | What is openclaw's outbound echo suppression window, and how many outbound message identities are retained?channels | 114,058 | 9,559 | 91.6% |
| Q029 | medium | How many consecutive typing-indicator failures does openclaw tolerate by default before backing off, and which file holds that default?channels | 162,778 | 9,557 | 94.1% |
| Q030 | hard | What are openclaw's default thread-binding idle and max-age policy values, and what does the max-age default mean?channels | 283,107 | 10,161 | 96.4% |
| Q031 | medium | What are the default minimum and maximum draft streaming chunk sizes in openclaw's draft streaming chunker?channels | 108,892 | 9,671 | 91.1% |
| Q032 | medium | What is openclaw's default initial delay before a progress draft is emitted, and what is the channel feedback-reflection cooldown default?channels | 163,225 | 9,691 | 94.1% |
| Q033 | hard | What limits does openclaw apply when probing inbound media on a channel event: how many probes, at what concurrency, and within what time budget?channels | 234,429 | 9,781 | 95.8% |
| Q034 | medium | What is openclaw's default run-activity heartbeat interval in the channel run state machine?channels | 240,226 | 9,456 | 96.1% |
| Q035 | hard | In openclaw's channel plugin loading, what caps exist on bundled channel load contexts and boundary roots, and what is the read-only channel plugin resolution cache size?channels | 332,979 | 9,849 | 97.0% |
| Q036 | medium | What timeout does openclaw apply when waiting for a configured binding route to become ready?channels | 195,721 | 9,587 | 95.1% |
| Q037 | medium | What are openclaw's default and maximum cron script timeouts (in seconds), and the default and maximum cron script tool budgets?cron | 158,648 | 9,596 | 94.0% |
| Q038 | hard | What wall-clock limit and tool budget does openclaw apply to a headless cron trigger evaluation, and how many trigger evaluations may run concurrently?cron | 339,549 | 9,819 | 97.1% |
| Q039 | medium | How many transient retries does openclaw's cron timer trigger default to, and after how many consecutive failures does the cron service raise a failure alert by default?cron | 276,266 | 9,549 | 96.5% |
| Q040 | hard | What are the default, minimum and maximum cron stream batch intervals and batch byte sizes in openclaw?cron | 106,057 | 9,606 | 90.9% |
| Q041 | hard | In openclaw's cron command runner, what is the default command timeout, and what value is used for the "effectively unbounded" timeout constant?cron | 194,672 | 9,571 | 95.1% |
| Q042 | medium | What is the cron expression evaluation cache size in openclaw, and what is the maximum length of a declarative cron job label?cron | 361,703 | 18,959 | 94.8% |
| Q043 | hard | What caps does openclaw apply to cron run diagnostics normalization (entries, chars per entry, summary chars), and what tail size does it keep for exec diagnostics?cron | 304,274 | 30,800 | 89.9% |
| Q044 | hard | When openclaw builds the "current context" block for an isolated cron agent run, how many messages does it include and what per-line and per-block character caps apply?cron | 165,243 | 9,568 | 94.2% |
| Q045 | hard | What limits does openclaw's session diff module enforce on files, untracked files, per-file patch bytes, total patch bytes and total changed lines?sessions | 129,430 | 15,030 | 88.4% |
| Q046 | medium | How long does openclaw wait for session work admission to drain, and which constant carries that value?sessions | 252,036 | 9,477 | 96.2% |
| Q047 | medium | What filename does openclaw's memory host SDK treat as the memory root file, and what is the current memory chunking version?memory | 202,781 | 9,492 | 95.3% |
| Q048 | medium | What are the default read-lines and max-chars limits when openclaw's memory host reads a memory file?memory | 945,440 | 9,534 | 99.0% |
| Q049 | hard | Name the SQLite table-name constants openclaw's memory host SDK defines for its index (base schema plus FTS schema).memory | 113,417 | 10,530 | 90.7% |
| Q050 | medium | How many memory items can a single openclaw memory migration carry, and which schema file defines that limit?memory | 190,647 | 9,457 | 95.0% |
| Q051 | hard | What error code does openclaw's agent core use to signal that a transcript cannot be continued, and where is it defined?agent-core | 142,551 | 9,573 | 93.3% |
| Q053 | hard | In openclaw's durable outbound delivery completion module, what are the two completion "kind" variants and what states can a completion result report?delivery | 189,111 | 19,433 | 89.7% |
| Q054 | expert | openclaw defines a CHARS_PER_TOKEN_ESTIMATE constant with the same value in two different workspace packages. Name both files and the value, and say what that duplication implies about package layering.tokenization | 129,499 | 9,760 | 92.5% |
| Q055 | hard | What exact marker string does openclaw use as the system prompt cache boundary, and which file owns it?ai-providers | 118,278 | 9,563 | 91.9% |
| Q056 | medium | What Claude Code version string does openclaw's Anthropic model contract pin, and in which file?ai-providers | 142,183 | 9,517 | 93.3% |
| Q057 | hard | What Anthropic beta header value does openclaw use for server-side fallback, and what value does it send for the fallback selection?ai-providers | 145,094 | 9,545 | 93.4% |
| Q058 | medium | What placeholder text does openclaw substitute when Anthropic reasoning content is omitted from a replayed message?ai-providers | 124,006 | 9,450 | 92.4% |
| Q059 | medium | What is openclaw's default Azure OpenAI API version, and what first-event timeout does it apply to Azure Responses streams?ai-providers | 570,821 | 15,509 | 97.3% |
| Q060 | medium | What default instructions string does openclaw send for OpenAI Codex Responses requests, and what does it use for empty input text?ai-providers | 2,498,016 | 9,681 | 99.6% |
| Q061 | hard | openclaw normalizes differences between OpenAI Responses and Azure Responses streaming events. What are the two content-part type strings and the two delta event type strings it maps between?ai-providers | 122,149 | 10,022 | 91.8% |
| Q062 | medium | What are the maximum lengths openclaw enforces for an OpenAI prompt-cache key and for a replayed Responses item id?ai-providers | 408,702 | 19,273 | 95.3% |
| Q064 | medium | What source id does openclaw assign to built-in API providers, and which file registers it?ai-providers | 1,972,677 | 9,475 | 99.5% |
| Q065 | hard | How does openclaw compare secret strings without leaking length information, and what does its comparison return when both inputs are empty strings?security | 107,297 | 14,144 | 86.8% |
| Q066 | medium | In openclaw's safe-regex checker, what are the cache size and test window constants?security | 111,080 | 9,487 | 91.5% |
| Q067 | hard | What does openclaw's stripUrlUserInfo do, which package owns it, and how does it behave when the input is not a parseable URL?security | 139,566 | 9,735 | 93.0% |
| Q068 | hard | Which helpers does openclaw's skill env-override runtime use to block dangerous host environment variables, and where do those helpers live?security | 294,711 | 14,371 | 95.1% |
| Q070 | expert | openclaw's gateway agents-workspace server method is deliberately read-only. What does its header comment say about mutations, and which API is named as the place edits belong?gateway | 141,809 | 20,056 | 85.9% |
| Q072 | hard | In openclaw's host env security module, which two git-related environment variables are singled out by name, and what default protocol allowlist does it define?security | 183,171 | 9,620 | 94.7% |
| Q101 | expert | When a cron job's isolated agent run finally reaches delivery (for example after a long gateway outage delayed the run), how does the delivery-dispatch pipeline decide the scheduled announce/direct delivery is too stale to send, and what is the exact threshold?cron-delivery | 270,785 | 14,789 | 94.5% |
| Q102 | extreme | What retry delay does openclaw apply after a cron job's 4th consecutive execution failure, and what happens to the delay once a job has failed 8 times in a row?cron-retry-backoff | 335,960 | 14,808 | 95.6% |
| Q103 | extreme | For an isolated cron agent run configured with a 90-second wall-clock job timeout, how long does the pre-execution watchdog wait (once the runner has started but before the run reaches an execution-stage phase) before it fires a timeout, and what formula produces that number?cron-watchdog | 113,546 | 14,947 | 86.8% |
| Q104 | expert | On gateway restart, when the cron service replays cron jobs that were missed while it was down, how many missed jobs does it run immediately by default, what happens to the rest, and why are missed agentTurn jobs handled differently from other missed jobs?cron-startup-catchup | 162,596 | 10,424 | 93.6% |
| Q105 | expert | When openclaw runs a user-configured 'safe' regex against an untrusted string, how does it avoid quadratic-time scanning on very long inputs, and what happens if a match only occurs in the middle of a long string but not within the tested windows?security-regex | 135,149 | 10,046 | 92.6% |
| Q106 | hard | When openclaw masks a credential-like value for logging, how many characters does it reveal from the start and end of the secret, and does that count change based on the secret's length?security-secret-masking | 273,080 | 9,891 | 96.4% |
| Q107 | extreme | The gateway's auth rate limiter gives loopback (localhost) callers an escalating delay instead of a hard lockout after repeated failed auth attempts. How many of those failed-attempt timestamps does it retain per loopback key before trimming older ones, and how is that number derived from the delay constants rather than being a separate hardcoded value?auth-rate-limit | 217,820 | 15,346 | 93.0% |
| Q108 | hard | How many 'unauthorized role' errors is a single WebSocket connection allowed to trigger by default before the gateway forcibly closes it, and is that closing event guaranteed to be logged even when it doesn't fall on the periodic log-sampling boundary?gateway-flood-guard | 96,983 | 20,360 | 79.0% |
| Q109 | expert | In the memory host SDK's SQLite schema, chunk-level recall metadata (importance, triggers, project_key) lives in its own table rather than on the chunk table itself. What numeric range is enforced for the importance column, and what does the schema-ensure function do if it finds those same three columns still present directly on the canonical chunks table from an unreleased build?memory-schema | 98,011 | 15,074 | 84.6% |
| Q110 | extreme | The memory host SDK's chunk provenance table restricts origin_class and session_kind to closed enums and cascades deletes from the canonical chunk table. When the schema-ensure function backfills provenance rows for chunks that predate the provenance table, what exact origin_class and session_kind values does it assign, and where does it source the observed_at timestamp from?memory-schema | 92,926 | 9,970 | 89.3% |
| Q111 | extreme | The memory host SDK's FTS schema helper for the derived chunk-body index refuses to build if its target table name collides with a specific other table -- which one, and why does the check use a case-insensitive comparison rather than a plain equality check?memory-schema-fts | 101,788 | 10,252 | 89.9% |
| Q112 | hard | During the memory host SDK's legacy-table-to-canonical migration, which two legacy tables get temporary indexes created before the row copy runs, and on exactly which columns?memory-schema-migration | 109,345 | 9,727 | 91.1% |
| Q113 | extreme | The gateway's background loop that automatically retries restart-aborted main sessions uses an exponential backoff. Using its default constants, what is the sleep delay immediately before the third (final) recovery attempt?restart-recovery | 307,170 | 10,044 | 96.7% |
| Q114 | extreme | There are two different user-facing 'gateway restart' notice strings for an interrupted main session -- one telling the user their turn couldn't safely resume, and one telling them to run /new or /reset. What distinguishes the two code paths that decide which one gets sent, and how does each affect the session's stored recovery state going forward?restart-recovery | 641,366 | 19,915 | 96.9% |
| Q115 | hard | When the recovery scanner reads a running main session's transcript to decide how to resume it after a restart, what exact message-count and byte-size caps does it pass to the transcript reader?session-transcript | 693,278 | 13,523 | 98.0% |
| Q116 | hard | There is a fixed threshold in the agents runtime for warning about a tool-call loop, and a comment ties its non-configurability to a specific past PR. What is that threshold value and what does the comment say was retired?agents-runtime | 125,670 | 9,744 | 92.2% |
| Q117 | extreme | The gateway rejects a worker's connect request with reason "rpc-set-mismatch" under two distinct conditions, not just when the worker's declared rpcSetVersion disagrees with the stored credential. What is the second condition, and which of the initial admission checks (bundle hash, version, session, owner epoch, rpc set, protocol features) is never re-checked again once the worker is admitted and starts sending heartbeats/RPCs?worker-admission | 132,225 | 31,906 | 75.9% |
| Q118 | extreme | In openclaw's WorkerConnection client, which close reasons put the connection into the terminal "fenced" state versus the retryable reconnect path, and how can a successful heartbeat response, with no socket close at all, still trigger fencing?worker-connection-lifecycle | 263,351 | 10,805 | 95.9% |
| Q119 | hard | The worker protocol enforces two different byte ceilings for JSON payload content depending on RPC family. What are the two values, and which constant deliberately reuses the smaller one to bound opaque provider-replay state blobs?worker-protocol-limits | 452,475 | 15,476 | 96.6% |
| Q120 | expert | What exact string format must a worker's build-identity bundle hash satisfy in the gateway protocol schema, and against how many independent server-side sources does the gateway cross-check that same hash before admitting the connection?worker-admission | 158,698 | 14,440 | 90.9% |
| Q121 | extreme | Before validating a worker inference frame (start/cancel/event/terminal) against its typebox schema, openclaw runs a separate hand-rolled JSON safety walk. What three classes of value does this walk reject that the schema check alone wouldn't necessarily catch, and what is the exact nesting-depth limit it enforces?worker-inference-validation | 186,658 | 15,089 | 91.9% |
| Q122 | hard | Why does the helper that builds the local WebSocket URL for a worker-to-gateway connection explicitly throw when the socket path contains a colon?worker-connection-transport | 218,844 | 15,148 | 93.1% |
| Q123 | expert | What are the exact reconnect-backoff parameters and overall admission deadline a worker uses when repeatedly reconnecting to its gateway after a retryable close, and what pre-existing gateway-client timeout constant does the per-attempt admission timeout default to?worker-connection-retry | 487,260 | 10,108 | 97.9% |
| Q124 | extreme | When a new cloud worker turn tries to claim a session whose previous turn's workspace result is still reconciling (rather than truly active elsewhere), how long does the gateway wait for the prior claim to release before giving up, what error message does it surface, and what git-ref shape must a stored workspace-conflict record satisfy before it's trusted as durable?worker-turn-lifecycle | 302,725 | 20,875 | 93.1% |
| Q125 | extreme | When Anthropic's server-side fallback beta reroutes a refused stream onto Opus 5 or 4.8 mid-generation, what happens to the pre-fallback assistant content blocks in openclaw's transcript, and what fixed cost table does it substitute for billing that turn?llm-provider-anthropic | 158,654 | 10,509 | 93.4% |
| Q126 | extreme | What are the default thresholds openclaw's bot-to-bot channel reply loop guard uses before it starts suppressing messages, and does it count a message from bot A to bot B separately from a reply from bot B to bot A?channels-loop-guard | 132,670 | 19,836 | 85.0% |
| Q127 | extreme | The before_compaction and after_compaction plugin hooks get a 30-second default timeout in openclaw's hook runner, the same as agent_end, instead of the tighter budget used by other modifying hooks. What specific mechanism in the codex agent harness does the code cite as the reason a hung compaction handler would be especially dangerous?plugin-hooks | 126,305 | 27,303 | 78.4% |
| Q128 | extreme | For the OpenAI ChatGPT Responses transport, openclaw defines two separate byte caps for reading the HTTP response body from the same endpoint family: one for error responses and one for streamed success bodies. What are the two constant values, and by what factor do they differ?llm-provider-streaming | 109,526 | 9,992 | 90.9% |
| Q129 | extreme | What are the default request-count and time-window limits for openclaw's in-memory webhook rate limiter, and how does its sibling anomaly tracker decide when to actually emit a log line for repeated failures from the same key?plugin-sdk-webhook | 97,014 | 14,937 | 84.6% |
| Q130 | hard | When creating an armable transport stall watchdog without an explicit check interval, what formula does openclaw use to derive the poll interval from the idle timeout, and what floor does it ultimately get clamped to?channels-transport | 95,373 | 9,806 | 89.7% |
| Q131 | extreme | What exact literal marker does openclaw insert into a system prompt to separate the cache-stable prefix from dynamic runtime additions, and what does prependSystemPromptAdditionAfterCacheBoundary do differently when that marker is absent from the prompt versus when it is present with no dynamic suffix after it?llm-provider-prompt-cache | 92,613 | 10,294 | 88.9% |
Both arms are the same model driving the same CLI against the same checkout of openclaw. The only variable is how the agent is allowed to find things.
code-answer endpoint. Retrieval and synthesis happen server-side against the code graph; only the answer crosses back into the model's context.Questions span four difficulty tiers, from “which file defines this constant” up to multi-hop questions about why two code paths disagree. No question was written after seeing a result.
Splitting by tier is the most useful cut of this data, because it shows exactly where the token savings come from. Reduction barely moves across tiers; it stays above 92% even on the hardest, most multi-hop questions.
| Tier | Questions | Reading files | CodeMesh | Token reduction |
|---|---|---|---|---|
| medium | 32 | 360,313 | 10,496 | 97.1% |
| hard | 40 | 215,333 | 12,330 | 94.3% |
| extreme | 17 | 200,725 | 15,696 | 92.2% |
| expert | 9 | 193,294 | 14,609 | 92.4% |
| All tiers | 98 | 258,115 | 12,524 | 95.1% |
Token columns are per-question averages; the reduction column is computed from tier totals. On the medium tier, the shape of question developers ask most often, 32 of them here, CodeMesh answers at 97.1% fewer tokens. Even on extreme and expert questions, reduction stays above 92%.
All 98 questions bucketed by how much less context the CodeMesh run consumed. The mass sits between 90% and 98%; this is a consistent effect, not a long tail carried by a few spectacular outliers.
7 questions fall below 85%. Those are the ones where retrieval escalated through several rounds before answering, the system spending tokens on purpose rather than returning a guess.
Cheap answers only matter if they are right, so the same run was graded. Every answer was judged against hand-written ground truth, median of three votes. Nothing is dropped: the 11 the system got outright wrong are in the table above with the rest.
Real measured API spend for this run, not a projection from a rate card. The same 98 answers cost $23.67 the file-reading way and $1.65 through CodeMesh.
| Reading files | CodeMesh | Saved | |
|---|---|---|---|
| Total tokens, 98 questions | 25,295,273 | 1,227,371 | 95.1% |
| Per question, average | 258,115 | 12,524 | 92.7% |
| Measured API spend | $23.67 | $1.65 | $22.03 |
A 258,115-token average is not a rounding error: it is most of a context window spent finding the answer before any thinking happens. That is the cost CodeMesh removes.
We deliberately benchmarked a public open-source codebase rather than our own, so nothing here rests on trusting us. openclaw is open: the questions, the ground truth and the numbers can all be checked independently. The number that matters for you, though, is the one from your repository, on your questions.
These numbers come from one public codebase. The fastest way to know what CodeMesh does for yours is to connect it.