Skip to content

Benchmarking Agent Systems

How to measure systems that do cognitive work with LLMs so the numbers survive scrutiny. kgbench is the worked example; the hazards are the point.

The hazard

Every control described here exists because a run without it produced a plausible, publishable, wrong result. Among them: six distinct channels by which answers leaked into the systems being tested, a comparison that turned out to be one configuration measured against itself, a token column that silently double-counted, and a report that published a grading model which had never graded anything.

Each of those produced a tidy table. That is what makes this hard โ€” the failure mode is not an error message, it is a believable number.

The vocabulary

Five terms carry the whole framework:

Term Means
Arm One retrieval strategy under test
Cell One (arm ร— agent ร— model ร— question ร— rep) โ€” a single measured run
Axis A dimension the matrix varies
Containment What a system is allowed to read
Baseline The token floor that must be subtracted before arms are comparable

The one rule

Measure the thing you claim to be measuring. Almost every failure below is a case of something else varying at the same time โ€” the environment, the harness overhead, the grader, or the answer being reachable by a route nobody accounted for.

The Glossary defines every term precisely. The Tutorial is hands-on. Experimental Design enumerates each control and the failure that motivated it.

Three reasons this is hard

A benchmark over a deterministic system is mostly bookkeeping. A benchmark over an LLM agent is not, for three reasons that compound.

The system can reach the answer by routes you did not design. An agent has tools. If the answer exists anywhere it can read โ€” the repository, its own prior output, an environment variable, a cached artefact โ€” then an arm that was supposed to be handicapped is not. Six such channels were found here, one of them structurally invisible to the sandbox built to prevent exactly that class of leak.

Cost is not a single number. Raw token totals conflate the strategy's cost with the harness overhead around it. Two arms with identical retrieval can differ in totals purely by how much scaffolding each needs, which is why a measured floor must be subtracted before arms are compared.

Grading is itself a system under test. A grader that is wrong produces results that are wrong in a consistent direction, which reads as signal.

The shape of a run

A cell is the atom: one arm, one agent, one model, one question, one repetition. The matrix is the product of those axes, and the report aggregates cells that share a question across repetitions.

Each cell runs in its own git worktree of the benchmark commit. That containment is not hygiene โ€” it is what allows the claim that an arm saw only what its strategy permits. A run outside a worktree cannot make that claim, so the harness refuses to start.

What comes out, and what to distrust

Results stream to disk as cells complete, and full answers are stored rather than only verdicts. That choice is load-bearing: it means a grading defect can be corrected and re-applied offline, and a result that changes under a corrected grader can be recognised as never having been a finding.

Three questions to ask of any table this produces:

  1. How many cells actually ran? Refused combinations are reported by preflight and are easy to skim past. A matrix missing cells still renders as a complete-looking table.
  2. Was the token floor subtracted? Without it, the comparison includes the harness.
  3. Has the grader itself been checked against known-good and known-bad answers? A grader that has never rejected anything has not been tested.

Applying this elsewhere

The transferable parts are not the scripts. They are: contain the system so it can only reach what you intend; measure the floor so cost comparisons are about the thing that varies; store raw output so grading is revisable; make refusals loud so an incomplete matrix cannot pass as a complete one; and treat any result that looks clean on the first run as a hypothesis about your harness rather than about the world.

This is a guide to measuring systems that do cognitive work with LLMs โ€” coding agents, retrieval layers, knowledge graphs, RAG pipelines โ€” in a way whose numbers survive scrutiny.

It documents kgbench, the benchmark harness in this repository. But the harness is the example, not the point. Nearly everything here was learned by getting it wrong first: six distinct channels by which the answers leaked โ€” one of them structurally invisible to the sandbox built to prevent exactly that โ€” a comparison that turned out to be one configuration measured against itself, a token column that silently double-counted, and a report that published a grading model which had never graded anything. Each of those produced a plausible-looking table. That is the hazard this guide is about.

New to this? Start with the Glossary. Every term used here โ€” arm, cell, axis, ungated, leak, containment โ€” is defined there in plain language before it is used anywhere else.

Companion documents

Document What it covers
Experimental Design The deep treatment: how bias is excluded and how token counts are made comparable
Tutorial Hands-on: run an experiment, write a question, add a retrieval backend, add an agent
Operator Reference Flags, prerequisites, reading a report
Measurement & Judging Lessons Case notes on grading failures and what each one taught

Why this is harder than it looks

The obvious way to compare two retrieval strategies is: ask both the same questions, score the answers, compare. Every part of that sentence hides a trap.

A leaked answer produces a correct answer. If the thing being measured can read the questions โ€” or a previous run's published report, or a session log in which the questions were written โ€” it will score well, and the score will look exactly like retrieval working. Leaks are invisible in the scores. They are only visible if you go looking for them before you believe the numbers.

A configuration flag may not configure anything. The first full run of this benchmark compared a "grep" strategy against a "code graph" strategy. The flag intended to restrict each one to its own tools did nothing under the mode the harness was using. Both strategies ran with the full tool surface. The comparison was one configuration measured against itself, and the tell โ€” both scoring identically on every question class โ€” had already been read as "the questions are too easy."

A missing number and a zero are not the same thing. Two of the three agents here cannot report their own token usage. Recording 0 for them would have made the least measurable agent look the cheapest. Recording null is correct but useless. Neither is a measurement.

Averages hide the fact that you measured two different things. Once cells can be run by different agents, the same arm label covers a cell that was strictly confined to three tools and a cell that kept every tool it shipped with. Their median is arithmetic, not meaningful.

The framework's design is a direct response to each of these.


Glossary

Read this first. These terms are used precisely throughout, and several of them mean something narrower than their everyday sense.

The units of measurement

Question โ€” one task put to the system, with a machine-checkable definition of a right answer. Not just a prompt: a question also carries its answer key (what must appear in a correct answer), its class, and its provenance (which files in the repository actually contain the answer, verified by a human).

Question class โ€” the kind of reasoning a question demands. This benchmark's set uses five:

Class What it asks for Why it is separate
lookup Find one specific fact โ€” a file, a variable, a value The easiest case; any strategy should manage it
structural How pieces relate โ€” what calls what, what is defined where Where an index should start to pay off
blast Blast radius โ€” what breaks if this changes Needs the graph of relationships, not just text
arch Architecture and intent โ€” why a thing is built this way Needs synthesis across many files
abstain A thing that does not exist in the repository Tests whether the system says "not here" or invents an answer

The abstain class matters more than its share of the question count suggests. A system that scores well everywhere else and fabricates confidently on absence is not usable, and nothing but an explicit absence question will reveal it.

Arm โ€” one way of answering. An arm is defined by the tool surface it is allowed to use, and nothing else. Same model, same prompt, same repository; only the tools differ. That is what makes the comparison about retrieval strategy rather than about model quality.

This benchmark's arms:

Arm Tools it may use What it represents
grep Text search + file read The baseline any retrieval layer has to beat โ€” what shipping agents already do
graphify File read + the graphify code-graph tools, no text search A graph-based index
codegraph File read + the CodeGraph tools, no text search A different graph-based index
hybrid Everything The only production-shaped arm โ€” an agent that chooses

Note what makes graphify and codegraph what they are: they withhold text search. An arm whose identity is a restriction can only exist where that restriction is enforceable. This has consequences โ€” see ungated below.

Axis โ€” a dimension the matrix varies. There are five:

arm  ร—  agent  ร—  model  ร—  question  ร—  repetition

An axis exists so that one factor can be varied while everything else is held still. Adding an axis multiplies the run: 4 arms ร— 3 agents ร— 1 model ร— 16 questions ร— 3 repetitions is 192 cells before refusals.

Cell โ€” one point in that matrix: one arm, on one agent, with one model, answering one question, once. A cell is the atomic unit of execution and the atomic unit of a result row. Each cell gets its own isolated environment and its own result record.

Repetition (rep) โ€” the same cell run again, unchanged. LLMs are non-deterministic, so a single observation is an anecdote. Repetitions turn it into a distribution you can put a median and a spread on. Three is the working minimum here; it is enough to notice variance, not enough to characterise it.

Agent โ€” which coding CLI actually drives the cell (claude, copilot, opencode). This is an axis in its own right because agents differ in far more than the model they call: their system prompts, their tool implementations, how they decide a task is finished.

Model โ€” which LLM answers. Named in the benchmark's own canonical spelling; each agent's dialect is derived, because the same model is called claude-sonnet-4-6 by one CLI and rapid-proxy/claude-sonnet-4.6 by another.

Control and enforcement

Gated / ungated โ€” a gated agent can be held to an arm's tool surface: you can tell it "you may use these tools and no others", and it obeys. An ungated agent cannot; you can configure which retrieval backends it reaches, but its built-in file and search tools are always present and cannot be withheld.

This asymmetry is the single most consequential fact about cross-agent measurement here. Of the three agents, only one is gated. On the others, an arm defined by withholding text search cannot be honoured โ€” the cell would search anyway, while wearing a label that says it did not. Those combinations are refused rather than run.

Control surface โ€” the set of things the harness can actually hold constant. Naming it explicitly matters, because the boundary between "held constant" and "hoped constant" is where invalid comparisons come from. Here the control surface is: the repository contents, the tool grant, the retrieval backends reachable, the model, the prompt, the working directory, the credentials, and the inherited configuration.

Enforcement โ€” the mechanism that makes a restriction real (a CLI flag, a config file). Every cell records which mechanism applied to it, in two parts, because the honest answer differs between them:

  • mcp_servers โ€” which retrieval backends were reachable. Enforceable on every agent.
  • builtins โ€” which built-in tools were available. Enforceable on one agent only.

A single boolean "was this enforced?" would have to lie about one of them.

Audit โ€” the check performed after the fact: compare the tools the agent actually executed against the tools it was granted. Enforcement is the mechanism; the audit is the guarantee. A flag can be wrong, a tool can be renamed, a new built-in can appear upstream โ€” each of those fails silently, and each produces a run that looks fine and compares nothing. A cell that used a tool it was not granted becomes tool_escape and cannot be scored.

Where no tool trace exists (agents that do not emit one), the cell records tool_audit: "unavailable" โ€” deliberately distinct from an empty violation list, which would read as "audited, clean".

Isolation

Sandbox โ€” a throwaway copy of the repository, created per run as a git worktree of the exact commit under test, with sensitive paths removed. The agents search this, never the live repository. It is discarded afterwards.

Containment โ€” the property that the sandbox does not contain the answers to the questions being asked. Not assumed: verified, by scanning the tree for each question's own prompt before any cell runs. The harness refuses to hand back a tree it could not verify.

Leak โ€” any path by which the answer reaches the system other than by retrieval. The obvious one is the answer key. The non-obvious ones found here, all real:

  • telemetry exports that echo the prompts, because this project records its own sessions
  • a previously published report of an earlier run
  • session logs from the sessions in which the questions were written
  • source-code comments explaining a previous leak, which quoted the thing they were explaining
  • the project's own agent instructions, which told agents to prefer one arm's tool over another
  • a prompt-injection hook that never touches the tree at all โ€” a user-level hook that retrieved this project's knowledge base against the question and prepended the answer. A sandbox is structurally incapable of catching this one, and it was feeding the agent digests of previous benchmark runs answering the same question. See Experimental Design

Leak term โ€” a string a question declares must appear nowhere in the tree. One occurrence aborts the run. This exists because the general leak scan matches five-word windows from the prompt, and a paraphrase of a question shares four words but never five in a row. The scan was built to catch a copy of a question; it has no way to catch a description of one.

Sandbox escape โ€” the agent operating outside the sandbox despite being placed inside it. Observed here twice: once by using a shell to fetch content the containment scan never saw, and once by reading an environment variable that still pointed at the live repository, and writing its answer there.

Scoring

Checklist โ€” the deterministic answer key: a list of facts a correct answer must contain, each with a matcher. Produces a score between 0 and 1 with no LLM involved, which means it is reproducible and can be re-applied offline to stored answers.

Matcher โ€” how one checklist fact is recognised in an answer. Types include exact path matching, any-of alternatives, and proximity matching that binds a claim to a subject.

Forbidden fact โ€” something a correct answer must not assert. Used mostly for absence questions, where the failure mode is confidently naming a file that does not do what is claimed. Matched only in assertive segments, so that an answer explaining what a path is not does not trip it.

Judge โ€” a second, LLM-based scorer that reads the answer and rates it. It never overrides the deterministic score. Its purpose is to disagree: a gap between the two is an alarm that something needs looking at.

Disagreement โ€” a cell where checklist and judge differ materially. Worth stating plainly what this is and is not: it names a symptom and never a cause. Across every investigation on this question set, the cause was a rubric, a false answer key, a regex, a shared match token, or a matcher that was simultaneously too loose and too narrow โ€” and never a badly written question. Twice the arms were right and the key was wrong.

Hallucination flag โ€” the answer asserted something the question declared forbidden, or fabricated where it should have abstained.

Contamination flag โ€” the answer cited the benchmark's own ground truth. Scored null, with the raw score kept as score_if_clean, so a leak can never rank as a win.

Outcomes and reliability

Every cell ends in exactly one outcome. The set is closed, which is what stops failures from quietly vanishing from the averages:

Outcome Meaning
ok Ran and produced an answer
timeout Exceeded its wall-clock budget โ€” a fact about the arm
host_stalled The timer fired far past its deadline: the machine was starved, not the arm slow. Void, not scored, not counted against anything
no_result Ran and produced no answer
api_error The model was unreachable โ€” credit, auth, availability
spawn_error The CLI could not be launched
tool_escape Used a tool it was not granted; unscorable

Distinguishing timeout from host_stalled matters: the first belongs in the arm's failure rate, the second is a fact about your laptop and belongs nowhere near the results.

Hard fail โ€” a cell that ended in any non-ok outcome. Counted, never dropped. An arm that stalls is not cheap; it is unavailable, and averaging only its successes reports the opposite.

Token accounting

Token source โ€” where a cell's token count came from. Recorded per cell, because the sources differ in kind and not merely in precision:

Source Meaning
stream-json The agent reported its own usage โ€” first-party and exact
proxy-db-taskid The request carried the cell's identifier โ€” exact, reconstructed from proxy telemetry
proxy-db-window Proxy rows recorded while the cell was running โ€” a time join, weaker than a tag
unmeasured No rows found. The field stays null, never 0

Baseline (token floor) โ€” what a session costs before any retrieval happens: system prompt, tool schemas, boilerplate. Measured by asking a trivial question that needs no tools.

Content tokens โ€” total minus the baseline. This is the number that actually distinguishes retrieval strategies. Whole-session totals are dominated by a fixed floor that compresses every ratio toward 1.0 and makes different strategies look identical.


The mental model

Everything the framework does follows from one sentence:

A cell is one measurement, and the only difference between two cells should be the axis you are varying.

Read backwards, that sentence generates the entire design. If the only difference should be the axis, then everything else must be held still โ€” the repository contents (so: a sandbox), the tool grant (so: enforcement and an audit), the retrieval backends (so: per-agent MCP restriction), the working directory and credentials (so: environment pinning), and the way the answer is scored (so: a deterministic grader that can be re-run offline).

And where something cannot be held still โ€” the elicitation difference between agents, the token accounting difference between sources โ€” the framework's obligation is to record the difference next to the number, not to average over it.


Architecture

kgbench architecture

Seven layers, each with one job:

1. Declaration. What to measure, as data rather than code. arms.json declares retrieval strategies, questions/<set>.json the questions and their answer keys, code-graph.json which retrieval backends exist. Adding an arm or a question is a config edit, not a code change.

2. Orchestration. kgbench-run.mjs walks the matrix. kgbench-supervise.sh runs it detached and resumes it after a signal death โ€” a full matrix runs for hours, and it must not be a child of anything that might tidy it up.

3. Isolation. The bias-control layer, covered in depth in Experimental Design.

4. Execution. Per-agent adapters translate a cell into that CLI's command line and know how to get an answer back out of it. All model traffic goes through one local proxy.

5. Measurement. Token attribution, ranked by source, with an offline backfill for numbers that arrive after the cell has ended.

6. Scoring. Deterministic first, LLM judge second, disagreements surfaced.

7. Reporting. Aggregation with provenance attached to every figure.

The isolation layers

kgbench isolation layers

Six layers stand between a cell and a meaningless number. They are ordered by when they act: L1โ€“L2 before the cell runs, L3โ€“L5 as it runs, L6 after. Each is explained in Experimental Design.


How a cell runs

graph TD
    A[Cell: arm x agent x model x question x rep] --> B{Already in<br/>results.jsonl?}
    B -->|yes| C[Skip โ€” resume is idempotent]
    B -->|no| D[Write per-agent MCP config<br/>into the sandbox]
    D --> E[Compose task id<br/>run--agent-model--arm-question-rep]
    E --> F[Build cell environment:<br/>strip keys, pin proxy,<br/>pin PWD, drop inherited config]
    F --> G[Spawn the agent CLI<br/>detached process group]
    G --> H{Answer<br/>obtained?}
    H -->|stream-json| I[Parse usage + tool trace]
    H -->|answer file| J[Read the file<br/>tool trace unavailable]
    H -->|neither| K[no_result]
    I --> L{Executed a tool<br/>it was not granted?}
    L -->|yes| M[tool_escape โ€” unscorable]
    L -->|no| N[Resolve tokens by source]
    J --> N
    N --> O[Deterministic grading]
    O --> P{Judge enabled<br/>and applicable?}
    P -->|yes| Q[LLM cross-check<br/>record disagreement]
    P -->|no| R[Append row to results.jsonl]
    Q --> R
    M --> R
    K --> R
    C --> S[Next cell]
    R --> S

Two properties of this flow are worth calling out.

Failure is a row, not an exception. Every path ends in a result record. A cell that timed out, escaped its tool surface, or never answered still produces a row with an outcome. Nothing is dropped, because a dropped failure is an arm that looks better than it is.

Cleanup is unconditional. The per-agent MCP config is removed after every cell, whatever happened. One of those files is written inside the measured tree, so leaving it behind would make the next cell's containment check see a file the run itself created.

How the matrix is built

graph TD
    A[--arms / --agents / --models / --set] --> B[Resolve arms<br/>expand backend tool tokens]
    B --> C[Preflight each arm<br/>is its backend actually up?]
    C -->|any fail| D[ABORT โ€” a down backend is<br/>indistinguishable from a<br/>backend that answers badly]
    C -->|all ok| E{For each arm x agent}
    E --> F{Can this agent honour<br/>this arm's identity?}
    F -->|no| G[REFUSE โ€” record reason<br/>in run.json]
    F -->|yes| H[For each model: a combination]
    G --> I[Print the whole decision<br/>before spending anything]
    H --> I
    I --> J[Build the sandbox<br/>verify containment]
    J --> K[Discover the real tool surface<br/>from the CLI itself]
    K --> L[Measure a token floor<br/>per arm x agent x model]
    L --> M{Any denied tool<br/>still available?}
    M -->|yes| N[ABORT โ€” the arms are not isolated]
    M -->|no| O[Run the matrix]

The shape of this is deliberate: everything that can invalidate a run is checked before the run spends anything. A down backend, an unenforceable arm, a leaked tree, a tool that should have been denied and is not โ€” each aborts or refuses up front. The alternative is discovering it after four hours and 200 cells.


What comes out

The results file

One JSON object per cell, appended as it completes, so a run can be inspected mid-flight and resumed. Full answers are stored untruncated โ€” which is what makes it possible to fix a grader and re-score offline instead of re-running the matrix.

Each row carries not only the numbers but their provenance: which agent, which model, which enforcement applied, how the answer was elicited, where the token count came from, and whether the tool audit was even possible.

The report

Aggregated per arm, and โ€” when the run used more than one agent โ€” per (arm, agent), because pooling those is averaging two different experiments. Every figure carries the provenance of the rows behind it, and arms whose cells were not all enforced are marked wherever their numbers appear rather than in a footnote at the bottom.

The report also declines to declare a winner unless the effect is real: a median gap of at least 1.25ร— and non-overlapping spread. Anything weaker prints "tie". At these sample sizes a 1.3ร— gap is not a result.

Where the dashboard fits

kgbench has no dashboard view of its own. Its outputs are a markdown report and SVG figures. What the dashboard does show is the shared telemetry the benchmark's token attribution is built on โ€” and that view is directly useful when diagnosing a measurement.

Token usage dashboard

Every LLM call in this environment routes through one local proxy, which records it. The treemap above is that record, broken down by process. Visible in it: kgbench judge (the benchmark's own second scorer, 1.9M tokens), and the per-agent token adapters โ€” Token adapter ยท copilot, Token adapter ยท claude, OpenCode agent (fg). Those adapters are exactly the mechanism that makes cross-agent token accounting possible at all, because two of the three agents cannot report their own usage. When a cell's tokens come back unmeasured, this is where you look to find out whether the rows exist at all.

The sibling /experiment harness โ€” a different system, for comparing agents on authoring tasks rather than retrieval โ€” does have a dashboard view:

Performance dashboard

It is shown here for contrast. It shares the proxy, the token database and the sandbox discipline with kgbench, but answers a different question: not "which retrieval strategy finds the answer" but "which agent does the work better".


Applying this beyond kgbench

If you are measuring some other system that does cognitive work with an LLM, the transferable parts are these, in rough order of how much grief they save:

  1. Verify containment, do not assume it. Whatever your system reads, check that it does not contain your answers โ€” programmatically, every run, and fail the run rather than warn. Assume you will leak; the question is only whether you find out.

  2. Enumerate what is injected into the prompt, not just what is in the corpus. Hooks, memory systems, retrieval augmentation, system prompts and prior-session state all add content your data inspection cannot see. The worst leak found here arrived this way, and it was self-reinforcing: the benchmark's own past answers were fed back into later runs.

  3. Audit what actually happened, do not trust the configuration. Record what your system actually did and compare it to what it was allowed to do. Every silent-configuration failure in this project was caught by that check and by nothing else.

  4. Make "not measured" distinguishable from "zero". In the data, in the aggregation, and in the rendered output. A zero that means "unknown" will find its way into a median.

  5. Record the provenance of every number next to the number. Not in a methodology section. A reader forms a conclusion at the moment they see the figure.

  6. Close the outcome set. Every execution ends in exactly one recorded state, including the failures, including the ones that are your machine's fault rather than the system's.

  7. Store enough to re-derive offline. Full outputs, timestamps, identifiers. Re-running a trial to fix a scoring bug changes what you are measuring; re-scoring stored outputs does not.

  8. Refuse rather than approximate. When a combination cannot be measured honestly, decline it loudly and record the refusal. A matrix that quietly shrinks is worse than one that says what it will not do.


Reading on