Code-retrieval benchmark: coding-v1¶
Question set coding-v1 (16 questions, 1 reps/arm) against 4 arms. Repo at 7924e45bd, model claude-sonnet-5, rapid-proxy/claude-sonnet-5, generated 2026-08-19T10:42:57.636Z.
Continuation budget: 2 turn(s) per answer-file agent (inert for claude, whose loop is unbounded). Runs at different budgets are not comparable.
Secondary scorer: claude-sonnet-5 via gh-copilot.
Arms searched a sandboxed worktree of 7924e45bd with 36 path(s) removed (answer key, telemetry exports, agent rule files), verified to contain no question prompt or provenance note.
Overall¶
Pooled across 3 agents (
claude,copilot,opencode). Each row below mixes cells that were run by different agents, elicited differently, and — for everything except claude — not confined to the arm's tool surface. The medians are arithmetic, not comparable. Per agent is the table to read; this one is kept only so arm-level reliability counts have somewhere to live.
| Arm | ranked | correctness (median) | content tokens (median) | total tokens (median) | tool calls | latency s | hard-fail | hallucination |
|---|---|---|---|---|---|---|---|---|
| grep † | 47/48 | 1.00 | 99912 | 132628 | 5.0 | 33.5 | 2% | 0% |
| graphify | 16/16 | 1.00 | 218368 | 241533 | 9.5 | 40.9 | 0% | 0% |
| codegraph | 16/16 | 1.00 | 74829 | 97128 | 2.0 | 16.8 | 0% | 0% |
| hybrid † | 47/48 | 1.00 | 96758 | 136511 | 4.0 | 24.8 | 2% | 2% |
content tokens = total minus that arm's measured empty-run baseline. Whole-session totals are dominated by a fixed floor of system prompt + tool schemas, which compresses every ratio; content tokens are what separate retrieval strategies.
† This arm's tool surface was not enforced on every cell. Only claude can be confined to an arm's tools (--allowedTools/--disallowedTools/--strict-mcp-config); on the other agents the arm's MCP servers are restricted but the built-in file and search tools stay open. The row above is what the cells did, under a label describing what they were asked to be confined to. See Measurement provenance.
Per agent¶
These are the numbers the run actually produced. The Overall table above pools them per arm, and pooling across agents is only meaningful for hybrid, whose surface is "everything" on every agent.
| Arm | Agent | ranked | correctness | content tokens | total tokens | tool calls | latency s | hard-fail | built-ins | answer via | tokens from |
|---|---|---|---|---|---|---|---|---|---|---|---|
| grep | claude | 16/16 | 1.00 | 79338 | 102065 | 4.0 | 17.3 | 0% | enforced | stream-json | stream-json×16 |
| grep | copilot | 15/16 | 1.00 | — | 156733 | 6.0 | 34.2 | 6% | not_enforced | answer-file | unmeasured×15, proxy-db-window×1 |
| grep | opencode | 16/16 | 1.00 | 129717 | 161590 | 4.0 | 46.3 | 0% | ungated | answer-file | stream-json×16 |
| graphify | claude | 16/16 | 1.00 | 218368 | 241533 | 9.5 | 40.9 | 0% | enforced | stream-json | stream-json×16 |
| codegraph | claude | 16/16 | 1.00 | 74829 | 97128 | 2.0 | 16.8 | 0% | enforced | stream-json | stream-json×16 |
| hybrid | claude | 16/16 | 1.00 | 81612 | 107933 | 4.0 | 17.1 | 0% | enforced | stream-json | stream-json×16 |
| hybrid | copilot | 16/16 | 1.00 | 70541 | 134797 | 6.0 | 34.3 | 0% | not_enforced | answer-file | unmeasured×15, proxy-db-session×1 |
| hybrid | opencode | 15/16 | 1.00 | 114406 | 146765 | 3.0 | 29.6 | 6% | ungated | answer-file | stream-json×16 |
tool calls is a dash only where the cell produced no tool trace at all — a dash means not measured, never zero. claude reports one via stream-json; copilot and opencode now report one parsed from their own JSON event streams. Their counts are in THEIR OWN tool vocabulary (view/create, read/write) rather than the arm's, which is why those cells carry tool_audit: observed — what ran is recorded, but conformance to the arm is not decidable from it. copilot's count includes one task_complete autopilot sentinel per cell, reported separately as control calls; claude has no equivalent, so subtract it before comparing tool counts across agents.
Winner by question class¶
| Class | grep | graphify | codegraph | hybrid | winner |
|---|---|---|---|---|---|
| abstain | 1.00 | 1.00 | 1.00 | 1.00 | tie — tie (ratio 1.00x < 1.25x) |
| arch | 1.00 | 1.00 | 0.73 | 1.00 | tie — tie (ratio 1.00x < 1.25x) |
| blast | 1.00 | 1.00 | 1.00 | 1.00 | tie — tie (ratio 1.00x < 1.25x) |
| lookup | 1.00 | 1.00 | 1.00 | 1.00 | tie — tie (ratio 1.00x < 1.25x) |
| structural | 1.00 | 1.00 | 1.00 | 1.00 | tie — tie (ratio 1.00x < 1.25x) |
A winner is declared only at a ≥1.25x median gap with non-overlapping IQR. Anything weaker prints "tie" — at these sample sizes a 1.3x gap is not a result.
Reliability¶
| Arm | runs | ranked | ungraded | failed | retry rate | hard-fail rate |
|---|---|---|---|---|---|---|
| grep † | 48 | 47 | 0 | 1 | 4% | 2% |
| graphify | 16 | 16 | 0 | 0 | 0% | 0% |
| codegraph | 16 | 16 | 0 | 0 | 0% | 0% |
| hybrid † | 48 | 47 | 0 | 1 | 2% | 2% |
Failed runs are counted, never dropped. An arm that stalls is not cheap — it is unavailable, and averaging only its successes would report the opposite.
Measurement provenance¶
Rows in this report come from more than one run. Every source below used the same sandboxed corpus, question set, reps and continuation budget — that is what makes them poolable — but they were executed separately, so anything that varies between runs (model nondeterminism, provider routing at the time) varies across these rows too.
| Run | Supplies | Why |
|---|---|---|
kgv1-telemetry | claude (4 arms) and copilot (grep, hybrid) — 96 cells | the primary run; its opencode cells are superseded and excluded |
kgv1-opencode-fixed | opencode (grep, hybrid) — 32 cells | re-run after a telemetry defect undercounted opencode tool calls 3.3x; same pinned corpus 7924e45bd |
Only claude cells were tool-enforced. --allowedTools, --disallowedTools and --strict-mcp-config are claude flags. For copilot and opencode an arm's MCP servers are restricted by writing the config file each CLI reads, but their built-in file and search tools cannot be withheld — so on those agents an arm name describes the retrieval strategy the cell was asked to use, not one it was confined to. Arms whose identity depends on withholding built-in search are refused outright on those agents rather than run under a label they would not honour.
Answers were elicited differently. claude streams its answer as structured JSON; the others are told to write it to a file, because an analysis-shaped prompt makes copilot exit in seconds and opencode yield on its first toolless step, both "succeeding" having answered nothing. That difference is a confound in every cross-agent comparison here, and it is not removable — it is what makes those cells produce an answer at all.
Where the token numbers came from.
| Source | cells | what it means |
|---|---|---|
stream-json | 96 | the agent reported its own usage — first-party and exact |
unmeasured | 30 | no rows found; the field is null, never 0 |
proxy-db-window | 1 | proxy rows stamped while the cell ran — a time join that cannot separate a neighbour's tail from this cell |
proxy-db-session | 1 | whole proxy sessions that BEGAN while the cell ran — attributed per session, so a neighbour's trailing calls are not charged here |
A cell whose tokens are unmeasured still ranks on correctness; it is only absent from the token medians. Reporting 0 there would make the least measurable agent look the cheapest.
1 cell(s) had more than one session of the same agent start inside a SINGLE attempt, or a session that started between their attempts. A cell that was retried owns one session per attempt and is priced correctly without appearing here, so these are cells where something is genuinely unaccounted for. Two causes produce it and the count alone does not separate them: another session of that agent running alongside the benchmark, or a recorded attempt window that is wrong. Read
token_ambiguityon the row — it names the attempt — and compare that attempt's span against the session timestamps in the proxy token DB before quoting or dropping the cost.
Agents in this run: claude, copilot, opencode.
Checklist vs judge disagreements¶
| Question | Arm | checklist | judge | note |
|---|---|---|---|---|
| A4 | grep | 1.00 | 0.67 | checklist_higher |
| A4 | grep | 1.00 | 0.67 | checklist_higher |
| A4 | graphify | 0.33 | 0.00 | checklist_higher |
| B2 | hybrid | 1.00 | 0.50 | checklist_higher |
| B2 | grep | 1.00 | 0.50 | checklist_higher |
| L2 | hybrid | 0.65 | 0.00 | checklist_higher |
judge_higher usually means the checklist matcher is too strict (the answer paraphrased a path) — fix the matcher and re-grade offline. checklist_higher usually means correct strings were padded into a wrong narrative, which is a real quality signal.
This table is an alarm, not a diagnosis. It says two graders differ; it does not say which is wrong, and the answer has not once been the obvious one. Across every investigation on this set the causes were a judge rubric, a false answer key, a regex, a shared match token, and a matcher that was too loose and too narrow at the same time — never a badly written question. Twice the arms were right and the key was wrong. And the detector is blind to the most common defect of all: because the judge's prompt is built from the same checklist, a WRONG KEY makes both graders agree and produces zero disagreements. See Measurement and judging lessons before concluding a question is at fault.
Limitations¶
- 1 reps per cell on one repository with one scorer, across 3 agents whose cells are not equivalently enforced (see above).
- Arms other than
hybridare FORCED onto a single retrieval strategy, which is not how an agent works in practice. Read them againsthybrid, not against each other. - Indexing cost is excluded from per-query numbers; it is reported separately per backend.
- Corpus scope differs between backends (graphify indexes docs and PDFs; code-only backends do not), so node/edge counts are not comparable at face value.