Skip to content

Code-retrieval benchmark: coding-v1

Question set coding-v1 (16 questions, 1 reps/arm) against 4 arms. Repo at 7924e45bd, model claude-sonnet-5, rapid-proxy/claude-sonnet-5, generated 2026-08-19T10:42:57.636Z.

Continuation budget: 2 turn(s) per answer-file agent (inert for claude, whose loop is unbounded). Runs at different budgets are not comparable.

Secondary scorer: claude-sonnet-5 via gh-copilot.

Arms searched a sandboxed worktree of 7924e45bd with 36 path(s) removed (answer key, telemetry exports, agent rule files), verified to contain no question prompt or provenance note.

Overall

Pooled across 3 agents (claude, copilot, opencode). Each row below mixes cells that were run by different agents, elicited differently, and — for everything except claude — not confined to the arm's tool surface. The medians are arithmetic, not comparable. Per agent is the table to read; this one is kept only so arm-level reliability counts have somewhere to live.

Arm ranked correctness (median) content tokens (median) total tokens (median) tool calls latency s hard-fail hallucination
grep † 47/48 1.00 99912 132628 5.0 33.5 2% 0%
graphify 16/16 1.00 218368 241533 9.5 40.9 0% 0%
codegraph 16/16 1.00 74829 97128 2.0 16.8 0% 0%
hybrid † 47/48 1.00 96758 136511 4.0 24.8 2% 2%

content tokens = total minus that arm's measured empty-run baseline. Whole-session totals are dominated by a fixed floor of system prompt + tool schemas, which compresses every ratio; content tokens are what separate retrieval strategies.

† This arm's tool surface was not enforced on every cell. Only claude can be confined to an arm's tools (--allowedTools/--disallowedTools/--strict-mcp-config); on the other agents the arm's MCP servers are restricted but the built-in file and search tools stay open. The row above is what the cells did, under a label describing what they were asked to be confined to. See Measurement provenance.

Per agent

These are the numbers the run actually produced. The Overall table above pools them per arm, and pooling across agents is only meaningful for hybrid, whose surface is "everything" on every agent.

Arm Agent ranked correctness content tokens total tokens tool calls latency s hard-fail built-ins answer via tokens from
grep claude 16/16 1.00 79338 102065 4.0 17.3 0% enforced stream-json stream-json×16
grep copilot 15/16 1.00 — 156733 6.0 34.2 6% not_enforced answer-file unmeasured×15, proxy-db-window×1
grep opencode 16/16 1.00 129717 161590 4.0 46.3 0% ungated answer-file stream-json×16
graphify claude 16/16 1.00 218368 241533 9.5 40.9 0% enforced stream-json stream-json×16
codegraph claude 16/16 1.00 74829 97128 2.0 16.8 0% enforced stream-json stream-json×16
hybrid claude 16/16 1.00 81612 107933 4.0 17.1 0% enforced stream-json stream-json×16
hybrid copilot 16/16 1.00 70541 134797 6.0 34.3 0% not_enforced answer-file unmeasured×15, proxy-db-session×1
hybrid opencode 15/16 1.00 114406 146765 3.0 29.6 6% ungated answer-file stream-json×16

tool calls is a dash only where the cell produced no tool trace at all — a dash means not measured, never zero. claude reports one via stream-json; copilot and opencode now report one parsed from their own JSON event streams. Their counts are in THEIR OWN tool vocabulary (view/create, read/write) rather than the arm's, which is why those cells carry tool_audit: observed — what ran is recorded, but conformance to the arm is not decidable from it. copilot's count includes one task_complete autopilot sentinel per cell, reported separately as control calls; claude has no equivalent, so subtract it before comparing tool counts across agents.

Winner by question class

Class grep graphify codegraph hybrid winner
abstain 1.00 1.00 1.00 1.00 tie — tie (ratio 1.00x < 1.25x)
arch 1.00 1.00 0.73 1.00 tie — tie (ratio 1.00x < 1.25x)
blast 1.00 1.00 1.00 1.00 tie — tie (ratio 1.00x < 1.25x)
lookup 1.00 1.00 1.00 1.00 tie — tie (ratio 1.00x < 1.25x)
structural 1.00 1.00 1.00 1.00 tie — tie (ratio 1.00x < 1.25x)

A winner is declared only at a ≥1.25x median gap with non-overlapping IQR. Anything weaker prints "tie" — at these sample sizes a 1.3x gap is not a result.

Reliability

Arm runs ranked ungraded failed retry rate hard-fail rate
grep † 48 47 0 1 4% 2%
graphify 16 16 0 0 0% 0%
codegraph 16 16 0 0 0% 0%
hybrid † 48 47 0 1 2% 2%

Failed runs are counted, never dropped. An arm that stalls is not cheap — it is unavailable, and averaging only its successes would report the opposite.

Measurement provenance

Rows in this report come from more than one run. Every source below used the same sandboxed corpus, question set, reps and continuation budget — that is what makes them poolable — but they were executed separately, so anything that varies between runs (model nondeterminism, provider routing at the time) varies across these rows too.

Run Supplies Why
kgv1-telemetry claude (4 arms) and copilot (grep, hybrid) — 96 cells the primary run; its opencode cells are superseded and excluded
kgv1-opencode-fixed opencode (grep, hybrid) — 32 cells re-run after a telemetry defect undercounted opencode tool calls 3.3x; same pinned corpus 7924e45bd

Only claude cells were tool-enforced. --allowedTools, --disallowedTools and --strict-mcp-config are claude flags. For copilot and opencode an arm's MCP servers are restricted by writing the config file each CLI reads, but their built-in file and search tools cannot be withheld — so on those agents an arm name describes the retrieval strategy the cell was asked to use, not one it was confined to. Arms whose identity depends on withholding built-in search are refused outright on those agents rather than run under a label they would not honour.

Answers were elicited differently. claude streams its answer as structured JSON; the others are told to write it to a file, because an analysis-shaped prompt makes copilot exit in seconds and opencode yield on its first toolless step, both "succeeding" having answered nothing. That difference is a confound in every cross-agent comparison here, and it is not removable — it is what makes those cells produce an answer at all.

Where the token numbers came from.

Source cells what it means
stream-json 96 the agent reported its own usage — first-party and exact
unmeasured 30 no rows found; the field is null, never 0
proxy-db-window 1 proxy rows stamped while the cell ran — a time join that cannot separate a neighbour's tail from this cell
proxy-db-session 1 whole proxy sessions that BEGAN while the cell ran — attributed per session, so a neighbour's trailing calls are not charged here

A cell whose tokens are unmeasured still ranks on correctness; it is only absent from the token medians. Reporting 0 there would make the least measurable agent look the cheapest.

1 cell(s) had more than one session of the same agent start inside a SINGLE attempt, or a session that started between their attempts. A cell that was retried owns one session per attempt and is priced correctly without appearing here, so these are cells where something is genuinely unaccounted for. Two causes produce it and the count alone does not separate them: another session of that agent running alongside the benchmark, or a recorded attempt window that is wrong. Read token_ambiguity on the row — it names the attempt — and compare that attempt's span against the session timestamps in the proxy token DB before quoting or dropping the cost.

Agents in this run: claude, copilot, opencode.

Checklist vs judge disagreements

Question Arm checklist judge note
A4 grep 1.00 0.67 checklist_higher
A4 grep 1.00 0.67 checklist_higher
A4 graphify 0.33 0.00 checklist_higher
B2 hybrid 1.00 0.50 checklist_higher
B2 grep 1.00 0.50 checklist_higher
L2 hybrid 0.65 0.00 checklist_higher

judge_higher usually means the checklist matcher is too strict (the answer paraphrased a path) — fix the matcher and re-grade offline. checklist_higher usually means correct strings were padded into a wrong narrative, which is a real quality signal.

This table is an alarm, not a diagnosis. It says two graders differ; it does not say which is wrong, and the answer has not once been the obvious one. Across every investigation on this set the causes were a judge rubric, a false answer key, a regex, a shared match token, and a matcher that was too loose and too narrow at the same time — never a badly written question. Twice the arms were right and the key was wrong. And the detector is blind to the most common defect of all: because the judge's prompt is built from the same checklist, a WRONG KEY makes both graders agree and produces zero disagreements. See Measurement and judging lessons before concluding a question is at fault.

Limitations

  • 1 reps per cell on one repository with one scorer, across 3 agents whose cells are not equivalently enforced (see above).
  • Arms other than hybrid are FORCED onto a single retrieval strategy, which is not how an agent works in practice. Read them against hybrid, not against each other.
  • Indexing cost is excluded from per-query numbers; it is reported separately per backend.
  • Corpus scope differs between backends (graphify indexes docs and PDFs; code-only backends do not), so node/edge counts are not comparable at face value.