Skip to content

tool-selection.md — data appendix

Generated by scripts/analysis/tool-selection-analysis.mjs. Do not hand-edit.

Runs analysed: coding-v1-r6, coding-v1-r7, coding-v1-x2, coding-v1-r8, coding-v1-r9. Arm: hybrid unless stated.

1. Tool telemetry by agent

A run with 0 recorded calls is not a run with no tool use — it is a run with no instrumentation. Every later section is therefore claude-only.

run agent cells recorded tool calls tool_audit
coding-v1-r9 claude 48 227 audited
coding-v1-r9 copilot 48 0 unavailable
coding-v1-r9 opencode 48 0 unavailable
coding-v1-r8 claude 48 230 audited
coding-v1-r8 copilot 48 0 unavailable
coding-v1-r8 opencode 48 0 unavailable

2. Graph uptake per run

unanimous counts questions whose reps all agree (all used a graph, or none did). It stays at 14–16 of 16 in every run: tool choice is a near-deterministic function of the question rather than a rate.

run tool calls graph calls rate cells w/ graph questions w/ graph unanimous
coding-v1-r6 322 3 0.93% 3/76 2/16 14/16
coding-v1-r7 306 5 1.63% 5/76 2/16 14/16
coding-v1-x2 226 3 1.33% 3/48 1/16 16/16
coding-v1-r8 230 6 2.61% 6/48 3/16 14/16
coding-v1-r9 227 20 8.81% 18/48 7/16 14/16

3. Did uptake really rise from r8 to r9?

unit of analysis comparison test p
48 cells (as published) 6/230 vs 20/227 two-proportion z, z=2.862 4.21e-3
16 questions 3/16 vs 7/16 Fisher exact 0.252
16 questions, paired 4 gained, 0 lost McNemar exact 0.125

The cell-level test treats 48 correlated cells as independent and overstates its own precision. Clustered on the question the shift is directionally clean — nothing lost the graph — but not resolvable at n=16.

Per-question detail

question r8 reps using a graph r9 reps using a graph
A1 0/3 1/3
A2 0/3 0/3
A3 0/3 0/3
A4 2/3 3/3
B1 3/3 3/3
B2 1/3 3/3
B3 0/3 0/3
L1 0/3 0/3
L2 0/3 2/3
L3 0/3 3/3
S1 0/3 0/3
S2 0/3 0/3
S3 0/3 3/3
T1 0/3 0/3
T3 0/3 0/3
T4 0/3 0/3

4. When is the choice made?

First executed tool per cell. In both runs exactly one cell of 48 begins with text search and later reaches for a graph — the agent commits on turn 0 and does not revise.

coding-v1-r8 — opened with a graph tool: 5; switched to one later: 1

first tool cells
Grep 35
Glob 6
mcp__codegraph__codegraph_explore 4
Read 2
mcp__graphify__query_graph 1

coding-v1-r9 — opened with a graph tool: 17; switched to one later: 1

first tool cells
Grep 22
mcp__codegraph__codegraph_explore 17
Read 5
Glob 4

5. Executed-tool histogram — coding-v1-r9

tool calls
Grep 140
Read 50
mcp__codegraph__codegraph_explore 20
Glob 17

6. Cost — the cross-cell comparison and the paired one

Comparing cells directly appears to vindicate the preference decisively:

cells n median content tokens median cost median score
used a graph tool 18 124312 $0.167 1.00
grep only 30 54201 $0.076 1.00

That comparison is confounded: the graph is called on a stable subset of questions, so this compares questions rather than strategies. The unconfounded version restricts to same-run, same-question pairs where some reps used a graph and some did not.

run / question graph reps, mean tokens grep reps, mean tokens ratio graph score grep score
coding-v1-r6 / A4 115299 103852 1.11× 1.00 1.00
coding-v1-r6 / B1 241819 219115 1.10× 1.00 1.00
coding-v1-r7 / A1 102137 94991 1.08× 1.00 1.00
coding-v1-r7 / A4 61006 90524 0.67× 0.91 0.88
coding-v1-r8 / A4 59690 114938 0.52× 0.91 0.82
coding-v1-r8 / B2 93101 140677 0.66× 1.00 1.00
coding-v1-r9 / A1 34589 82153 0.42× 0.65 1.00
coding-v1-r9 / L2 71148 79461 0.90× 1.00 1.00
coding-v1-abdesc-actionable / A1 172420 86227 2.00× 1.00 1.00
coding-v1-abdesc-actionable / A2 148157 170776 0.87× 1.00 1.00
coding-v1-abdesc-actionable / T3 61099 63485 0.96× 1.00 1.00

Paired comparisons available: 11.

statistic value
mean token ratio graph : grep 0.935×
95% CI [0.685, 1.186]
graph cheaper in 7/11 pairs
sign test, two-sided p = 0.549
mean score delta -0.021

The interval spans 1.0 and the sign test is nowhere near significance. Controlling for the question, the token difference between reaching for a graph and not reaching for one is indistinguishable from zero — the graph is neither reliably cheaper nor reliably dearer. That is the honest reading, and it supersedes an earlier n=8 pass of this same comparison which gave 0.81× and was written up as a reversal. Three more pairs moved the mean to 0.94× and widened the interval across 1.0; a mean of eight ratios quoted without dispersion read as a result it never was.

What survives is the negative half: the cross-cell table above is confounded by question difficulty and cannot support the claim that using a graph costs 2.29× more. That inference is retired. It is NOT replaced by the opposite one.

7. Are the questions keyword-addressable?

For each question, a literal token appearing verbatim in its own prompt is searched across the repository and the hit set checked against that question's ground-truth evidence paths. git grep -l is used, so "rank" is path order, not relevance — the claim being tested is reachability, not ranking.

question token from the prompt files hit ground truth reachable
L1 MANAGED_MCP_KEYS 18 yes
L2 summaryStats 20 yes
L3 system-health 231 yes
S1 config/code-graph.json 30 yes
S2 supervisord 77 yes
S3 transport 164 yes
B1 mcp.tools 13 yes
B2 12435 258 yes
B3 .codegraph 24 yes
A1 .observations 240 yes
A2 content tokens 33 yes
A3 code-graph.json 33 yes
A4 CodeGraph 252 yes
T1 Memgraph 107 no
T3 payment reconciliation 12 — (abstain: nothing to find)
T4 CODEGRAPH_MAX_DEPTH 15 — (abstain: nothing to find)

13 of 14 gradeable non-abstain questions have their ground truth in the hit set of a single literal grep of a token the question hands the agent. This is a benchmark of keyword-addressable retrieval, and grep is the optimal instrument for that.

8. Is the answer even in the index?

Graphify index at commit b6ed72ff68894caa077fe08f49112105afe8182a: 58124 nodes over 4347 distinct source files.

question evidence files in the index missing
L1 1 1 —
L2 3 3 —
L3 1 1 —
S1 2 1 config/code-graph.json
S2 2 1 docker/supervisord.conf
S3 1 0 config/code-graph.json
B1 3 3 —
B2 3 3 —
B3 2 1 docker/docker-compose.yml
A1 1 0 docker/docker-compose.yml
A2 2 2 —
A3 2 0 config/code-graph.json, config/kgbench/arms.json
A4 1 0 config/code-graph.json
T1 1 1 —

17/25 ground-truth evidence files are present in the Graphify index. 7 of 14 questions have at least one required fact the graph cannot see at any price.

Does coverage drive the choice? (no)

questions graph calls / total calls
ground truth fully in the index 13 / 140 = 9.29%
ground truth partially or not in the index 7 / 73 = 9.59%

No difference. The agent is not avoiding the graph where the graph is blind — per section 4 it commits before it could possibly find out.