tool-selection.md — data appendix¶
Generated by scripts/analysis/tool-selection-analysis.mjs. Do not hand-edit.
Runs analysed: coding-v1-r6, coding-v1-r7, coding-v1-x2, coding-v1-r8, coding-v1-r9. Arm: hybrid unless stated.
1. Tool telemetry by agent¶
A run with 0 recorded calls is not a run with no tool use — it is a run with no instrumentation. Every later section is therefore claude-only.
| run | agent | cells | recorded tool calls | tool_audit |
|---|---|---|---|---|
coding-v1-r9 | claude | 48 | 227 | audited |
coding-v1-r9 | copilot | 48 | 0 | unavailable |
coding-v1-r9 | opencode | 48 | 0 | unavailable |
coding-v1-r8 | claude | 48 | 230 | audited |
coding-v1-r8 | copilot | 48 | 0 | unavailable |
coding-v1-r8 | opencode | 48 | 0 | unavailable |
2. Graph uptake per run¶
unanimous counts questions whose reps all agree (all used a graph, or none did). It stays at 14–16 of 16 in every run: tool choice is a near-deterministic function of the question rather than a rate.
| run | tool calls | graph calls | rate | cells w/ graph | questions w/ graph | unanimous |
|---|---|---|---|---|---|---|
coding-v1-r6 | 322 | 3 | 0.93% | 3/76 | 2/16 | 14/16 |
coding-v1-r7 | 306 | 5 | 1.63% | 5/76 | 2/16 | 14/16 |
coding-v1-x2 | 226 | 3 | 1.33% | 3/48 | 1/16 | 16/16 |
coding-v1-r8 | 230 | 6 | 2.61% | 6/48 | 3/16 | 14/16 |
coding-v1-r9 | 227 | 20 | 8.81% | 18/48 | 7/16 | 14/16 |
3. Did uptake really rise from r8 to r9?¶
| unit of analysis | comparison | test | p |
|---|---|---|---|
| 48 cells (as published) | 6/230 vs 20/227 | two-proportion z, z=2.862 | 4.21e-3 |
| 16 questions | 3/16 vs 7/16 | Fisher exact | 0.252 |
| 16 questions, paired | 4 gained, 0 lost | McNemar exact | 0.125 |
The cell-level test treats 48 correlated cells as independent and overstates its own precision. Clustered on the question the shift is directionally clean — nothing lost the graph — but not resolvable at n=16.
Per-question detail¶
| question | r8 reps using a graph | r9 reps using a graph |
|---|---|---|
| A1 | 0/3 | 1/3 |
| A2 | 0/3 | 0/3 |
| A3 | 0/3 | 0/3 |
| A4 | 2/3 | 3/3 |
| B1 | 3/3 | 3/3 |
| B2 | 1/3 | 3/3 |
| B3 | 0/3 | 0/3 |
| L1 | 0/3 | 0/3 |
| L2 | 0/3 | 2/3 |
| L3 | 0/3 | 3/3 |
| S1 | 0/3 | 0/3 |
| S2 | 0/3 | 0/3 |
| S3 | 0/3 | 3/3 |
| T1 | 0/3 | 0/3 |
| T3 | 0/3 | 0/3 |
| T4 | 0/3 | 0/3 |
4. When is the choice made?¶
First executed tool per cell. In both runs exactly one cell of 48 begins with text search and later reaches for a graph — the agent commits on turn 0 and does not revise.
coding-v1-r8 — opened with a graph tool: 5; switched to one later: 1
| first tool | cells |
|---|---|
Grep | 35 |
Glob | 6 |
mcp__codegraph__codegraph_explore | 4 |
Read | 2 |
mcp__graphify__query_graph | 1 |
coding-v1-r9 — opened with a graph tool: 17; switched to one later: 1
| first tool | cells |
|---|---|
Grep | 22 |
mcp__codegraph__codegraph_explore | 17 |
Read | 5 |
Glob | 4 |
5. Executed-tool histogram — coding-v1-r9¶
| tool | calls |
|---|---|
Grep | 140 |
Read | 50 |
mcp__codegraph__codegraph_explore | 20 |
Glob | 17 |
6. Cost — the cross-cell comparison and the paired one¶
Comparing cells directly appears to vindicate the preference decisively:
| cells | n | median content tokens | median cost | median score |
|---|---|---|---|---|
| used a graph tool | 18 | 124312 | $0.167 | 1.00 |
| grep only | 30 | 54201 | $0.076 | 1.00 |
That comparison is confounded: the graph is called on a stable subset of questions, so this compares questions rather than strategies. The unconfounded version restricts to same-run, same-question pairs where some reps used a graph and some did not.
| run / question | graph reps, mean tokens | grep reps, mean tokens | ratio | graph score | grep score |
|---|---|---|---|---|---|
coding-v1-r6 / A4 | 115299 | 103852 | 1.11× | 1.00 | 1.00 |
coding-v1-r6 / B1 | 241819 | 219115 | 1.10× | 1.00 | 1.00 |
coding-v1-r7 / A1 | 102137 | 94991 | 1.08× | 1.00 | 1.00 |
coding-v1-r7 / A4 | 61006 | 90524 | 0.67× | 0.91 | 0.88 |
coding-v1-r8 / A4 | 59690 | 114938 | 0.52× | 0.91 | 0.82 |
coding-v1-r8 / B2 | 93101 | 140677 | 0.66× | 1.00 | 1.00 |
coding-v1-r9 / A1 | 34589 | 82153 | 0.42× | 0.65 | 1.00 |
coding-v1-r9 / L2 | 71148 | 79461 | 0.90× | 1.00 | 1.00 |
coding-v1-abdesc-actionable / A1 | 172420 | 86227 | 2.00× | 1.00 | 1.00 |
coding-v1-abdesc-actionable / A2 | 148157 | 170776 | 0.87× | 1.00 | 1.00 |
coding-v1-abdesc-actionable / T3 | 61099 | 63485 | 0.96× | 1.00 | 1.00 |
Paired comparisons available: 11.
| statistic | value |
|---|---|
| mean token ratio graph : grep | 0.935× |
| 95% CI | [0.685, 1.186] |
| graph cheaper in | 7/11 pairs |
| sign test, two-sided | p = 0.549 |
| mean score delta | -0.021 |
The interval spans 1.0 and the sign test is nowhere near significance. Controlling for the question, the token difference between reaching for a graph and not reaching for one is indistinguishable from zero — the graph is neither reliably cheaper nor reliably dearer. That is the honest reading, and it supersedes an earlier n=8 pass of this same comparison which gave 0.81× and was written up as a reversal. Three more pairs moved the mean to 0.94× and widened the interval across 1.0; a mean of eight ratios quoted without dispersion read as a result it never was.
What survives is the negative half: the cross-cell table above is confounded by question difficulty and cannot support the claim that using a graph costs 2.29× more. That inference is retired. It is NOT replaced by the opposite one.
7. Are the questions keyword-addressable?¶
For each question, a literal token appearing verbatim in its own prompt is searched across the repository and the hit set checked against that question's ground-truth evidence paths. git grep -l is used, so "rank" is path order, not relevance — the claim being tested is reachability, not ranking.
| question | token from the prompt | files hit | ground truth reachable |
|---|---|---|---|
| L1 | MANAGED_MCP_KEYS | 18 | yes |
| L2 | summaryStats | 20 | yes |
| L3 | system-health | 231 | yes |
| S1 | config/code-graph.json | 30 | yes |
| S2 | supervisord | 77 | yes |
| S3 | transport | 164 | yes |
| B1 | mcp.tools | 13 | yes |
| B2 | 12435 | 258 | yes |
| B3 | .codegraph | 24 | yes |
| A1 | .observations | 240 | yes |
| A2 | content tokens | 33 | yes |
| A3 | code-graph.json | 33 | yes |
| A4 | CodeGraph | 252 | yes |
| T1 | Memgraph | 107 | no |
| T3 | payment reconciliation | 12 | — (abstain: nothing to find) |
| T4 | CODEGRAPH_MAX_DEPTH | 15 | — (abstain: nothing to find) |
13 of 14 gradeable non-abstain questions have their ground truth in the hit set of a single literal grep of a token the question hands the agent. This is a benchmark of keyword-addressable retrieval, and grep is the optimal instrument for that.
8. Is the answer even in the index?¶
Graphify index at commit b6ed72ff68894caa077fe08f49112105afe8182a: 58124 nodes over 4347 distinct source files.
| question | evidence files | in the index | missing |
|---|---|---|---|
| L1 | 1 | 1 | — |
| L2 | 3 | 3 | — |
| L3 | 1 | 1 | — |
| S1 | 2 | 1 | config/code-graph.json |
| S2 | 2 | 1 | docker/supervisord.conf |
| S3 | 1 | 0 | config/code-graph.json |
| B1 | 3 | 3 | — |
| B2 | 3 | 3 | — |
| B3 | 2 | 1 | docker/docker-compose.yml |
| A1 | 1 | 0 | docker/docker-compose.yml |
| A2 | 2 | 2 | — |
| A3 | 2 | 0 | config/code-graph.json, config/kgbench/arms.json |
| A4 | 1 | 0 | config/code-graph.json |
| T1 | 1 | 1 | — |
17/25 ground-truth evidence files are present in the Graphify index. 7 of 14 questions have at least one required fact the graph cannot see at any price.
Does coverage drive the choice? (no)¶
| questions | graph calls / total calls |
|---|---|
| ground truth fully in the index | 13 / 140 = 9.29% |
| ground truth partially or not in the index | 7 / 73 = 9.59% |
No difference. The agent is not avoiding the graph where the graph is blind — per section 4 it commits before it could possibly find out.