Generated by scripts/analysis/tool-description-ab.mjs. Do not hand-edit.
| control | treatment |
| run | coding-v1-r9 | coding-v1-abdesc-actionable |
GRAPHIFY_TOOL_DESCRIPTIONS | terse | actionable |
| cells (hybrid / claude) | 48 | 48 |
| corpus commit | d8a9b0647 | d8a9b0647 (pinned) |
1. Primary outcome — Graphify uptake
| metric | terse | actionable | change |
| Graphify calls | 0 | 11 | 0 → 11 (0.00% → 5.09%) |
| cells using Graphify | 0 | 10 | 0 → 10 |
| CodeGraph calls | 20 | 15 | 0.75× (8.81% → 6.94%) |
| all graph calls | 20 | 26 | 1.30× (8.81% → 12.04%) |
| cells using any graph | 18 | 24 | 1.33× |
| total tool calls | 227 | 216 | — |
Grep | 140 | 130 | 0.93× |
Read | 50 | 43 | 0.86× |
Glob | 17 | 17 | 1.00× |
2. Significance, clustered on the question
Questions eliciting at least one Graphify call: 0/16 terse → 6/16 actionable.
| test | value |
| Fisher exact, two-sided | 0.018 |
| McNemar exact, paired (6 gained, 0 lost) | 0.031 |
| question | terse reps w/ Graphify | actionable reps w/ Graphify |
| A1 | 0/3 | 2/3 ← new |
| A2 | 0/3 | 2/3 ← new |
| A3 | 0/3 | 0/3 |
| A4 | 0/3 | 0/3 |
| B1 | 0/3 | 0/3 |
| B2 | 0/3 | 2/3 ← new |
| B3 | 0/3 | 0/3 |
| L1 | 0/3 | 0/3 |
| L2 | 0/3 | 1/3 ← new |
| L3 | 0/3 | 1/3 ← new |
| S1 | 0/3 | 0/3 |
| S2 | 0/3 | 0/3 |
| S3 | 0/3 | 0/3 |
| T1 | 0/3 | 0/3 |
| T3 | 0/3 | 2/3 ← new |
| T4 | 0/3 | 0/3 |
| first tool | terse | actionable |
Glob | 4 | 2 |
Grep | 22 | 19 |
Read | 5 | 6 |
mcp__codegraph__codegraph_explore | 17 | 14 |
mcp__graphify__query_graph | 0 | 7 |
4. Did it cost correctness, tokens or time?
Scored by the deterministic checklist grader. The judge was disabled for the treatment run (--no-judge): the proxy could not serve the same judge model the control used, and a judge column graded by a different model is not comparable to one that was not.
Latency is NOT comparable between these two runs and is shown for completeness only. The control was measured on an otherwise idle machine; the treatment ran while a background consolidator held a second agent process open. kgbench runs its cells strictly sequentially precisely because CPU contention corrupts wall-clock — see "When the machine lied" in the report. Token counts and tool counts are unaffected by contention; wall seconds are. The primary outcome of this A/B is the tool counts in section 1.
| metric | terse | actionable |
| median score | 1.00 | 1.00 |
| mean score | 0.960 | 0.989 |
| cells answered | 48/48 | 48/48 |
| median content tokens | 78855 | 88608 |
| median wall seconds | 14.4 | 16.4 |
| median tool calls | 4 | 4 |
| hallucinations | 1 | 0 |