Skip to content

Tool-description A/B — does advertising change uptake?

Generated by scripts/analysis/tool-description-ab.mjs. Do not hand-edit.

control treatment
run coding-v1-r9 coding-v1-abdesc-actionable
GRAPHIFY_TOOL_DESCRIPTIONS terse actionable
cells (hybrid / claude) 48 48
corpus commit d8a9b0647 d8a9b0647 (pinned)

1. Primary outcome — Graphify uptake

metric terse actionable change
Graphify calls 0 11 0 → 11 (0.00% → 5.09%)
cells using Graphify 0 10 0 → 10
CodeGraph calls 20 15 0.75× (8.81% → 6.94%)
all graph calls 20 26 1.30× (8.81% → 12.04%)
cells using any graph 18 24 1.33×
total tool calls 227 216 —
Grep 140 130 0.93×
Read 50 43 0.86×
Glob 17 17 1.00×

2. Significance, clustered on the question

Questions eliciting at least one Graphify call: 0/16 terse → 6/16 actionable.

test value
Fisher exact, two-sided 0.018
McNemar exact, paired (6 gained, 0 lost) 0.031
question terse reps w/ Graphify actionable reps w/ Graphify
A1 0/3 2/3 ← new
A2 0/3 2/3 ← new
A3 0/3 0/3
A4 0/3 0/3
B1 0/3 0/3
B2 0/3 2/3 ← new
B3 0/3 0/3
L1 0/3 0/3
L2 0/3 1/3 ← new
L3 0/3 1/3 ← new
S1 0/3 0/3
S2 0/3 0/3
S3 0/3 0/3
T1 0/3 0/3
T3 0/3 2/3 ← new
T4 0/3 0/3

3. First tool call — is the change in the opening move?

first tool terse actionable
Glob 4 2
Grep 22 19
Read 5 6
mcp__codegraph__codegraph_explore 17 14
mcp__graphify__query_graph 0 7

4. Did it cost correctness, tokens or time?

Scored by the deterministic checklist grader. The judge was disabled for the treatment run (--no-judge): the proxy could not serve the same judge model the control used, and a judge column graded by a different model is not comparable to one that was not.

Latency is NOT comparable between these two runs and is shown for completeness only. The control was measured on an otherwise idle machine; the treatment ran while a background consolidator held a second agent process open. kgbench runs its cells strictly sequentially precisely because CPU contention corrupts wall-clock — see "When the machine lied" in the report. Token counts and tool counts are unaffected by contention; wall seconds are. The primary outcome of this A/B is the tool counts in section 1.

metric terse actionable
median score 1.00 1.00
mean score 0.960 0.989
cells answered 48/48 48/48
median content tokens 78855 88608
median wall seconds 14.4 16.4
median tool calls 4 4
hallucinations 1 0