Swept every benchmark doc for figures produced off `result.usage` and fixed
the ones that had raw logs to re-derive from.
residual-context-occupancy.md — the sonnet 3-turn throughput table. Re-derived
from the preserved logs: tokens saved 23% -> 56%, and vscode's "98% MORE tokens
with codegraph" was never real, it is 41% fewer. Cost, time and tool calls were
never affected by this field and are unchanged. The occupancy table itself is
measured off the timeline, so every number in it stands -- including the 82%
higher residual, which is the finding the document exists for.
call-sequence-analysis.md — this doc DIAGNOSED the bug and its reproduce block
claimed the aggregator summed per-turn tokens. It did not, until 04c0f8e. Noted,
with the three wrong results the gap produced: the excalidraw cut recorded here,
the sonnet campaign, and the Opus re-measure that invented a token regression.
answer-directly-vs-explore-agent.md — build 0.9.4, 2026-05-24, raw logs gone.
Cannot be re-derived, so flagged rather than silently left or invented: its
token figure is indicative, its turn/read/context findings do not depend on the
broken field and stand.
The remaining benchmark docs (allocation-ab-1500, dedup-cg20, allocation-
efficiency, feedback-metrics) carry no throughput tables — checked, clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replace the stale "## CodeGraph" example block (NEVER call explore directly /
ALWAYS spawn an Explore agent) and the How-It-Works diagram with the validated
"answer directly" guidance, and add codegraph_context/trace/explore to the tool
table. Interactive A/B (Excalidraw + VS Code, n=3/arm) shows direct codegraph
answering beats Explore-agent delegation at every scale: main-session context is
~scale-invariant (~50k), with 0 reads vs 17-26 and ~28% fewer tokens. Record the
writeup under docs/benchmarks/answer-directly-vs-explore-agent.md.
Docs-only; stays on 0.9.4 (no version bump).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>