docs: one entry point for the three feedback metrics, and how to run them (CG-11)
Three per-metric docs told a maintainer what each number means; none said which one answers which question, which harness produces it, or how to read the arm table. agent-eval-feedback-metrics.md is that page — the metric → question map, when to reach for ab-new-vs-baseline.sh (isolates a change, both arms codegraph-on) versus run-all.sh (with vs without, a different question) versus bench-readme.sh, the worked CG-22 express table where all three read together, and the bucket → fix mapping. Not a fourth restatement: the derivations stay where they are and each doc now points here. The caveats that change how the summary table is read are carried over rather than dropped — allocation efficiency is relative (attribution is by citation, so same-question builds only, and never "codegraph wastes N%"), occupancy shares are Claude Code / 200k and do not transfer between hosts while the arm ratio does, sufficient is not correct, small-n throughout. Plus the contamination row, which means different things in the two harnesses and is the first thing to look at in both. Also records that the CG-8 7-repo bucket block no longer re-derives: bench-readme.sh overwrites /tmp/ab-readme, so the swept logs are gone. The current logs give a different distribution over the same 62 calls, and the CG-8-era and current classifiers agree exactly on them — so nothing moved under the metric, the corpus did. CG-13 re-establishes the baseline.
This commit is contained in:
@@ -58,6 +58,12 @@ scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>
|
||||
codegraph-tool calls, duration, **total cost**.
|
||||
- Interactive (`parse-session.mjs`): the `VERDICT: codegraph_explore used Nx |
|
||||
Read N | Grep/Bash N` and `TOKENS:` lines.
|
||||
- Both paths also print the three feedback metrics — residual context occupancy,
|
||||
explore sufficiency, allocation efficiency — and a headless A/B ends with a
|
||||
side-by-side `ARM COMPARISON` table. Report that table, and check its
|
||||
contamination row first: `CLI calls that RETURNED output` > 0 means the arm
|
||||
reached codegraph through Bash and its numbers are void. How to read the rest:
|
||||
`docs/benchmarks/agent-eval-feedback-metrics.md`.
|
||||
|
||||
Lead with cost + tool/Read counts — they are the reliable signals; raw token
|
||||
in/out are confounded by subagent delegation and prompt caching. State whether
|
||||
|
||||
Reference in New Issue
Block a user