All agent A/B arms now run claude --model sonnet --effort high by default
(MODEL/EFFORT env overrides exist). Sonnet is the deliberate floor model:
codegraph's users attach whatever host they already run (Cursor Composer,
Gemini, ...), and a stronger model's tool-use masks the salience problems a
weaker one exposes — what lands on Sonnet generalizes up; Opus/Fable-only
wins don't generalize down. Policy recorded in CLAUDE.md's validation
methodology; 11 hardcoded --model opus call sites across 9 eval scripts
switched to the env-overridable default.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Replaces the old interactive publish.js script with two Claude skills and
a full agent-evaluation harness:
- `.claude/skills/audit/` — `/audit` skill drives `scripts/agent-eval/audit.sh`
to benchmark retrieval quality (with vs. without codegraph) on a chosen
real-world repo from the new `corpus.json` (17 repos across 14 languages).
- `.claude/skills/publish/` — `/publish` skill orchestrates the full release
workflow (preflight → changelog → confirmation gate → bump/build → npm
publish → GitHub release), replacing `publish.js`.
- `scripts/agent-eval/` — headless (`run-agent.sh`, `run-all.sh`) and
interactive tmux (`itrun.sh`) harnesses with stream-json parsers
(`parse-run.mjs`, `parse-session.mjs`) that report tool calls, token
usage, and a VERDICT line summarising codegraph_explore vs. Read/Grep counts.
- `run-interactive-test.md` — documents the two harnesses, idle-detection
approach, and what "good" agent behavior looks like after explore-first
guidance.