chore(agent-eval): standing A/B model policy — sonnet + high effort, never Opus/Fable (#816)
All agent A/B arms now run claude --model sonnet --effort high by default (MODEL/EFFORT env overrides exist). Sonnet is the deliberate floor model: codegraph's users attach whatever host they already run (Cursor Composer, Gemini, ...), and a stronger model's tool-use masks the salience problems a weaker one exposes — what lands on Sonnet generalizes up; Opus/Fable-only wins don't generalize down. Policy recorded in CLAUDE.md's validation methodology; 11 hardcoded --model opus call sites across 9 eval scripts switched to the env-overridable default. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
823ffd1c3d
commit
0682681175
@@ -64,7 +64,7 @@ run(){ # label, withCodegraph(0/1)
|
||||
prewarm "$tgt"
|
||||
else cp "$OUT/mcp-empty.json" "$cfg"; fi
|
||||
( cd "$tgt" && claude -p "$Q" --output-format stream-json --verbose \
|
||||
--permission-mode bypassPermissions --model opus --max-budget-usd 4 \
|
||||
--permission-mode bypassPermissions --model "${MODEL:-sonnet}" --effort "${EFFORT:-high}" --max-budget-usd 4 \
|
||||
--strict-mcp-config --mcp-config "$cfg" </dev/null > "$OUT/$label-$i.jsonl" 2>"$OUT/$label-$i.err" )
|
||||
echo "[$label] run $i:"; analyze "$OUT/$label-$i.jsonl"
|
||||
if [ -n "$BUILD_CMD" ]; then ( cd "$tgt" && eval "$BUILD_CMD" >/dev/null 2>&1 && echo " build: PASS" || echo " build: FAIL" ); fi
|
||||
|
||||
Reference in New Issue
Block a user