Files
codegraph/scripts/agent-eval/run-all.sh
T
Colby McHenry 4080b7501e test(agent-eval): measure residual context occupancy, over multi-turn sessions (CG-7)
The A/B arms reported cost, tokens, time and tool counts for one headless
question. They could not report what issue #1500 actually measured: how much of
the context window a tool's responses still occupy once the question is
answered, which every later turn is then charged for.

parse-run.mjs now measures that. Tokens are measured, not estimated: for each
assistant request, input + cache_read + cache_creation is the exact token count
of its whole prompt, so consecutive requests differ by exactly what was appended
between them. That delta is priced against the characters in the gap, calibrated
on gaps that are >=80% tool result. Explore output lands near 2.3 chars/token, so
the usual bytes/4 estimate would have under-counted it by ~40%.

Content also leaves the window, so residual is tracked apart from contributed:
a compact_boundary clears the resident set, and a mid-run context drop is
micro-compaction, which sheds the oldest tool results first and is applied FIFO.

run-all.sh takes "Q1||Q2||Q3" and runs them as one resumed session, one segment
file per turn; parse-run.mjs stitches the segments back together. bench-readme.sh
now runs each README repo as a three-turn session (CG_TURNS=1 restores the
single-question form). parse-bench-readme.mjs reports the arms' retrieval
residual side by side -- codegraph's responses against the without-arm's
Read/Grep/Bash -- in absolute tokens, share of context, and share of window, and
says so explicitly when the rows it aggregated were single-turn.

Two transcript traps are handled and documented at the call site: Claude Code
emits one assistant event per content block, all carrying the same usage (summing
per event double-counts every turn with both thinking and a tool_use), and the
streamed output_tokens is a partial snapshot.

Occupancy lives in parse-run.mjs and is imported by the aggregator rather than
extracted to a module -- a new scripts/agent-eval/*.mjs scores into the
self-query fixture's own corpus and moves its numbers.
2026-08-04 14:35:56 -05:00

123 lines
5.4 KiB
Bash
Executable File

#!/usr/bin/env bash
# With/without A/B (and optional interactive) eval for a codegraph version on a
# repo. Codegraph is the ONLY variable: both arms launch claude with
# --strict-mcp-config — with = codegraph-only MCP (pointed at $CG_BIN),
# without = empty MCP. Built-in Read/Grep/Bash stay available in both arms.
#
# Usage: run-all.sh <repo-path> "<question>" [headless|tmux|all]
#
# MULTI-TURN: separate questions with "||" to run them as ONE session —
# run-all.sh <repo> "How does X work?||Where is Y handled in that path?"
# Turn 1 runs normally; every later turn `--resume`s the same session, so the
# earlier turns' tool output is still in the window (that is the whole point:
# residual context occupancy, the cost a single-question run cannot see).
# Segments land in run-<label>.jsonl, run-<label>.t2.jsonl, … and parse-run.mjs
# stitches them back into one session.
#
# Env: CG_BIN codegraph binary (default: command -v codegraph)
# AGENT_EVAL_OUT output dir (default: /tmp/agent-eval)
# MODEL / EFFORT claude model/effort (default: sonnet / high — the
# standing A/B policy; see CLAUDE.md, don't raise)
set -uo pipefail
REPO="${1:?usage: run-all.sh <repo-path> \"<question>\" [headless|tmux|all]}"
Q="${2:?question required}"
MODE="${3:-headless}"
# Split "Q1||Q2||Q3" into turns (kept bash-3.2-safe: macOS ships 3.2).
TURNS=()
rest="$Q"
while [ "$rest" != "${rest#*||}" ]; do
TURNS+=("${rest%%||*}")
rest="${rest#*||}"
done
TURNS+=("$rest")
CG_BIN="${CG_BIN:-$(command -v codegraph)}"
OUT="${AGENT_EVAL_OUT:-/tmp/agent-eval}"
HARNESS="$(cd "$(dirname "$0")" && pwd)"
mkdir -p "$OUT"
# Neutralize any ambient CodeGraph prompt-hook (~/.claude) in BOTH arms:
# the hook injects codegraph context into every prompt, which contaminates
# the without-arm (free structural context) and double-counts the with-arm.
# The A/B's only variable must be the MCP server wired below.
export CODEGRAPH_NO_PROMPT_HOOK=1
[ -n "$CG_BIN" ] || { echo "no codegraph binary on PATH (set CG_BIN)"; exit 1; }
[ -d "$REPO/.codegraph" ] || { echo "no .codegraph index at $REPO — index it first"; exit 1; }
case "$MODE" in headless|tmux|all) ;; *) echo "mode must be headless|tmux|all (got '$MODE')"; exit 1;; esac
# MCP config files (path form avoids inline-JSON quoting through tmux).
cat > "$OUT/mcp-codegraph.json" <<JSON
{"mcpServers":{"codegraph":{"command":"$CG_BIN","args":["serve","--mcp","--path","$REPO"]}}}
JSON
echo '{"mcpServers":{}}' > "$OUT/mcp-empty.json"
echo "###### codegraph: $CG_BIN"
echo "###### repo: $REPO"
echo "###### turns: ${#TURNS[@]}"
for t in "${TURNS[@]}"; do echo "###### - $t"; done
echo
# Pull the session id out of a segment's result event so the next turn can
# --resume it (rather than minting a --session-id, which needs a valid uuid).
session_id_of() {
node -e '
const fs=require("fs");
for (const l of fs.readFileSync(process.argv[1],"utf8").split("\n").reverse()) {
if (!l) continue; let e; try { e=JSON.parse(l) } catch { continue }
if (e.session_id) { console.log(e.session_id); break }
}' "$1" 2>/dev/null
}
# Headless arm: claude -p with stream-json -> exact tool sequence + tokens/cost
# + residual context occupancy. One session, one segment file per turn.
headless() {
local label="$1" cfg="$2"
echo "############################## HEADLESS [$label] ##############################"
local sid="" seg=0 out="" files=()
: > "$OUT/run-$label.err"
for q in "${TURNS[@]}"; do
seg=$((seg + 1))
out="$OUT/run-$label.jsonl"
[ "$seg" -gt 1 ] && out="$OUT/run-$label.t$seg.jsonl"
local resume=()
[ -n "$sid" ] && resume=(--resume "$sid")
( cd "$REPO" && claude -p "$q" \
--output-format stream-json --verbose \
--permission-mode bypassPermissions \
--model "${MODEL:-sonnet}" --effort "${EFFORT:-high}" \
--max-budget-usd 4 \
--strict-mcp-config --mcp-config "$cfg" \
${resume[@]+"${resume[@]}"} \
</dev/null > "$out" 2>>"$OUT/run-$label.err" )
echo "exit $? -> $out ($(wc -l < "$out" | tr -d ' ') lines) [turn $seg/${#TURNS[@]}]"
files+=("$out")
sid="$(session_id_of "$out")"
if [ -z "$sid" ] && [ "$seg" -lt "${#TURNS[@]}" ]; then
echo " WARN: no session_id in $out — later turns would start a FRESH context; stopping this arm"
break
fi
done
tail -2 "$OUT/run-$label.err" 2>/dev/null
node "$HARNESS/parse-run.mjs" "${files[@]}" 2>&1 || true
echo
}
if [ "$MODE" = headless ] || [ "$MODE" = all ]; then
headless "headless-with" "$OUT/mcp-codegraph.json"
headless "headless-without" "$OUT/mcp-empty.json"
fi
if [ "$MODE" = tmux ] || [ "$MODE" = all ]; then
echo "############################## INTERACTIVE [with] ##############################"
CLAUDE_EXTRA_ARGS="--model ${MODEL:-sonnet} --effort ${EFFORT:-high} --strict-mcp-config --mcp-config $OUT/mcp-codegraph.json" \
bash "$HARNESS/itrun.sh" "$REPO" "int-with" "${TURNS[0]}" 2>&1 || echo "[itrun WITH failed]"
echo
echo "############################## INTERACTIVE [without] ##############################"
CLAUDE_EXTRA_ARGS="--model ${MODEL:-sonnet} --effort ${EFFORT:-high} --strict-mcp-config --mcp-config $OUT/mcp-empty.json" \
bash "$HARNESS/itrun.sh" "$REPO" "int-without" "${TURNS[0]}" 2>&1 || echo "[itrun WITHOUT failed]"
echo
fi
echo "############################## RUN-ALL COMPLETE ##############################"