Running scripts/agent-eval against a `claude -p` spawned from within a Claude Code session (nested, e.g. from a Bash tool call) makes the codegraph MCP attach unreliable: the server is healthy (full handshake ~165ms) but the nested client marks it status:"pending"/0-tools under CPU/timing contention, so the agent silently runs with no codegraph. NO_DAEMON + `< /dev/null` don't fix it — it's the nested client, not the server. Documented in CLAUDE.md's validation methodology. Adds ab-new-vs-baseline.sh: A/Bs a retrieval/steering change as new-build vs baseline-build (both codegraph-on, isolating the change — vs run-all.sh's with-vs-without), on a throwaway copy of an indexed repo. Run it in a real terminal. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>