perf(mcp): answer-directly steering — ~35% cheaper, ~70% fewer tool calls (#224)
* perf(mcp): steer agents to answer directly instead of delegating to subagents CodeGraph beats native grep/read on cost only when the agent queries it directly. When the agent delegates to file-reading sub-agents, those sub-agents read files regardless of the index, so CodeGraph becomes net overhead on top of the reads. The install templates even told agents to "spawn a subagent for explore-class questions" — the expensive path. Changes: - server-instructions + both install templates: add an "Answer directly — don't delegate exploration" directive; reposition codegraph_explore as the efficient one-call multi-symbol tool (was: "spawn a subagent for it"). - codegraph_explore: hard-cap output to its adaptive budget (it overran, ~30k vs a 28k cap) and tighten the medium tier (28k->13k). - codegraph_node: return a member outline for container kinds instead of the full class body. Rigorous N>=4-per-arm warm-block benchmark (median total_cost_usd): excalidraw (~600 files): WITH $0.54 vs native $1.02 (-47%) vscode (~10k files): WITH $0.41 vs native $0.72 (-42%) ky (~25 files): WITH $0.46 vs native $0.44 (wash) Answers were equal-or-better (correct, file:line-cited) with ~6x fewer tool calls; the directive drove the direct path on 14/14 codegraph runs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(readme): rebuild benchmark with real-world repos + cost/token/time/tool savings Replace the "Claude Code (Python+Rust/Java)" rows — which benchmarked the Claude Code CLI repo, not real codebases in those languages — with real open-source projects per language: Django (Python), Tokio (Rust), OkHttp (Java), Gin (Go), plus Alamofire (Swift) and the existing TypeScript repos (VS Code, Excalidraw). The table now reports all four savings the change targets — cost, tokens, time, tool calls — as the median of 4 runs per arm (Claude Opus 4.7, headless claude -p, with vs empty MCP config). Averages across the 7 repos: 35% cheaper, 59% fewer tokens, 49% faster, 70% fewer tool calls. Adds a methodology note and raw WITH->WITHOUT medians. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
a47355780b
commit
f5bbc26c60
@@ -19,16 +19,17 @@ Use codegraph for **structural** questions — what calls what, what would break
|
||||
| "What would break if I changed Z?" | `codegraph_impact` |
|
||||
| "Show me Y's signature / source / docstring" | `codegraph_node` |
|
||||
| "Give me focused context for a task/area" | `codegraph_context` |
|
||||
| "Survey an unfamiliar module/topic" | `codegraph_explore` |
|
||||
| "See several related symbols' source at once" | `codegraph_explore` |
|
||||
| "What files exist under path/" | `codegraph_files` |
|
||||
| "Is the index healthy?" | `codegraph_status` |
|
||||
|
||||
### Rules of thumb
|
||||
|
||||
- **Answer directly — don't delegate exploration.** For "how does X work" / architecture / trace questions, answer with 2-3 codegraph calls: `codegraph_context` first, then ONE `codegraph_explore` for the source of the symbols it surfaces. Codegraph IS the pre-built index, so spawning a separate file-reading sub-task/agent — or running a grep + read loop — repeats work codegraph already did and costs more for the same answer.
|
||||
- **Trust codegraph results.** They come from a full AST parse. Do NOT re-verify them with grep — that's slower, less accurate, and wastes context.
|
||||
- **Don't grep first** when looking up a symbol by name. `codegraph_search` is faster and returns kind + location + signature in one call.
|
||||
- **Don't chain `codegraph_search` + `codegraph_node`** when you just want context — `codegraph_context` is one call.
|
||||
- **`codegraph_explore` is the heavy hitter** for unfamiliar areas — it returns full source from all relevant files in one call, but is token-heavy. If your harness supports parallel subagents (e.g., Claude Code's Task tool), spawn one for explore-class questions to keep main session context clean.
|
||||
- **Don't loop `codegraph_node` over many symbols** — one `codegraph_explore` call returns several symbols' source grouped in a single capped call, while each separate node/Read call re-reads the whole context and costs far more.
|
||||
- **Index lag**: the file watcher debounces ~500ms behind writes; don't re-query immediately after editing a file in the same turn.
|
||||
|
||||
### If `.codegraph/` doesn't exist
|
||||
|
||||
Reference in New Issue
Block a user