Commit Graph
866 Commits
Author SHA1 Message Date
Colby MchenryandGitHub 222f82b9a5 Merge pull request #1440 from colbymchenry/feat/copilot-installer-targets
feat(installer): GitHub Copilot targets — VS Code, Copilot CLI, JetBrains
2026-08-07 13:50:33 -05:00
Colby McHenry 493d4210f1 Merge branch 'main' into feat/copilot-installer-targets
# Conflicts:
#	CHANGELOG.md
2026-08-07 13:49:22 -05:00
Colby MchenryandGitHub c84ce55855 Merge pull request #1528 from colbymchenry/issue/671-mcp-supported-languages
feat(mcp): surface supported languages in MCP server instructions (recut of #678)
2026-08-07 13:40:20 -05:00
Colby McHenryandClaude Fable 5 2962e7e1f4 feat(mcp): surface supported languages in MCP server instructions (#671)
Recut of #678 against the current instructions — the original predated the
explore-first rewrite and conflicted in both files it touched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:40:15 -05:00
Colby MchenryandGitHub 1b36132a89 Merge pull request #1498 from colbymchenry/bugfix/CG-16
fix(telemetry-dashboard): accept Origin: null on login — no-referrer policy locked Chromium out (CG-16)
2026-08-07 13:38:07 -05:00
Colby MchenryandGitHub 99f2ebf0d1 Merge pull request #1527 from colbymchenry/bugfix/CG-38
CG-38: guarantee an agent-named symbol renders, wherever it sits
2026-08-07 13:19:25 -05:00
Colby MchenryandGitHub fa8a3d7226 Merge pull request #1526 from colbymchenry/feature/CG-35
CG-33/CG-35: converge incremental sync with a full rebuild
2026-08-07 13:19:21 -05:00
Colby McHenry 2c708caf7c Merge branch 'main' into feature/CG-35 2026-08-06 21:17:56 -05:00
Colby McHenry 89c53ddf24 fix(explore): guarantee an agent-named symbol renders, wherever it sits (CG-38)
`codegraph_explore` never returned `queueMessage` (L1087) or
`flushQueuedMessages` (L1102) from a 1,414-line file, on a symbol bag or a
prose question, even with that file at rank #1 holding 67% of the envelope —
the agent got a same-stem `QueuedMessage` interface at L70 and had to Read the
file for the functions it had named. Pre-existing at every build including
pre-epic (controlled bisect, index held fixed).

Two independent causes:

1. `buildFlowFromNamedSymbols` returns the Flow prose AND the set of node ids
   the agent named — and the latter is the whole guarantee, since it injects a
   named def into its file's cluster ranges at importance 9. Its bail-outs
   returned EMPTY, zeroing the identity whenever there was nothing to PRINT.
   Two sibling closures that never call each other produce no chain, no synth
   hop and no boundary, so both defs lost importance 9 and the file rendered
   from its head. `identityOnly()` now separates the two, gated on
   shape-precise tokens so a prose word that exact-matches a callable cannot
   promote itself.

2. The ceiling trim filled in SOURCE order, so an over-ceiling render always
   dropped the END of a large file first. The shrink HAD kept both symbols
   (1022-1121); the trim cut back to 839. `windowToCeiling` now takes the
   spine call site plus every importance>=9 member as focus lines, tries the
   full ceiling first, and splits the held-back reserve evenly with
   carry-forward — greedy-in-source-order reproduced the bug one level down.

The shrink's loose size estimate is left alone deliberately, and the comment
now says why: making it exact was built and measured WORSE (it stops at the
last member that fits whole and the released bytes carry forward to
lower-ranked files, costing payroll-go's `s.store.Upsert`). `bound()` clamps to
the ceiling anyway, so the slack costs no bytes; it just must not pick the
survivors, which is what the trim now handles.

The measurement gap this closes: every existing probe is aggregate — envelope
share, per-file spend, source totals, file counts — and all are green on a
response that returns 25K from the right file and omits the named function.
`probe-named-symbol.mjs` checks the definition LINE against the response's
rendered lines, per symbol.

Suite envelope byte-identical to main on all six repos; probe-allocation 4/4,
no starvation flags; 180 files / 2,997 tests green. Fixture: 7/7 fail on main,
7/7 pass here, deterministic over 4 runs per arm.
2026-08-06 21:11:51 -05:00
Colby MchenryandGitHub 969ea1ec37 Merge pull request #1525 from colbymchenry/feature/CG-24
CG-24: explore response noise — allocation fixes, generated-file detection, and index-drift convergence
2026-08-06 20:34:28 -05:00
Colby McHenryandClaude Opus 5 07338ff12e docs(benchmarks): record CG-38 as open, and correct the regression claim
The epic record said nothing was open. CG-38 is: agent-named symbols in the
tail of a large file never render, which the epic's probes cannot see because
none of them measures whether the named symbol appeared.

Also corrects a wrong claim made while investigating it. The epic was said to
have regressed its own motivating query; that comparison varied the index as
well as the engine. A controlled bisect holding the index fixed shows the
pre-epic engine rendering 12 lines and CG-36 rendering 463 — the epic strictly
improves the case, and the symbols render at neither.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 20:29:14 -05:00
Colby McHenryandClaude Opus 5 8a4623463d merge: shrink a later cluster into the remainder instead of dropping it (CG-36)
A file whose top-ranked cluster was trivial kept it, dropped the cluster
carrying the answer WHOLE, and left most of its reservation unspent — because
only the first-chosen cluster could be shrunk. CG-31's carry-forward then
correctly handed that slack down the rank order, so the budget was not merely
unspent but REDIRECTED to weaker files.

django's sql/query.py (score 83, reserved 7,947) went from 1,923 delivered
chars to 10,082, and its envelope share from 7.7% to 40.4%; contrib/admin/
filters.py (score 18) went from 8,057 at 355% of its reservation down to 2,198
at 97%. All 8 starvation flags across the suite clear. Net +1,012 source chars.

The issue named the wrong fix point and the measurement said so: both real
cases lost on maxImportance, NOT on the density tiebreak the issue and its
duplicate (CG-37) suspected. Cluster ranking was left untouched, so the
Session.swift case density-first exists for still works — now pinned by a
dense-header fixture.

Accepted cost: okhttp trades its rank-6 file (score 21, reserved 1,999) for
+7,196 chars in the two files that answer the question, taking it from 6
delivered files to 5. django -159, okhttp -219 and tokio -25 source chars
against the epic tip; gin +1,176, alamofire +187, excalidraw +52. django also
stops cutting its epilogue.

Ships probe-file-spend.mjs, a standing suite-wide probe for reservation vs
spend, so this stays measurable — the original evidence came from ad-hoc
instrumentation that no longer existed and had to be re-derived by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 15:17:41 -05:00
Colby McHenry 10f1ac601a docs(benchmarks): record the CG-36 cluster-starvation measurement
The issue blamed the density tiebreak; both real cases lost on maxImportance,
so ranking was left alone. Full before/after table, the one cost (okhttp's
rank-6 file, squeezed out by reservations that were already structurally
over-subscribed), and what ships to keep it measurable.
2026-08-06 15:13:35 -05:00
Colby McHenry eed16447c3 fix(explore): shrink a later cluster into the remainder instead of dropping it (CG-36)
A file's ranked clusters were all-or-nothing past the first one: the top-ranked
cluster was taken (shrunk to fit when it had to be) and every cluster below it
was rendered whole, then either fit the remainder or was dropped entirely. On a
file whose top-ranked cluster is TRIVIAL that discards the answer — django's
`db/models/sql/query.py` kept a 22-line glue cluster and dropped the 624-line
`Query` body, spending 1,923 of a 7,947 reservation; okhttp's
`RealInterceptorChain.kt` did the same behind its import header.

The response stayed full, which is why this was invisible: the unspent
reservation carried forward exactly as designed and a file scoring a fifth as
much took the bytes.

Two sites, the same rule — hold the remainder while it is still worth a section
(CG-26's between-FILES lesson, applied between CLUSTERS):

- selection now shrinks a later cluster into what is left of the file's budget,
  by the same whole-member rule the first cluster already used;
- the ceiling trim re-renders the weakest cluster into the room that remains
  before dropping it. On excalidraw's `typeChecks.ts` the section-cost estimate
  missed by 13 chars and a 1,512-char cluster — the file's highest-SCORING one —
  was thrown away to pay for it.

Cluster RANKING is untouched: measured, both real cases lost on `maxImportance`,
not on the density tiebreak the issue suspected, and density-first is what keeps
Alamofire's `Session.swift` from burying its methods under the property list.

Suite (6 repos, clean-rebuilt indexes): all 8 starvation flags cleared,
+1,012 source chars net. django's `sql/query.py` 1,923 -> 10,082 of 7,947,
okhttp's `RealInterceptorChain.kt` 1,474 -> 6,038 of 6,058, gin's
`routergroup.go` 3,273 -> 5,632. okhttp trades its rank-6 file (score 21) for
+7,196 chars in the two files that answer the question.

Ships two fixtures pulling in opposite directions (`starved-cluster-ts` and
`dense-header-ts`), a `spendShareAtLeast` gate in probe-allocation, and
probe-file-spend.mjs — a standing per-file reservation-vs-delivered sweep.
2026-08-06 15:10:46 -05:00
Colby McHenryandClaude Opus 5 76ab1fe130 docs(benchmarks): record the CG-24 epic resolution
Four shipped fixes, one open defect (CG-36), and five issues closed because
measurement contradicted them. The headline is that the reported symptom was
not an explore bug at all — it was a degraded index (CG-33), and the reported
query answers correctly on a clean rebuild with no explore change.

Records the two traps that cost real time and are now guarded in tooling: the
nonexistent .codegraph/graph.db path that sqlite3 silently creates, and
ab-new-vs-baseline.sh swapping src/ mid-run so a commit captures baseline
sources.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 14:46:11 -05:00
Colby McHenryandClaude Opus 5 ed6c56f3fa merge: damp undepended-on ambient declaration files on flow queries (CG-28)
Both halves of the issue were measured on a hermetic fixture of four
declaration-shaped files varying on banner and depended-on-ness.

CG-25 already handles the motivating file: the Wrangler worker-configuration.d.ts
that opened this issue is demoted by the generated penalty alone, worth 15-46
points of envelope share across four flow queries. No new mechanism for it.

The narrower gap is real. A declaration file with NO banner carried pen 1.00,
took rank #1 and 51% of delivered source on a prose flow query, and displaced
the flow's own entry file out of the response entirely.

The rule is deliberately narrow, and both conditions were derived by survey
rather than guessed. 'Declares no callable and calls nothing' flags 1.1-18.0%
of files across the corpus and catches real source — okhttp's SocketPolicy.kt,
tokio/src/runtime/mod.rs, Alamofire's umbrella file, django's locale format
tables. Requiring every symbol to be type-level drops that to 0-4%. The
'nothing depends on it' condition was added after the broader version demoted a
pure-interface file with 13 inbound imports and broke the CG-31 displacement
gate — a different invariant entirely.

Does NOT stack with the generated penalty: rankPenalty takes Math.min of the
two, so a file that is both takes the stronger, never the product. A query that
NAMES a declaration symbol exempts its file entirely, so asking about a type
still reaches it at full weight.

Six-repo envelope is byte-identical to the pre-change tip — the rule does not
fire on any benchmark repo, consistent with the 0-4% survey.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 14:40:25 -05:00
Colby McHenryandClaude Opus 5 9efae0f8f2 fix(explore): damp ambient declaration files on flow queries (CG-28)
A file that declares nothing but types and that nothing in the index depends
on — a hand-written ambient `.d.ts` of global shims, vendored typings, module
augmentation — cannot answer a flow question: no bodies, no call edges, no
behaviour, nothing typed by it. But the identifiers it declares are exactly the
generic ones a prose question uses (`Body`, `Message`, `ImageMetadata`,
`ReadableStream`), so on term overlap it out-scored the implementation. Measured
on the new fixture: rank #1 and 51% of delivered source, with the flow's own
entry file pushed out of the response entirely.

Measured first, per the issue: the Wrangler `worker-configuration.d.ts` that
opened this is already handled by CG-25's banner detection, worth 15-46 points
of envelope share across four flow queries. CG-25 credited; only the un-bannered
case needed anything.

`rankPenalty` now multiplies score and graph mass by 0.5 for such files, taken
as the STRONGER of it and the generated penalty rather than multiplied — one
property two signals see must not be charged twice. Detection is structural, not
by extension, and four conditions deep. Two of them were forced by measurement:
requiring every symbol to be type-level takes the corpus flag rate from 1-18%
(which swept in Kotlin sealed classes, Rust mod.rs re-exports and django's
locale tables) down to 0-4%; requiring that nothing depends on the file
separates an ambient shim from a working types module, and without it the rule
demoted displacement-ts's pipeline `types.ts` and broke the CG-31 gate.

A query that NAMES a declared type is exempt, so a question about a type still
reaches its declaration at full weight. Precise tokens only, so "…the file
body…" cannot exempt a `Body` interface it never meant to name; this needs its
own set because `namedSeedIds` is callable-only and a type never becomes one.

Regression evidence in docs/benchmarks/explore-declaration-only-cg28.md:
6-repo envelope sweep byte-identical against a clean baseline build, zero
ambient files reach the candidate set on VS Code across five queries, corpus
flag rate 0-0.74%, both allocation fixtures PASS, full suite 2,978 green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 14:35:56 -05:00
Colby McHenryandClaude Opus 5 463f6e7844 merge: factory-closure envelope premise measured and rejected (CG-27)
CG-27 proposed adding function/method to ENVELOPE_KINDS so a factory closure
spanning most of its file stops merging every inner symbol into one cluster.
The issue required the ranking claim be measured before any fix. It was, on a
hermetic fixture built to make the pattern maximally visible, and it does not
hold — nothing shipped to src/.

The literal change is a large regression: dropping the enclosing range SPLITS
the file into a trivial cluster (a type alias plus a helper, span 7) and the
answer-bearing one (every closure, span 359). Cluster ranking breaks the equal
maxImportance tie on density, so the trivial cluster wins, is taken first, and
is the only one that may be shrunk; the answer-bearing cluster then does not fit
and is dropped whole. Rank #1 fell from 7,539 delivered chars to 397, and from
7 of 11 inner closures to 0. The enclosing range was holding the file together
as one cluster, inside which shrinkCluster already did the per-symbol ranking
the issue asked for.

A better mechanism reaching the same intent — deferring the envelope member
inside shrinkCluster, leaving clustering untouched — is noise: 69 vs 68 inner
definitions across nine query shapes, one better, one worse, seven unchanged.

The one configuration where the envelope IS selected (the factory as sole
top-tier member) is already absorbed by CG-30, which windows it on whole lines:
a contiguous readable head carrying 6 of 9 closures, bounded and never empty.

Kept: the fixture, the deterministic probe, and the measurement record.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 14:11:01 -05:00
Colby McHenry 91cb5b4317 measure(explore): the factory-closure envelope premise does not hold (CG-27)
CG-27 asked whether the >50%-of-file envelope drop should cover `function` /
`method`, so a `createFoo()` factory returning an object of closures stops
merging every closure inside it into one cluster. Measured on a hermetic
fixture, it should not, and the issue is closed as obsolete with CG-30 credited.

Two mechanisms already absorb the shape. shrinkCluster orders members by
(importance desc, size ASC) and refuses any member that overruns the cap once
something is kept, so a file-spanning member is only selected when it is the
sole member of the top importance tier — eight of nine query shapes never
selected it at all. When it IS selected, CG-30 windows it on whole lines, so
the file still delivers bounded, readable source (6 of 9 closure definitions
in that configuration).

Dropping the range instead SPLITS the file, and only the first-chosen cluster
may be shrunk: a trivial 7-line cluster won the density tiebreak and the
answer-bearing cluster was dropped whole — rank-#1 file 7,539 chars and 7 of 11
closures to 397 and none. Reaching the same intent more carefully (defer the
envelope MEMBER inside shrinkCluster, leaving clustering untouched) is noise:
69 vs 68 closure definitions across nine query shapes. Nothing shipped.

Adds the fixture, the probe, a standing gate on the outcome, and the record —
including a real defect the measurement exposed on the epic tip: django's
query.py leaves 8,212 of 10,135 unspent and drops a score-290 cluster to keep a
score-14 one. Filed separately.

No behaviour change, so no CHANGELOG entry.
2026-08-06 14:07:22 -05:00
Colby McHenry d49265043c test(explore): add the factory-closure fixture and its selection probe (CG-27)
A file whose top-level symbol spans almost all of it — createFoo() returning
an object of closures — is how Svelte 5 rune stores, React custom-hook modules,
IIFE module-pattern JS and Zustand's create((set,get)=>({…})) are all written.
probe-factory-closure.mjs measures what such a file DELIVERS from within: which
inner symbols' definitions reach the agent, not how many bytes did.
2026-08-06 13:58:36 -05:00
Colby McHenryandClaude Opus 5 dc4fd755ef merge: recognize Wrangler-style generated banners (CG-25)
A generated Cloudflare Wrangler ambient-types file was not flagged generated, so
it ranked with no penalty and competed with hand-written source on generic token
overlap. The banner shape it uses — "Generated by <tool> by running <command>" —
matched none of the existing content patterns, all of which require DO NOT EDIT,
a standalone @generated, or the "auto(matically) generated by" phrasings.

Precision is held by requiring TWO 'by' clauses: the banner must name a tool and
then say 'by running'. Ordinary prose ("the report is generated by running the
nightly job") has only one and does not match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:17:15 -05:00
Colby McHenryandClaude Opus 5 8bb0f53bca merge: explore allocation — bounded overshoot, displacement guard, exact budget (CG-30, CG-31, CG-26)
Lands the three-branch allocation stack. Every admitted file now receives at
least its reservation before any file draws on carry-forward slack, on every
render path — cluster, whole-file GRACE, and whole-file BUY.

CG-30 bounded how far an oversize cluster member may overshoot (windowed on
whole lines past 1.5x rather than emitted whole or dropped). CG-31 gave the
cluster path the `owedBelow` displacement guard the BUY arm always had, holding
back only the prefix of what is owed below that the response can actually pay.
CG-26 closed the three remaining holes: the whole-file arms had no displacement
guard at all, section overhead was charged at a flat 200 against a real 300-500,
and `owedPayableBelow` held all-or-nothing where it should hold partially.

Deterministic across the 6-repo suite, clean-rebuilt indexes, both builds: no
repo truncates, no repo loses a file, okhttp gains one, and every repo lands at
or under the 25,000 hard ceiling.

Accepted trade (maintainer decision): excalidraw -552 and okhttp -164 source
chars against the CG-31 tip, in exchange for the trailing pointer list surviving
instead of being discarded whole. Those bytes existed at the CG-31 tip only
because it over-filled a ceiling it mis-measured and then dropped the entire
epilogue; a pointer the agent can act on beats a few hundred chars on the
last-ranked file.

Two issues opened during this work were closed as invalid rather than fixed:
CG-32 (named-file ordering) and CG-34 (allocator over-reservation). Both were
filed on diagnoses that did not survive measurement — CG-32's symptom was index
drift (CG-33), and CG-34's premise was overturned by CG-31's own results.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 11:17:15 -05:00
Colby McHenryandClaude Opus 5 57e0854213 fix(explore): recognize Wrangler-style "generated by … by running" banners (CG-25)
Cloudflare Wrangler's `worker-configuration.d.ts` (~12k lines of ambient
types) carried no banner any GENERATED_CONTENT_PATTERNS entry matched:
every existing marker requires `DO NOT EDIT`, a standalone `@generated`,
`<auto-generated>`, or the literal `automatically/auto-generated by`
phrasings. Wrangler emits a bare `Generated by Wrangler by running
`wrangler types``, so the file ranked with pen 1.00 and won 79.4% of an
explore envelope on generic token overlap alone (CG-24).

The discriminator is the reproduction instruction, not the word
"generated": the banner must name a tool AND then say `by running`, i.e.
two separate "by" clauses. That keeps prose out — "the nightly summary is
generated by running the ETL job" has only one — while catching every
CLI-driven emitter that tells you how to regenerate.

Precision swept over 441,856 files across the whole local source tree: 5
hits, all genuine Wrangler output, no false positives.

Isolated before/after on the CG-24 repro (same query, same index, only
the `files.generated` flag differing):

  before  pen 1.00  score 115.0  share 79.4%  3 files rendered
  after   pen 0.30  score  35.4  share 21.1%  4 files rendered

The new pattern stays in the existing table position, below the header
window the detector scans, so the module still does not classify itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:54:05 -05:00
Colby McHenryandClaude Opus 5 02ee151e46 CG-35: give the sync-convergence suite teeth against the rebind pass
The suite passed unchanged with `CODEGRAPH_NO_REBIND=1`, so the larger half
of CG-33 — the rebind pass — had no coverage at all.

The cause was the ground truth, not the cases: `rebuildEdgeSet` called
`indexAll()` on the live handle. That is not a rebuild. Every file hashes
identical, so the store writes nothing (`nodesCreated: 0`), no reference is
re-created, and every edge survives — the comparison read the synced index
against itself and could never fail. It now goes through `CodeGraph.recreate`,
which deletes the database file the way the CLI's `index` command does.

With a real rebuild, three existing cases fail under the kill switch. Adds two
more for the rules that carry the risk:

- an edge with no `refName` stamp (older engine) and a synthesized
  (`provenance='heuristic'`) edge are never deleted — both planted directly,
  and each verified load-bearing by mutation;
- a name over the 500-edge ceiling is declined losslessly rather than
  rebound in part, with a rare name in the same sync as the control that
  proves the pass ran.

The per-file-vs-batch-wide delta rule is likewise confirmed by mutation: a
batch-wide name set fails its case.

CODEGRAPH_NO_REBIND=1 now fails 4 cases; unset is green; full suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:45:37 -05:00
Colby McHenryandClaude Opus 5 03893b0ab9 CG-33: converge incremental sync with a full rebuild
A live, auto-synced index did not converge to a clean rebuild of the same
tree — 4.3% of distinct edges wrong in both directions on this repo's own
index, overwhelmingly `calls`, which is what flow queries traverse and what
explore's file ranking weights. Silent: nothing warned, and the symptom read
as "codegraph isn't very good" rather than "this index needs rebuilding."

Two causes, and the fix needed both. Resolution binds a reference to one of
the same-named definitions PROJECT-WIDE, so a definition appearing or
vanishing changes the correct answer for references in files the sync never
touches — and those references resolved successfully once, which deletes
their unresolved_refs row, leaving nothing to revisit them with (#1240's
retry only revisits refs parked as failed). Separately, when nothing
disambiguated the candidates the winner came down to rowid, i.e. the order
files happened to be WRITTEN, which differs between a scan-order full index
and a sync that appends each file as it changes. That second one is why
re-resolution alone could not converge: re-resolving against the identical
graph still picked a different candidate.

So getNodesByName now orders by (file_path, start_line) — a property of the
code, not of the write order — and sync computes a definitionDelta and
re-opens the resolution edges whose answer it may have invalidated,
re-inserting each as the reference that created it for the orphan sweep to
bind against the post-sync graph.

The delta compares `file\0name` pairs per file rather than one name set over
the batch: a commit that adds `collect` to a new file while an unrelated
changed file already defines `collect` cancels out of a batch-wide set, and
that miss was the largest residual class in the first measurement.

Conservative where the failure modes are asymmetric — a wrong deletion is a
permanent edge loss, a missed rebind is only residual drift. Edges without a
refName stamp are never touched (nothing to restore them from), sources the
sync already re-extracted are skipped, and a per-name ceiling declines the
generic names. Edges are deleted before the sweep re-inserts, since
INSERT OR IGNORE against idx_edges_identity would otherwise keep both rows
when a reference rebinds elsewhere.

Replaying real commits of this repo through sync, then diffing against a
rebuild: 16 commits 48 -> 0; 80 commits 1,634 -> 361, with the actively
misleading direction (stale edges the index keeps asserting) 671 -> 2.
Index and sync wall-clock are unchanged; the ORDER BY costs 18% per uncached
name lookup, which never reaches wall-clock because the resolver memoizes it.

The 357-edge residual at 80 commits is one pre-existing class: refs to
generic names (`push`, `join`) parked above #1240's per-name retry ceiling,
which a rebuild resolves into cross-language garbage — a TS test file
"calling" an R method. Converging there would mean manufacturing wrong edges,
so it is left alone. And no drift metric in `codegraph status`: it cannot be
computed without the rebuild it would be recommending, and a proxy would fire
on that residual and train users to ignore it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 04:32:40 -05:00
Colby McHenryandClaude Opus 5 5f32478b57 docs(benchmarks): record the CG-26 A/B — the invariant holds on every path
Deterministic 6-repo table, the three agent A/Bs (django, excalidraw, okhttp,
2 runs/arm, Read 0 in all 12 runs), and an honest read of the two repos that
deliver a few hundred fewer source chars: at the CG-31 tip both were over-filled
by the flat-200 section overhead and paid for it by discarding their epilogue
whole.

Also: CHANGELOG entries for the two user-visible changes, and the memory note
now carries the fourth accounting gap plus the two lessons — hold the REMAINDER
when a full reservation no longer fits, and never skip a file over an accounting
difference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:58:18 -05:00
Colby McHenryandClaude Opus 5 7cbde95ce2 fix(explore): pay every admitted file on every render path (CG-26)
The invariant this closes: every admitted file receives at least its
reservation before any file draws on carry-forward slack. CG-30 bounded an
oversize cluster member and CG-31 gave the cluster path a displacement guard;
three holes were left, and each one starved a file that had been admitted,
reserved and — in the worst case — rendered.

1. The whole-file arms had no displacement guard. BUY's fit test read
   `renderCeiling - totalChars` (everyone's room) while its source-space
   sibling refused the same trade, and GRACE was not fit-tested at all.
   okhttp's CallServerInterceptor.kt shipped 8,499 chars on a 5,964 funded
   ceiling and the rank-6 file below it delivered nothing. Both arms now test
   the render they actually produce against `fundedHeadroom`, and a whole
   render that does not fit falls through to clustering instead of skipping
   the file.

2. Every section was charged a flat 200 chars while a real header runs
   300-500. The loop believed it had room it did not have — okhttp allocated
   26,601 against a 24,400 ceiling — so the final truncation threw a
   fully-rendered section away. Sections are charged their real cost now, the
   owed-below arithmetic uses a per-file overhead estimated from the file's own
   symbols, and a marginal overrun trims the weakest cluster (or windows the
   last one into the room that is left) rather than skipping the file over a
   rounding difference.

3. `owedPayableBelow` held all-or-nothing. When the last admitted file's FULL
   reservation no longer fit, nothing was held for it: on the precise-query
   fixture the rank-5 file took 4,134 chars against a 2,948 reservation while
   rank 6 — admitted, reserved 2,539 — was left 4 chars and skipped. It now
   holds the remainder while that remainder is still worth a section
   (MIN_CHARS).

And the epilogue is budgeted instead of discarded. The flat 600-char margin was
neither the epilogue's size (1,064 gin, 1,788 django, 2,231 excalidraw) nor a
bound on it, so four of six suite repos shipped with no pointer list and no
reminders at all. The loop now reserves the epilogue's FLOOR — the one line
that says an uncovered area exists, plus a pointer for every file whose bytes
were deliberately withheld (CG-12) — and the rest is fitted to the room that
actually remains, in priority order, entry by entry. Sized from the real
strings; no constant was swept against the suite.

Deterministic, same clean-rebuilt indexes, baseline = CG-31 tip:

  repo         base source   new source   files      ceiling
  django            20,791       20,878   6 -> 6     was discarding its epilogue
  tokio             21,521       21,607   5 -> 5     was discarding its epilogue
  okhttp            19,034       18,870   5 -> 6     +1 file delivered
  excalidraw        20,204       19,652   8 -> 8     keeps its pointer list
  gin               10,776       10,776   4 -> 4     byte-identical
  alamofire         11,662       11,662   2 -> 2     byte-identical

No repo truncates any more and none loses a file. okhttp and excalidraw trade
164 and 552 source chars on their LAST-ranked file for the pointer list naming
what the response could not cover — bytes the CG-31 tip only had because it
over-filled a ceiling it mis-measured and then discarded the epilogue whole.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:45:28 -05:00
Colby McHenryandClaude Opus 5 c54e0080c2 fix(explore): keep the drift warning out of the cuttable epilogue (CG-31)
The '⚠ changed on disk after the last index sync' banner is an honesty claim
about source we DID render — line refs elsewhere in the response may be
shifted — not a note about the response. Drawing the epilogue boundary after
it means the size cut can never be what silences it. Suite numbers unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:15:06 -05:00
Colby McHenryandClaude Opus 5 be7c968439 docs(benchmarks): record the CG-31 A/B — no regression, four repos stop truncating
Deterministic (6 repos, clean rebuilds, both builds): four deliver more source
and one more file each, two are byte-identical, none deliver less. Agent A/B
(django n=3, okhttp n=2, gin n=2, sonnet/effort high, both arms codegraph-on,
0 contamination): the new arm is faster on all three, Read at or below
baseline, occupancy lower.

Also records the two corrections the suite forced on the first cut of the
guard, and the two residuals CG-26 inherits — the render loop's 600-char
epilogue margin (a sweep was run and deliberately NOT shipped) and the BUY
arm's source-space-only guard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:12:30 -05:00
Colby McHenryandClaude Opus 5 f1fecb8232 fix(explore): fund the guard from room that exists, and cut the epilogue first (CG-31)
Two corrections found by measuring the first cut of the guard against the
6-repo suite. The first version held back the FULL sum of the reservations
below a file. On django that took 2,319 chars off a file the agent receives
and handed them to a section the hard ceiling then threw away — the guard's
own failure mode, one layer down. tokio lost 1,298 the same way.

1. `owedPayableBelow` — hold back only the prefix of what is owed below that
   the response can still PAY, in rank order. A promise the ceiling cannot
   reach is not a claim on this file's bytes.

2. The final truncation now spends the EPILOGUE before it spends a rendered
   file section. It used to cut at the last section header, dropping that
   section AND the trailing notes; dropping the notes alone is almost always
   enough. A section is source the agent otherwise has to Read; the epilogue
   is a pointer list and two reminders, and the note that replaces it carries
   the "explore these names" instruction forward.

Also count `flow.text` in `totalChars`. It is prepended to `lines` to make the
final output, so the render loop always spent against a ceiling it was ~2K
under on symbol-bag queries.

Deterministic, same clean-rebuilt indexes, both builds (baseline = CG-30 tip):

  repo         base source   new source      files
  django            20,033       20,791   5 trunc -> 6
  excalidraw        18,776       20,204   7 trunc -> 8
  okhttp            15,628       19,034   4 trunc -> 5
  tokio             20,340       21,521   4 trunc -> 5
  gin               10,776       10,776   4 -> 4   (byte-identical)
  alamofire         11,662       11,662   2 -> 2   (byte-identical)

No repo delivers less; four stop truncating. `funded` in the diagnostic now
reports the render CEILING the guard allows, which is what every render path
is actually bounded by.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 02:59:46 -05:00
Colby McHenryandClaude Opus 5 089dcc276f fix(explore): hold back what is still owed below a clustered render (CG-31)
Carry-forward slack let a file spend what the files ABOVE it left on the
table. Nothing held back what was promised BELOW it. The whole-file BUY arm
has always refused that trade (`owedBelow`); the cluster path read `headroom`
— what is left before the hard ceiling — instead of what is still owed, so
`fileBudget` and `SPINE_CEILING` could pay a 1.5x overshoot out of another
file's reservation.

`fundedHeadroom` is the same inequality in the units the cluster path spends
in: source PLUS the per-section overhead each unreached file will charge.
Floored at the file's own reservation — a kept promise is not a displacement —
and it is <= `headroom` by construction, so it is the only bound the three
render sites need. The skeleton path's `bodyCap` takes it too.

Measured on `__tests__/fixtures/displacement-ts` (a 4-stage pipeline padded
past 500 files, where the 24K envelope genuinely saturates the 24.4K render
ceiling):

  before  ingest.ts emitted 9,301 on a 6,289 spendable, then lost the whole
          section to the final ceiling — 0 delivered. types.ts and sink.ts
          skipped `budget-whole-file`. 3 of 6 admitted files delivered.
  after   ingest.ts bounded to the 4,913 actually free. 6 of 6 delivered,
          envelope 14,908 -> 22,066.

The self-query allocation fixture flips back to PASS with it, on a clean full
rebuild of this repo's index (CG-33). Its `afterCG30` verdict blamed an
over-RESERVED incidental file; the reservation was identical in both arms —
the file was over-SPENDING. Recorded honestly in `afterCG31`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 02:44:59 -05:00
Colby McHenryandClaude Opus 5 0d014a6582 docs(benchmarks): record the CG-30 A/B — deterministic win, no behavioural regression
Primary evidence is deterministic: on django, query.py rendered 2.12x its budget
on main and 1.49x with the bound, and the freed bytes reach the files below it
(+2,104 chars of source in the same five files). gin is a true control — the two
builds emit byte-identical explore output there, which is what makes its agent-run
deltas variance by construction.

Also records the harness trap that voided the first two batches: ab-new-vs-baseline
swaps src/ to the baseline ref mid-run, so a commit made while it runs captures
baseline sources. Check the "changed:" line before believing any run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 02:26:22 -05:00
Colby McHenryandClaude Opus 5 cd1ea27ea8 fix(explore): restore the CG-30 bound d652c14 reverted
d652c14 was committed while an ab-new-vs-baseline run had the engine checked out
at the BASELINE ref — that harness swaps src/ files mid-run and restores them on
exit — so it captured main's tools.ts and explore-diagnostics.ts and silently
undid 765c06a. Restored from 765c06a, with d652c14's doc tweak re-applied.

The A/B runs started after that commit are void with it (their "changed:" line
lists only explore-diagnostics.ts, i.e. both arms ran the same retrieval code)
and are re-run rather than reported.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 02:08:28 -05:00
Colby McHenryandClaude Opus 5 d652c148f6 docs(cg-30): changelog entry + record the self-query probe flip honestly
The self-query allocation probe fixture's delivered-share gates now fail. The
cause is not the new bound: allocation is unchanged between arms (parse-run.mjs
32.3% vs 33.9% on main) and tools.ts delivers the same 8,282 chars in both. What
changed is that the incidental file now DELIVERS — on main its whole section was
cut by the hard-ceiling truncation, so the fixture passed on truncation luck.
Every file on this repo obeys the new bound (max 1.40x of spendable).

Recorded as `afterCG30` with that reasoning rather than tuning the bound to
restore the pass. The over-reservation it exposes is epic CG-24's subject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 01:47:02 -05:00
Colby McHenryandClaude Opus 5 765c06aa40 fix(explore): bound how far an oversize cluster member may overshoot (CG-30)
shrinkCluster keeps an oversize cluster's highest-importance member WHOLE on
purpose — an empty file section sends the agent to Read, the outcome explore
exists to prevent. What it lacked was a bound, and "never empty" quietly meant
"never bounded": on the reporting repo one file emitted 22,376 chars against a
9,181-char reservation (2.44x), past both the per-file budget and the spine
ceiling. That overshoot is what collapses `headroom` for every file below it.

The same rule has a second face. When the top member is bigger than the whole
response ceiling, the file does not overshoot — it is dropped entirely at the
renderCeiling check, so the agent gets nothing for a file it named.

renderCluster now takes a ceiling (1.5x what the file may spend — the same
multiple SPINE_CEILING already draws, and never below the cap, so a cluster
that fits is untouched). Past it the member is WINDOWED on whole lines rather
than emitted whole or dropped: leading window plus, on a flow cluster, a window
on the spine's call site. A partial window shorter than 12 lines is dropped
instead — a sliver in the session record forces the next call's dedup to shred
the block around it or re-send it — unless nothing else was emitted, where the
never-empty floor wins.

Measured on the new fixture, pre-fix vs post-fix:
  monthly.ts    12,391 chars on a 3,334 budget (3.7x)  →  4,941 (1.48x)
  quarterly.ts  dropped, no headroom left               →  4,004 delivered

Also: the diagnostic now reports `spendable` (reservation + inherited slack)
alongside `reserved`. Every render bound reads the former, so reporting only
the latter makes an ordinary carry-forward read as a file spending over budget
— and it made the overshoot this issue is about unmeasurable. A windowed file
is now flagged `clipped` too, instead of presenting a window as the whole file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 01:44:45 -05:00
Colby McHenryandClaude Opus 5 2cf63fd114 CG-33: record index-drift measurement and add a drift diff tool
A live, auto-sync-maintained index does not converge to a clean full
rebuild of the identical tree. On codegraph's own repo, 4.3% of distinct
edges are wrong in both directions (751 missing, 476 stale), dominated by
`calls` — the edges flow queries traverse and that feed the RWR mass
explore ranks files by.

Raw edge rows differ by only +0.7%, because the divergence is
bidirectional and nets out; any drift check must compare edge SETS.
Rebuild-vs-rebuild is 0, so the indexer is deterministic and this is not
noise. Node sets are identical and every integrity check is 0 on both
indexes, so this is stale cross-file resolution, not accumulated residue.

`diff-index-drift.mjs` is read-only and takes two index paths — rebuilding
is the caller's job, so the tool can never clobber the artifact it is
measuring. It also refuses a missing path, since node:sqlite creates an
empty database rather than failing and an empty schema reads exactly like
a stale pre-migration index.

Diagnostic captures from the originating incident are deliberately NOT
committed: they contain verbatim source from a private repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 01:24:59 -05:00
Colby McHenryandClaude Opus 5 d6d17288be docs(benchmarks): re-derive the token figures the result.usage bug touched
Swept every benchmark doc for figures produced off `result.usage` and fixed
the ones that had raw logs to re-derive from.

residual-context-occupancy.md — the sonnet 3-turn throughput table. Re-derived
from the preserved logs: tokens saved 23% -> 56%, and vscode's "98% MORE tokens
with codegraph" was never real, it is 41% fewer. Cost, time and tool calls were
never affected by this field and are unchanged. The occupancy table itself is
measured off the timeline, so every number in it stands -- including the 82%
higher residual, which is the finding the document exists for.

call-sequence-analysis.md — this doc DIAGNOSED the bug and its reproduce block
claimed the aggregator summed per-turn tokens. It did not, until 04c0f8e. Noted,
with the three wrong results the gap produced: the excalidraw cut recorded here,
the sonnet campaign, and the Opus re-measure that invented a token regression.

answer-directly-vs-explore-agent.md — build 0.9.4, 2026-05-24, raw logs gone.
Cannot be re-derived, so flagged rather than silently left or invented: its
token figure is indicative, its turn/read/context findings do not depend on the
broken field and stand.

The remaining benchmark docs (allocation-ab-1500, dedup-cg20, allocation-
efficiency, feedback-metrics) carry no throughput tables — checked, clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 21:36:10 -05:00
Colby McHenryandClaude Opus 5 da1f6121fd docs(readme): benchmark table from the corrected 2026-08-05 re-measure
Re-measured on the CLI-blocked harness at the README's own stated methodology
(claude-opus-4-8, single question, median of 4, same 7 repos), with tokens
summed per turn per 04c0f8e.

What moved, and what did not:

  tool calls  89% -> 88% fewer   holds
  file reads  0 on all seven     holds
  tokens      69% -> 62% fewer   holds
  time        20% -> 53% faster  was badly understated
  cost        60% -> 44% cheaper genuinely lower

Time was the most wrong figure in the old table, in our favour: every repo is
faster with the graph, including the two that previously carried a footnote
excusing them for being slower. That footnote is gone. Cost is the one real
downgrade, and it varies with how much discovery a question demands rather
than with repo size -- 57-78% where the file-reading arm needed 28-43 tool
calls, near-even on Gin at 7. Said plainly instead of averaged away.

Methodology now records that the CLI is blocked in BOTH arms and why: on an
unblocked harness the control reached codegraph through Bash in 26 of 28 runs,
so it was not a control. All 28 attempted it here; all 28 were blocked.

The context-footprint note from a7db24d is unaffected -- occupancy is measured
off the timeline, not the token field that was wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 21:28:25 -05:00
Colby McHenryandClaude Opus 5 04c0f8eab5 test(agent-eval): sum tokens per turn — result.usage stopped being cumulative
"Tokens processed" was read off result.usage. That was correct when the README
figures were measured; in current Claude Code the field reports the LAST turn
only. Nothing here changed — the host did, silently — and the harness kept
reporting the smaller number.

The error is one-sided, which makes it worse than noise: it under-counts
whichever arm takes more turns, and that is always the WITHOUT arm. On the
2026-08-05 campaign it turned a real 62% token saving into 19% and invented a
token REGRESSION on tokio (-41%) and alamofire (-25%) that does not exist. Those
numbers were one push away from the README.

Now summed per assistant request and deduped by message.id, the same rule the
occupancy timeline already used — Claude Code emits one event per content block
carrying identical usage, so summing per event double-counts (~1.7x measured).

CLAUDE.md already warned about this field. The code did not follow; it does now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 21:27:02 -05:00
Colby McHenryandClaude Opus 5 a7db24d0c0 docs(readme): disclose the context-footprint side of the benchmark
The benchmark table measures throughput -- tokens processed, tools called,
cost to reach one answer. It has never measured what is still resident in the
window afterward, and on that axis codegraph costs more: ~80% more retrieval
context left behind than a file-reading agent, on all seven repos.

That is the axis issue #1500 reported, and it is structural rather than a
defect -- one dense payload that answers the question and stays, versus
grep-and-read churn that evicts. Worth stating plainly next to the cost note
rather than leaving a user to discover it in a long session.

The Opus 4.8 single-question figures in the table are deliberately untouched:
the occupancy campaign ran sonnet / 3-turn, a different regime, and nothing
measured there licenses restating them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 14:30:34 -05:00
Colby McHenryandClaude Opus 5 48cbc21a17 merge: cross-call explore session dedup (CG-2, #1500)
Never re-serve source this session already sent: session-scoped state (CG-17)
plus a precise back-reference in place of the bytes (CG-18). Duplicated source
drops 6.4% -> 0.74% of the response, at flat cost per call, with more unique
source in its place.

Bar 4 of CG-20 -- "residual occupancy must drop" -- is NOT met, and the gate
records why: CG-18's accepted rule spends reclaimed bytes on files not yet
shown rather than banking them, so a design that re-spends every byte cannot
lower the byte count. The two requirements were mutually unsatisfiable as
written. Bars 1-3 (no extra Reads, no abandonment, no bucket shift) pass across
24 runs on both arms.

Kept on that basis, and cheap to reverse: CODEGRAPH_EXPLORE_DEDUP=0 disables it
at runtime. Full gate: docs/benchmarks/explore-dedup-ab-cg20.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 14:29:20 -05:00
Colby McHenryandClaude Opus 5 2cfd321b23 docs: CG-20 — the dedup gate, three bars pass and the fourth cannot (CG-2)
Read is 0 in all 24 runs of both arms across client-go and excalidraw, no
isError, codegraph is the last tool in every run, and both "Read a file we
returned" / "did not return" buckets are empty — with back-references
demonstrably reaching the agent in 8 of the 9 multi-call runs.

Residual occupancy is flat, and the measurement shows it could not have been
anything else: CG-18 was accepted on the rule that reclaimed bytes get spent on
unseen files rather than banked, so the byte count cannot fall. What moves is
the duplicate fraction of that residual — 87% less across the agent runs, 86%
and 94% on two matched deterministic replays.

Also recorded: dedup.savedChars is a pre-clip figure (11,450 reported against
1,042 chars actually re-served), so tuning off it inflates the win ~7x.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 14:11:33 -05:00
Colby McHenry 7a7ea30cbd docs: cross-call dedup — its gates, where the bytes go, and the all-pointer guard (CG-18)
Extends the session-state design doc with the layer built on it: what gates a
withheld span (session, content fingerprint, two size floors, kill switch),
why the fingerprint and not #1474's drift flag, the two channels the
reclaimed bytes leave by, and why "no duplicate ranges across calls" holds
for every call that had something new to say rather than universally.
2026-08-05 13:47:29 -05:00
Colby McHenry ab38d1f090 feat(explore): point at source this session already sent, don't send it twice (CG-18)
A later explore call re-served whatever it re-ranked, so on the #1500 report
the 4th call spent its envelope on the spine the 1st call had already
delivered. CG-17 recorded what was served; this acts on it.

What a withheld span becomes is the whole design: a POINTER, never a silence.
An insufficient-feeling response is what sends an agent to Read, and one or
two of those early in a session teach it to abandon codegraph — so the
replacement names the file, the symbols and the line spans, and says both
that the source came from THIS conversation and that the file has not changed
since.

- Content fingerprint, not the drift flag, gates it. They answer different
  questions: two calls inside one drift window served the same current bytes,
  while a file edited AND re-synced between calls is never "stale" and yet
  the agent's copy is now wrong. An edited file re-emits in full.
- Only a covered run of >= 8 lines is replaced, and a remainder under 160
  chars folds into the pointer. Below those the pointer costs more than the
  source and the block reads as shredded — a fence holding `228\t` is a
  broken-looking response, which is the expensive failure.
- The reclaimed bytes go to files the agent has NOT seen, two ways: a smaller
  `sourceSpent` hands slack down CG-21's carry-forward pool, and a fully
  back-referenced file gives up its maxFiles slot the way a cliffed one does.
  Within a file, the cluster shrink now reads the DEDUPED length, so it never
  drops new symbols to make room for source it isn't sending.
- If dedup suppresses everything and nothing new takes its place, the top
  suppressed file is spliced back in whole. An all-pointer response is the
  shape that reads as "codegraph found nothing"; one re-served file is the
  cheaper mistake.

Kill switch: CODEGRAPH_EXPLORE_DEDUP=0.
2026-08-05 13:47:24 -05:00
Colby McHenry 4e94860f8f docs: the session-state layer — its four constraints and which way to be wrong (CG-17) 2026-08-05 13:24:06 -05:00
Colby McHenry fc31b1e2bf feat(mcp): remember what explore already served this session (CG-17)
Explore answers every call as if it were the first: no record of the files
and line ranges it already sent, so a 4th call re-serves the 1st call's
spine and the tier call budget can only be asked for, never enforced.

Track it per MCP session, per resolved project root — files, coalesced line
ranges, bytes, and the call's index in the session. Nothing reads it yet:
the response is byte-identical, which the suite pins against an untracked
call of the same query.

The daemon shares ONE ToolHandler and a pool of worker threads across every
connected client, so the state can live neither on the handler nor in a
worker. It lives on MCPSession and is handed to execute() per call; the
session's view rides DOWN on the args and the call's emission rides BACK on
the ToolResult, both as plain properties so they survive the structured
clone to and from a worker. execute() records the emission on the main
thread and deletes it unconditionally — including for callers that track
nothing, like the CLI — so it can never reach the wire. A view a client
spells itself is discarded rather than trusted.

Ranges are reported by the render loop itself (buildSection now returns the
spans it slices alongside the text), and only files that survive the final
hard-ceiling truncation are recorded. Where a bound forces a choice the
record keeps FEWER ranges than were emitted: under-reporting re-serves
something the agent has, over-reporting withholds source it never saw and
costs a Read.

Every bound caps detail only — callCount keeps counting past eviction, so
CG-19's decay can't reset itself every 8 calls.
2026-08-05 13:24:06 -05:00
Colby McHenryandClaude Opus 5 5dd4db68cd docs: the occupancy baseline says our residual is higher — write that down (CG-13)
Fills the empty RESULTS placeholder with the 2026-08-05 campaign
(bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min).

The finding is not the flattering one. Retrieval residual is 82% HIGHER
with codegraph and share-of-context 27% higher, on all seven repos —
vscode 67k resident against 18k. At the same time six of seven
without-arms *process* more total tokens (gin 660k vs 290k) while
leaving less behind. Both are true: one dense verbatim payload stays
resident where many small Read/Grep results evict. This corroborates
issue #1500 on our own harness; the aggregator used to print it as
"-82% lower with codegraph" until the sign bug at 520ed9d.

Also:

- States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the
  README's table is Opus 4.8 / single-question. Measured 24/23/20/84
  against the published 60/69/20/89 — model and turn count, not
  contamination. Records the two inversions honestly (vscode processes
  98% more tokens, django costs 17% more) and that 4 of 28 with-arm
  sessions still touched Read.
- Corrects the "Settled" section, which claimed a 7-repo baseline
  existed before one did, and adds the unclaimed Opus rerun.
- Records the contamination gate: 0 CLI calls returned output in 56
  sessions, but 29 attempts were blocked — 26 of 28 without-arm
  sessions tried. no-cli-shim.sh is load-bearing, not precautionary.
- Records the secondary readings as absolute, not before/after: 86.7%
  allocation efficiency pooled over 110 calls, read-of-a-file-we-
  returned 2%, explore-again 73% and ambiguous by construction.

README.md is deliberately untouched — restating its numbers from sonnet
3-turn data would be wrong. A proposed README paragraph is drafted at
the end of the benchmark doc for the maintainer to accept or reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 13:06:51 -05:00
Colby McHenryandClaude Opus 5 520ed9d933 test(agent-eval): report residual direction by sign, not by hope (CG-13)
The occupancy summary hardcoded "% lower with codegraph". pct(w, wo) is the
reduction going with->without, so a negative value means the with-arm's
residual is LARGER -- and the line printed "-82% lower with codegraph" for the
case where codegraph in fact occupies 82% MORE. A double negative that reads as
a win and inverts the headline of the whole metric.

Direction now follows the sign in words, and the negative case says what the
shape actually is: codegraph front-loads one large verbatim payload that stays
resident, where Read/Grep churn many small results that evict. Fewer total
tokens processed and a larger persistent footprint are both true at once --
that pair is the axis issue #1500 reported.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 13:03:08 -05:00
Colby McHenry 382791f11e docs: one entry point for the three feedback metrics, and how to run them (CG-11)
Three per-metric docs told a maintainer what each number means; none said
which one answers which question, which harness produces it, or how to read
the arm table. agent-eval-feedback-metrics.md is that page — the metric →
question map, when to reach for ab-new-vs-baseline.sh (isolates a change,
both arms codegraph-on) versus run-all.sh (with vs without, a different
question) versus bench-readme.sh, the worked CG-22 express table where all
three read together, and the bucket → fix mapping. Not a fourth restatement:
the derivations stay where they are and each doc now points here.

The caveats that change how the summary table is read are carried over rather
than dropped — allocation efficiency is relative (attribution is by citation,
so same-question builds only, and never "codegraph wastes N%"), occupancy
shares are Claude Code / 200k and do not transfer between hosts while the arm
ratio does, sufficient is not correct, small-n throughout. Plus the
contamination row, which means different things in the two harnesses and is
the first thing to look at in both.

Also records that the CG-8 7-repo bucket block no longer re-derives:
bench-readme.sh overwrites /tmp/ab-readme, so the swept logs are gone. The
current logs give a different distribution over the same 62 calls, and the
CG-8-era and current classifiers agree exactly on them — so nothing moved
under the metric, the corpus did. CG-13 re-establishes the baseline.
2026-08-05 00:59:58 -05:00
Colby McHenry 3e8922dfad test(agent-eval): report all three feedback metrics per arm, side by side (CG-11)
The three metrics existed but only run-all.sh printed them, one block per
run. ab-new-vs-baseline.sh — the harness that actually isolates a retrieval
change, both arms codegraph-on — grepped its parse output down to `by type`
and `Result`, so occupancy, sufficiency and allocation never reached the
maintainer running the A/B they were built for.

Both harnesses now print the three blocks under every run and end with one
compare-arms.mjs table: median [min–max] per arm across RUNS, sufficiency
pooled (it is per-CALL, so median-of-run-percentages would weight a 1-call
run like a 5-call one), allocation pooled by bytes and per run. The table is
"did it move?"; the per-run blocks stay the "why?" — only they name the query
that fell short and the file nothing cited. It reproduces the recorded CG-22
express result off logs already on disk: baseline 3/6 calls in the
`Read a file we returned` bucket at 82.0%, new 0/5 at 96.9%.

parse-bench-readme.mjs gets the same two metrics as a with-arm table, so the
CG-13 campaign aggregates all three rather than occupancy alone.

Also folds the CLI-block shim into no-cli-shim.sh and gives it to
ab-new-vs-baseline.sh. There it is not a with/without leak but an attribution
one, and it breaks all three metrics at once: output arriving through Bash is
charged to Bash in the occupancy table, and an explore issued through the CLI
is not a tool call at all, so it never reaches the sufficiency classifier or
the allocation parse. The run silently drops calls from every number.

The daemon pre-warm and the model policy are untouched.

Validated on one live gin arm (2 explores, 0 Read, all three blocks + table)
and against the cg22/cg15 and ab-readme logs. Selftest 68/68.
2026-08-05 00:58:54 -05:00