The epic record said nothing was open. CG-38 is: agent-named symbols in the
tail of a large file never render, which the epic's probes cannot see because
none of them measures whether the named symbol appeared.
Also corrects a wrong claim made while investigating it. The epic was said to
have regressed its own motivating query; that comparison varied the index as
well as the engine. A controlled bisect holding the index fixed shows the
pre-epic engine rendering 12 lines and CG-36 rendering 463 — the epic strictly
improves the case, and the symbols render at neither.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A file whose top-ranked cluster was trivial kept it, dropped the cluster
carrying the answer WHOLE, and left most of its reservation unspent — because
only the first-chosen cluster could be shrunk. CG-31's carry-forward then
correctly handed that slack down the rank order, so the budget was not merely
unspent but REDIRECTED to weaker files.
django's sql/query.py (score 83, reserved 7,947) went from 1,923 delivered
chars to 10,082, and its envelope share from 7.7% to 40.4%; contrib/admin/
filters.py (score 18) went from 8,057 at 355% of its reservation down to 2,198
at 97%. All 8 starvation flags across the suite clear. Net +1,012 source chars.
The issue named the wrong fix point and the measurement said so: both real
cases lost on maxImportance, NOT on the density tiebreak the issue and its
duplicate (CG-37) suspected. Cluster ranking was left untouched, so the
Session.swift case density-first exists for still works — now pinned by a
dense-header fixture.
Accepted cost: okhttp trades its rank-6 file (score 21, reserved 1,999) for
+7,196 chars in the two files that answer the question, taking it from 6
delivered files to 5. django -159, okhttp -219 and tokio -25 source chars
against the epic tip; gin +1,176, alamofire +187, excalidraw +52. django also
stops cutting its epilogue.
Ships probe-file-spend.mjs, a standing suite-wide probe for reservation vs
spend, so this stays measurable — the original evidence came from ad-hoc
instrumentation that no longer existed and had to be re-derived by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The issue blamed the density tiebreak; both real cases lost on maxImportance,
so ranking was left alone. Full before/after table, the one cost (okhttp's
rank-6 file, squeezed out by reservations that were already structurally
over-subscribed), and what ships to keep it measurable.
A file's ranked clusters were all-or-nothing past the first one: the top-ranked
cluster was taken (shrunk to fit when it had to be) and every cluster below it
was rendered whole, then either fit the remainder or was dropped entirely. On a
file whose top-ranked cluster is TRIVIAL that discards the answer — django's
`db/models/sql/query.py` kept a 22-line glue cluster and dropped the 624-line
`Query` body, spending 1,923 of a 7,947 reservation; okhttp's
`RealInterceptorChain.kt` did the same behind its import header.
The response stayed full, which is why this was invisible: the unspent
reservation carried forward exactly as designed and a file scoring a fifth as
much took the bytes.
Two sites, the same rule — hold the remainder while it is still worth a section
(CG-26's between-FILES lesson, applied between CLUSTERS):
- selection now shrinks a later cluster into what is left of the file's budget,
by the same whole-member rule the first cluster already used;
- the ceiling trim re-renders the weakest cluster into the room that remains
before dropping it. On excalidraw's `typeChecks.ts` the section-cost estimate
missed by 13 chars and a 1,512-char cluster — the file's highest-SCORING one —
was thrown away to pay for it.
Cluster RANKING is untouched: measured, both real cases lost on `maxImportance`,
not on the density tiebreak the issue suspected, and density-first is what keeps
Alamofire's `Session.swift` from burying its methods under the property list.
Suite (6 repos, clean-rebuilt indexes): all 8 starvation flags cleared,
+1,012 source chars net. django's `sql/query.py` 1,923 -> 10,082 of 7,947,
okhttp's `RealInterceptorChain.kt` 1,474 -> 6,038 of 6,058, gin's
`routergroup.go` 3,273 -> 5,632. okhttp trades its rank-6 file (score 21) for
+7,196 chars in the two files that answer the question.
Ships two fixtures pulling in opposite directions (`starved-cluster-ts` and
`dense-header-ts`), a `spendShareAtLeast` gate in probe-allocation, and
probe-file-spend.mjs — a standing per-file reservation-vs-delivered sweep.
Four shipped fixes, one open defect (CG-36), and five issues closed because
measurement contradicted them. The headline is that the reported symptom was
not an explore bug at all — it was a degraded index (CG-33), and the reported
query answers correctly on a clean rebuild with no explore change.
Records the two traps that cost real time and are now guarded in tooling: the
nonexistent .codegraph/graph.db path that sqlite3 silently creates, and
ab-new-vs-baseline.sh swapping src/ mid-run so a commit captures baseline
sources.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both halves of the issue were measured on a hermetic fixture of four
declaration-shaped files varying on banner and depended-on-ness.
CG-25 already handles the motivating file: the Wrangler worker-configuration.d.ts
that opened this issue is demoted by the generated penalty alone, worth 15-46
points of envelope share across four flow queries. No new mechanism for it.
The narrower gap is real. A declaration file with NO banner carried pen 1.00,
took rank #1 and 51% of delivered source on a prose flow query, and displaced
the flow's own entry file out of the response entirely.
The rule is deliberately narrow, and both conditions were derived by survey
rather than guessed. 'Declares no callable and calls nothing' flags 1.1-18.0%
of files across the corpus and catches real source — okhttp's SocketPolicy.kt,
tokio/src/runtime/mod.rs, Alamofire's umbrella file, django's locale format
tables. Requiring every symbol to be type-level drops that to 0-4%. The
'nothing depends on it' condition was added after the broader version demoted a
pure-interface file with 13 inbound imports and broke the CG-31 displacement
gate — a different invariant entirely.
Does NOT stack with the generated penalty: rankPenalty takes Math.min of the
two, so a file that is both takes the stronger, never the product. A query that
NAMES a declaration symbol exempts its file entirely, so asking about a type
still reaches it at full weight.
Six-repo envelope is byte-identical to the pre-change tip — the rule does not
fire on any benchmark repo, consistent with the 0-4% survey.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A file that declares nothing but types and that nothing in the index depends
on — a hand-written ambient `.d.ts` of global shims, vendored typings, module
augmentation — cannot answer a flow question: no bodies, no call edges, no
behaviour, nothing typed by it. But the identifiers it declares are exactly the
generic ones a prose question uses (`Body`, `Message`, `ImageMetadata`,
`ReadableStream`), so on term overlap it out-scored the implementation. Measured
on the new fixture: rank #1 and 51% of delivered source, with the flow's own
entry file pushed out of the response entirely.
Measured first, per the issue: the Wrangler `worker-configuration.d.ts` that
opened this is already handled by CG-25's banner detection, worth 15-46 points
of envelope share across four flow queries. CG-25 credited; only the un-bannered
case needed anything.
`rankPenalty` now multiplies score and graph mass by 0.5 for such files, taken
as the STRONGER of it and the generated penalty rather than multiplied — one
property two signals see must not be charged twice. Detection is structural, not
by extension, and four conditions deep. Two of them were forced by measurement:
requiring every symbol to be type-level takes the corpus flag rate from 1-18%
(which swept in Kotlin sealed classes, Rust mod.rs re-exports and django's
locale tables) down to 0-4%; requiring that nothing depends on the file
separates an ambient shim from a working types module, and without it the rule
demoted displacement-ts's pipeline `types.ts` and broke the CG-31 gate.
A query that NAMES a declared type is exempt, so a question about a type still
reaches its declaration at full weight. Precise tokens only, so "…the file
body…" cannot exempt a `Body` interface it never meant to name; this needs its
own set because `namedSeedIds` is callable-only and a type never becomes one.
Regression evidence in docs/benchmarks/explore-declaration-only-cg28.md:
6-repo envelope sweep byte-identical against a clean baseline build, zero
ambient files reach the candidate set on VS Code across five queries, corpus
flag rate 0-0.74%, both allocation fixtures PASS, full suite 2,978 green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CG-27 proposed adding function/method to ENVELOPE_KINDS so a factory closure
spanning most of its file stops merging every inner symbol into one cluster.
The issue required the ranking claim be measured before any fix. It was, on a
hermetic fixture built to make the pattern maximally visible, and it does not
hold — nothing shipped to src/.
The literal change is a large regression: dropping the enclosing range SPLITS
the file into a trivial cluster (a type alias plus a helper, span 7) and the
answer-bearing one (every closure, span 359). Cluster ranking breaks the equal
maxImportance tie on density, so the trivial cluster wins, is taken first, and
is the only one that may be shrunk; the answer-bearing cluster then does not fit
and is dropped whole. Rank #1 fell from 7,539 delivered chars to 397, and from
7 of 11 inner closures to 0. The enclosing range was holding the file together
as one cluster, inside which shrinkCluster already did the per-symbol ranking
the issue asked for.
A better mechanism reaching the same intent — deferring the envelope member
inside shrinkCluster, leaving clustering untouched — is noise: 69 vs 68 inner
definitions across nine query shapes, one better, one worse, seven unchanged.
The one configuration where the envelope IS selected (the factory as sole
top-tier member) is already absorbed by CG-30, which windows it on whole lines:
a contiguous readable head carrying 6 of 9 closures, bounded and never empty.
Kept: the fixture, the deterministic probe, and the measurement record.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CG-27 asked whether the >50%-of-file envelope drop should cover `function` /
`method`, so a `createFoo()` factory returning an object of closures stops
merging every closure inside it into one cluster. Measured on a hermetic
fixture, it should not, and the issue is closed as obsolete with CG-30 credited.
Two mechanisms already absorb the shape. shrinkCluster orders members by
(importance desc, size ASC) and refuses any member that overruns the cap once
something is kept, so a file-spanning member is only selected when it is the
sole member of the top importance tier — eight of nine query shapes never
selected it at all. When it IS selected, CG-30 windows it on whole lines, so
the file still delivers bounded, readable source (6 of 9 closure definitions
in that configuration).
Dropping the range instead SPLITS the file, and only the first-chosen cluster
may be shrunk: a trivial 7-line cluster won the density tiebreak and the
answer-bearing cluster was dropped whole — rank-#1 file 7,539 chars and 7 of 11
closures to 397 and none. Reaching the same intent more carefully (defer the
envelope MEMBER inside shrinkCluster, leaving clustering untouched) is noise:
69 vs 68 closure definitions across nine query shapes. Nothing shipped.
Adds the fixture, the probe, a standing gate on the outcome, and the record —
including a real defect the measurement exposed on the epic tip: django's
query.py leaves 8,212 of 10,135 unspent and drops a score-290 cluster to keep a
score-14 one. Filed separately.
No behaviour change, so no CHANGELOG entry.
A file whose top-level symbol spans almost all of it — createFoo() returning
an object of closures — is how Svelte 5 rune stores, React custom-hook modules,
IIFE module-pattern JS and Zustand's create((set,get)=>({…})) are all written.
probe-factory-closure.mjs measures what such a file DELIVERS from within: which
inner symbols' definitions reach the agent, not how many bytes did.
A generated Cloudflare Wrangler ambient-types file was not flagged generated, so
it ranked with no penalty and competed with hand-written source on generic token
overlap. The banner shape it uses — "Generated by <tool> by running <command>" —
matched none of the existing content patterns, all of which require DO NOT EDIT,
a standalone @generated, or the "auto(matically) generated by" phrasings.
Precision is held by requiring TWO 'by' clauses: the banner must name a tool and
then say 'by running'. Ordinary prose ("the report is generated by running the
nightly job") has only one and does not match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Lands the three-branch allocation stack. Every admitted file now receives at
least its reservation before any file draws on carry-forward slack, on every
render path — cluster, whole-file GRACE, and whole-file BUY.
CG-30 bounded how far an oversize cluster member may overshoot (windowed on
whole lines past 1.5x rather than emitted whole or dropped). CG-31 gave the
cluster path the `owedBelow` displacement guard the BUY arm always had, holding
back only the prefix of what is owed below that the response can actually pay.
CG-26 closed the three remaining holes: the whole-file arms had no displacement
guard at all, section overhead was charged at a flat 200 against a real 300-500,
and `owedPayableBelow` held all-or-nothing where it should hold partially.
Deterministic across the 6-repo suite, clean-rebuilt indexes, both builds: no
repo truncates, no repo loses a file, okhttp gains one, and every repo lands at
or under the 25,000 hard ceiling.
Accepted trade (maintainer decision): excalidraw -552 and okhttp -164 source
chars against the CG-31 tip, in exchange for the trailing pointer list surviving
instead of being discarded whole. Those bytes existed at the CG-31 tip only
because it over-filled a ceiling it mis-measured and then dropped the entire
epilogue; a pointer the agent can act on beats a few hundred chars on the
last-ranked file.
Two issues opened during this work were closed as invalid rather than fixed:
CG-32 (named-file ordering) and CG-34 (allocator over-reservation). Both were
filed on diagnoses that did not survive measurement — CG-32's symptom was index
drift (CG-33), and CG-34's premise was overturned by CG-31's own results.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Cloudflare Wrangler's `worker-configuration.d.ts` (~12k lines of ambient
types) carried no banner any GENERATED_CONTENT_PATTERNS entry matched:
every existing marker requires `DO NOT EDIT`, a standalone `@generated`,
`<auto-generated>`, or the literal `automatically/auto-generated by`
phrasings. Wrangler emits a bare `Generated by Wrangler by running
`wrangler types``, so the file ranked with pen 1.00 and won 79.4% of an
explore envelope on generic token overlap alone (CG-24).
The discriminator is the reproduction instruction, not the word
"generated": the banner must name a tool AND then say `by running`, i.e.
two separate "by" clauses. That keeps prose out — "the nightly summary is
generated by running the ETL job" has only one — while catching every
CLI-driven emitter that tells you how to regenerate.
Precision swept over 441,856 files across the whole local source tree: 5
hits, all genuine Wrangler output, no false positives.
Isolated before/after on the CG-24 repro (same query, same index, only
the `files.generated` flag differing):
before pen 1.00 score 115.0 share 79.4% 3 files rendered
after pen 0.30 score 35.4 share 21.1% 4 files rendered
The new pattern stays in the existing table position, below the header
window the detector scans, so the module still does not classify itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Deterministic 6-repo table, the three agent A/Bs (django, excalidraw, okhttp,
2 runs/arm, Read 0 in all 12 runs), and an honest read of the two repos that
deliver a few hundred fewer source chars: at the CG-31 tip both were over-filled
by the flat-200 section overhead and paid for it by discarding their epilogue
whole.
Also: CHANGELOG entries for the two user-visible changes, and the memory note
now carries the fourth accounting gap plus the two lessons — hold the REMAINDER
when a full reservation no longer fits, and never skip a file over an accounting
difference.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The invariant this closes: every admitted file receives at least its
reservation before any file draws on carry-forward slack. CG-30 bounded an
oversize cluster member and CG-31 gave the cluster path a displacement guard;
three holes were left, and each one starved a file that had been admitted,
reserved and — in the worst case — rendered.
1. The whole-file arms had no displacement guard. BUY's fit test read
`renderCeiling - totalChars` (everyone's room) while its source-space
sibling refused the same trade, and GRACE was not fit-tested at all.
okhttp's CallServerInterceptor.kt shipped 8,499 chars on a 5,964 funded
ceiling and the rank-6 file below it delivered nothing. Both arms now test
the render they actually produce against `fundedHeadroom`, and a whole
render that does not fit falls through to clustering instead of skipping
the file.
2. Every section was charged a flat 200 chars while a real header runs
300-500. The loop believed it had room it did not have — okhttp allocated
26,601 against a 24,400 ceiling — so the final truncation threw a
fully-rendered section away. Sections are charged their real cost now, the
owed-below arithmetic uses a per-file overhead estimated from the file's own
symbols, and a marginal overrun trims the weakest cluster (or windows the
last one into the room that is left) rather than skipping the file over a
rounding difference.
3. `owedPayableBelow` held all-or-nothing. When the last admitted file's FULL
reservation no longer fit, nothing was held for it: on the precise-query
fixture the rank-5 file took 4,134 chars against a 2,948 reservation while
rank 6 — admitted, reserved 2,539 — was left 4 chars and skipped. It now
holds the remainder while that remainder is still worth a section
(MIN_CHARS).
And the epilogue is budgeted instead of discarded. The flat 600-char margin was
neither the epilogue's size (1,064 gin, 1,788 django, 2,231 excalidraw) nor a
bound on it, so four of six suite repos shipped with no pointer list and no
reminders at all. The loop now reserves the epilogue's FLOOR — the one line
that says an uncovered area exists, plus a pointer for every file whose bytes
were deliberately withheld (CG-12) — and the rest is fitted to the room that
actually remains, in priority order, entry by entry. Sized from the real
strings; no constant was swept against the suite.
Deterministic, same clean-rebuilt indexes, baseline = CG-31 tip:
repo base source new source files ceiling
django 20,791 20,878 6 -> 6 was discarding its epilogue
tokio 21,521 21,607 5 -> 5 was discarding its epilogue
okhttp 19,034 18,870 5 -> 6 +1 file delivered
excalidraw 20,204 19,652 8 -> 8 keeps its pointer list
gin 10,776 10,776 4 -> 4 byte-identical
alamofire 11,662 11,662 2 -> 2 byte-identical
No repo truncates any more and none loses a file. okhttp and excalidraw trade
164 and 552 source chars on their LAST-ranked file for the pointer list naming
what the response could not cover — bytes the CG-31 tip only had because it
over-filled a ceiling it mis-measured and then discarded the epilogue whole.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The '⚠ changed on disk after the last index sync' banner is an honesty claim
about source we DID render — line refs elsewhere in the response may be
shifted — not a note about the response. Drawing the epilogue boundary after
it means the size cut can never be what silences it. Suite numbers unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Deterministic (6 repos, clean rebuilds, both builds): four deliver more source
and one more file each, two are byte-identical, none deliver less. Agent A/B
(django n=3, okhttp n=2, gin n=2, sonnet/effort high, both arms codegraph-on,
0 contamination): the new arm is faster on all three, Read at or below
baseline, occupancy lower.
Also records the two corrections the suite forced on the first cut of the
guard, and the two residuals CG-26 inherits — the render loop's 600-char
epilogue margin (a sweep was run and deliberately NOT shipped) and the BUY
arm's source-space-only guard.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two corrections found by measuring the first cut of the guard against the
6-repo suite. The first version held back the FULL sum of the reservations
below a file. On django that took 2,319 chars off a file the agent receives
and handed them to a section the hard ceiling then threw away — the guard's
own failure mode, one layer down. tokio lost 1,298 the same way.
1. `owedPayableBelow` — hold back only the prefix of what is owed below that
the response can still PAY, in rank order. A promise the ceiling cannot
reach is not a claim on this file's bytes.
2. The final truncation now spends the EPILOGUE before it spends a rendered
file section. It used to cut at the last section header, dropping that
section AND the trailing notes; dropping the notes alone is almost always
enough. A section is source the agent otherwise has to Read; the epilogue
is a pointer list and two reminders, and the note that replaces it carries
the "explore these names" instruction forward.
Also count `flow.text` in `totalChars`. It is prepended to `lines` to make the
final output, so the render loop always spent against a ceiling it was ~2K
under on symbol-bag queries.
Deterministic, same clean-rebuilt indexes, both builds (baseline = CG-30 tip):
repo base source new source files
django 20,033 20,791 5 trunc -> 6
excalidraw 18,776 20,204 7 trunc -> 8
okhttp 15,628 19,034 4 trunc -> 5
tokio 20,340 21,521 4 trunc -> 5
gin 10,776 10,776 4 -> 4 (byte-identical)
alamofire 11,662 11,662 2 -> 2 (byte-identical)
No repo delivers less; four stop truncating. `funded` in the diagnostic now
reports the render CEILING the guard allows, which is what every render path
is actually bounded by.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Carry-forward slack let a file spend what the files ABOVE it left on the
table. Nothing held back what was promised BELOW it. The whole-file BUY arm
has always refused that trade (`owedBelow`); the cluster path read `headroom`
— what is left before the hard ceiling — instead of what is still owed, so
`fileBudget` and `SPINE_CEILING` could pay a 1.5x overshoot out of another
file's reservation.
`fundedHeadroom` is the same inequality in the units the cluster path spends
in: source PLUS the per-section overhead each unreached file will charge.
Floored at the file's own reservation — a kept promise is not a displacement —
and it is <= `headroom` by construction, so it is the only bound the three
render sites need. The skeleton path's `bodyCap` takes it too.
Measured on `__tests__/fixtures/displacement-ts` (a 4-stage pipeline padded
past 500 files, where the 24K envelope genuinely saturates the 24.4K render
ceiling):
before ingest.ts emitted 9,301 on a 6,289 spendable, then lost the whole
section to the final ceiling — 0 delivered. types.ts and sink.ts
skipped `budget-whole-file`. 3 of 6 admitted files delivered.
after ingest.ts bounded to the 4,913 actually free. 6 of 6 delivered,
envelope 14,908 -> 22,066.
The self-query allocation fixture flips back to PASS with it, on a clean full
rebuild of this repo's index (CG-33). Its `afterCG30` verdict blamed an
over-RESERVED incidental file; the reservation was identical in both arms —
the file was over-SPENDING. Recorded honestly in `afterCG31`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Primary evidence is deterministic: on django, query.py rendered 2.12x its budget
on main and 1.49x with the bound, and the freed bytes reach the files below it
(+2,104 chars of source in the same five files). gin is a true control — the two
builds emit byte-identical explore output there, which is what makes its agent-run
deltas variance by construction.
Also records the harness trap that voided the first two batches: ab-new-vs-baseline
swaps src/ to the baseline ref mid-run, so a commit made while it runs captures
baseline sources. Check the "changed:" line before believing any run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
d652c14 was committed while an ab-new-vs-baseline run had the engine checked out
at the BASELINE ref — that harness swaps src/ files mid-run and restores them on
exit — so it captured main's tools.ts and explore-diagnostics.ts and silently
undid 765c06a. Restored from 765c06a, with d652c14's doc tweak re-applied.
The A/B runs started after that commit are void with it (their "changed:" line
lists only explore-diagnostics.ts, i.e. both arms ran the same retrieval code)
and are re-run rather than reported.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The self-query allocation probe fixture's delivered-share gates now fail. The
cause is not the new bound: allocation is unchanged between arms (parse-run.mjs
32.3% vs 33.9% on main) and tools.ts delivers the same 8,282 chars in both. What
changed is that the incidental file now DELIVERS — on main its whole section was
cut by the hard-ceiling truncation, so the fixture passed on truncation luck.
Every file on this repo obeys the new bound (max 1.40x of spendable).
Recorded as `afterCG30` with that reasoning rather than tuning the bound to
restore the pass. The over-reservation it exposes is epic CG-24's subject.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
shrinkCluster keeps an oversize cluster's highest-importance member WHOLE on
purpose — an empty file section sends the agent to Read, the outcome explore
exists to prevent. What it lacked was a bound, and "never empty" quietly meant
"never bounded": on the reporting repo one file emitted 22,376 chars against a
9,181-char reservation (2.44x), past both the per-file budget and the spine
ceiling. That overshoot is what collapses `headroom` for every file below it.
The same rule has a second face. When the top member is bigger than the whole
response ceiling, the file does not overshoot — it is dropped entirely at the
renderCeiling check, so the agent gets nothing for a file it named.
renderCluster now takes a ceiling (1.5x what the file may spend — the same
multiple SPINE_CEILING already draws, and never below the cap, so a cluster
that fits is untouched). Past it the member is WINDOWED on whole lines rather
than emitted whole or dropped: leading window plus, on a flow cluster, a window
on the spine's call site. A partial window shorter than 12 lines is dropped
instead — a sliver in the session record forces the next call's dedup to shred
the block around it or re-send it — unless nothing else was emitted, where the
never-empty floor wins.
Measured on the new fixture, pre-fix vs post-fix:
monthly.ts 12,391 chars on a 3,334 budget (3.7x) → 4,941 (1.48x)
quarterly.ts dropped, no headroom left → 4,004 delivered
Also: the diagnostic now reports `spendable` (reservation + inherited slack)
alongside `reserved`. Every render bound reads the former, so reporting only
the latter makes an ordinary carry-forward read as a file spending over budget
— and it made the overshoot this issue is about unmeasurable. A windowed file
is now flagged `clipped` too, instead of presenting a window as the whole file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swept every benchmark doc for figures produced off `result.usage` and fixed
the ones that had raw logs to re-derive from.
residual-context-occupancy.md — the sonnet 3-turn throughput table. Re-derived
from the preserved logs: tokens saved 23% -> 56%, and vscode's "98% MORE tokens
with codegraph" was never real, it is 41% fewer. Cost, time and tool calls were
never affected by this field and are unchanged. The occupancy table itself is
measured off the timeline, so every number in it stands -- including the 82%
higher residual, which is the finding the document exists for.
call-sequence-analysis.md — this doc DIAGNOSED the bug and its reproduce block
claimed the aggregator summed per-turn tokens. It did not, until 04c0f8e. Noted,
with the three wrong results the gap produced: the excalidraw cut recorded here,
the sonnet campaign, and the Opus re-measure that invented a token regression.
answer-directly-vs-explore-agent.md — build 0.9.4, 2026-05-24, raw logs gone.
Cannot be re-derived, so flagged rather than silently left or invented: its
token figure is indicative, its turn/read/context findings do not depend on the
broken field and stand.
The remaining benchmark docs (allocation-ab-1500, dedup-cg20, allocation-
efficiency, feedback-metrics) carry no throughput tables — checked, clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Re-measured on the CLI-blocked harness at the README's own stated methodology
(claude-opus-4-8, single question, median of 4, same 7 repos), with tokens
summed per turn per 04c0f8e.
What moved, and what did not:
tool calls 89% -> 88% fewer holds
file reads 0 on all seven holds
tokens 69% -> 62% fewer holds
time 20% -> 53% faster was badly understated
cost 60% -> 44% cheaper genuinely lower
Time was the most wrong figure in the old table, in our favour: every repo is
faster with the graph, including the two that previously carried a footnote
excusing them for being slower. That footnote is gone. Cost is the one real
downgrade, and it varies with how much discovery a question demands rather
than with repo size -- 57-78% where the file-reading arm needed 28-43 tool
calls, near-even on Gin at 7. Said plainly instead of averaged away.
Methodology now records that the CLI is blocked in BOTH arms and why: on an
unblocked harness the control reached codegraph through Bash in 26 of 28 runs,
so it was not a control. All 28 attempted it here; all 28 were blocked.
The context-footprint note from a7db24d is unaffected -- occupancy is measured
off the timeline, not the token field that was wrong.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Tokens processed" was read off result.usage. That was correct when the README
figures were measured; in current Claude Code the field reports the LAST turn
only. Nothing here changed — the host did, silently — and the harness kept
reporting the smaller number.
The error is one-sided, which makes it worse than noise: it under-counts
whichever arm takes more turns, and that is always the WITHOUT arm. On the
2026-08-05 campaign it turned a real 62% token saving into 19% and invented a
token REGRESSION on tokio (-41%) and alamofire (-25%) that does not exist. Those
numbers were one push away from the README.
Now summed per assistant request and deduped by message.id, the same rule the
occupancy timeline already used — Claude Code emits one event per content block
carrying identical usage, so summing per event double-counts (~1.7x measured).
CLAUDE.md already warned about this field. The code did not follow; it does now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The benchmark table measures throughput -- tokens processed, tools called,
cost to reach one answer. It has never measured what is still resident in the
window afterward, and on that axis codegraph costs more: ~80% more retrieval
context left behind than a file-reading agent, on all seven repos.
That is the axis issue #1500 reported, and it is structural rather than a
defect -- one dense payload that answers the question and stays, versus
grep-and-read churn that evicts. Worth stating plainly next to the cost note
rather than leaving a user to discover it in a long session.
The Opus 4.8 single-question figures in the table are deliberately untouched:
the occupancy campaign ran sonnet / 3-turn, a different regime, and nothing
measured there licenses restating them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Never re-serve source this session already sent: session-scoped state (CG-17)
plus a precise back-reference in place of the bytes (CG-18). Duplicated source
drops 6.4% -> 0.74% of the response, at flat cost per call, with more unique
source in its place.
Bar 4 of CG-20 -- "residual occupancy must drop" -- is NOT met, and the gate
records why: CG-18's accepted rule spends reclaimed bytes on files not yet
shown rather than banking them, so a design that re-spends every byte cannot
lower the byte count. The two requirements were mutually unsatisfiable as
written. Bars 1-3 (no extra Reads, no abandonment, no bucket shift) pass across
24 runs on both arms.
Kept on that basis, and cheap to reverse: CODEGRAPH_EXPLORE_DEDUP=0 disables it
at runtime. Full gate: docs/benchmarks/explore-dedup-ab-cg20.md
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Read is 0 in all 24 runs of both arms across client-go and excalidraw, no
isError, codegraph is the last tool in every run, and both "Read a file we
returned" / "did not return" buckets are empty — with back-references
demonstrably reaching the agent in 8 of the 9 multi-call runs.
Residual occupancy is flat, and the measurement shows it could not have been
anything else: CG-18 was accepted on the rule that reclaimed bytes get spent on
unseen files rather than banked, so the byte count cannot fall. What moves is
the duplicate fraction of that residual — 87% less across the agent runs, 86%
and 94% on two matched deterministic replays.
Also recorded: dedup.savedChars is a pre-clip figure (11,450 reported against
1,042 chars actually re-served), so tuning off it inflates the win ~7x.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Extends the session-state design doc with the layer built on it: what gates a
withheld span (session, content fingerprint, two size floors, kill switch),
why the fingerprint and not #1474's drift flag, the two channels the
reclaimed bytes leave by, and why "no duplicate ranges across calls" holds
for every call that had something new to say rather than universally.
A later explore call re-served whatever it re-ranked, so on the #1500 report
the 4th call spent its envelope on the spine the 1st call had already
delivered. CG-17 recorded what was served; this acts on it.
What a withheld span becomes is the whole design: a POINTER, never a silence.
An insufficient-feeling response is what sends an agent to Read, and one or
two of those early in a session teach it to abandon codegraph — so the
replacement names the file, the symbols and the line spans, and says both
that the source came from THIS conversation and that the file has not changed
since.
- Content fingerprint, not the drift flag, gates it. They answer different
questions: two calls inside one drift window served the same current bytes,
while a file edited AND re-synced between calls is never "stale" and yet
the agent's copy is now wrong. An edited file re-emits in full.
- Only a covered run of >= 8 lines is replaced, and a remainder under 160
chars folds into the pointer. Below those the pointer costs more than the
source and the block reads as shredded — a fence holding `228\t` is a
broken-looking response, which is the expensive failure.
- The reclaimed bytes go to files the agent has NOT seen, two ways: a smaller
`sourceSpent` hands slack down CG-21's carry-forward pool, and a fully
back-referenced file gives up its maxFiles slot the way a cliffed one does.
Within a file, the cluster shrink now reads the DEDUPED length, so it never
drops new symbols to make room for source it isn't sending.
- If dedup suppresses everything and nothing new takes its place, the top
suppressed file is spliced back in whole. An all-pointer response is the
shape that reads as "codegraph found nothing"; one re-served file is the
cheaper mistake.
Kill switch: CODEGRAPH_EXPLORE_DEDUP=0.
Explore answers every call as if it were the first: no record of the files
and line ranges it already sent, so a 4th call re-serves the 1st call's
spine and the tier call budget can only be asked for, never enforced.
Track it per MCP session, per resolved project root — files, coalesced line
ranges, bytes, and the call's index in the session. Nothing reads it yet:
the response is byte-identical, which the suite pins against an untracked
call of the same query.
The daemon shares ONE ToolHandler and a pool of worker threads across every
connected client, so the state can live neither on the handler nor in a
worker. It lives on MCPSession and is handed to execute() per call; the
session's view rides DOWN on the args and the call's emission rides BACK on
the ToolResult, both as plain properties so they survive the structured
clone to and from a worker. execute() records the emission on the main
thread and deletes it unconditionally — including for callers that track
nothing, like the CLI — so it can never reach the wire. A view a client
spells itself is discarded rather than trusted.
Ranges are reported by the render loop itself (buildSection now returns the
spans it slices alongside the text), and only files that survive the final
hard-ceiling truncation are recorded. Where a bound forces a choice the
record keeps FEWER ranges than were emitted: under-reporting re-serves
something the agent has, over-reporting withholds source it never saw and
costs a Read.
Every bound caps detail only — callCount keeps counting past eviction, so
CG-19's decay can't reset itself every 8 calls.
Fills the empty RESULTS placeholder with the 2026-08-05 campaign
(bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min).
The finding is not the flattering one. Retrieval residual is 82% HIGHER
with codegraph and share-of-context 27% higher, on all seven repos —
vscode 67k resident against 18k. At the same time six of seven
without-arms *process* more total tokens (gin 660k vs 290k) while
leaving less behind. Both are true: one dense verbatim payload stays
resident where many small Read/Grep results evict. This corroborates
issue #1500 on our own harness; the aggregator used to print it as
"-82% lower with codegraph" until the sign bug at 520ed9d.
Also:
- States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the
README's table is Opus 4.8 / single-question. Measured 24/23/20/84
against the published 60/69/20/89 — model and turn count, not
contamination. Records the two inversions honestly (vscode processes
98% more tokens, django costs 17% more) and that 4 of 28 with-arm
sessions still touched Read.
- Corrects the "Settled" section, which claimed a 7-repo baseline
existed before one did, and adds the unclaimed Opus rerun.
- Records the contamination gate: 0 CLI calls returned output in 56
sessions, but 29 attempts were blocked — 26 of 28 without-arm
sessions tried. no-cli-shim.sh is load-bearing, not precautionary.
- Records the secondary readings as absolute, not before/after: 86.7%
allocation efficiency pooled over 110 calls, read-of-a-file-we-
returned 2%, explore-again 73% and ambiguous by construction.
README.md is deliberately untouched — restating its numbers from sonnet
3-turn data would be wrong. A proposed README paragraph is drafted at
the end of the benchmark doc for the maintainer to accept or reject.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The occupancy summary hardcoded "% lower with codegraph". pct(w, wo) is the
reduction going with->without, so a negative value means the with-arm's
residual is LARGER -- and the line printed "-82% lower with codegraph" for the
case where codegraph in fact occupies 82% MORE. A double negative that reads as
a win and inverts the headline of the whole metric.
Direction now follows the sign in words, and the negative case says what the
shape actually is: codegraph front-loads one large verbatim payload that stays
resident, where Read/Grep churn many small results that evict. Fewer total
tokens processed and a larger persistent footprint are both true at once --
that pair is the axis issue #1500 reported.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three per-metric docs told a maintainer what each number means; none said
which one answers which question, which harness produces it, or how to read
the arm table. agent-eval-feedback-metrics.md is that page — the metric →
question map, when to reach for ab-new-vs-baseline.sh (isolates a change,
both arms codegraph-on) versus run-all.sh (with vs without, a different
question) versus bench-readme.sh, the worked CG-22 express table where all
three read together, and the bucket → fix mapping. Not a fourth restatement:
the derivations stay where they are and each doc now points here.
The caveats that change how the summary table is read are carried over rather
than dropped — allocation efficiency is relative (attribution is by citation,
so same-question builds only, and never "codegraph wastes N%"), occupancy
shares are Claude Code / 200k and do not transfer between hosts while the arm
ratio does, sufficient is not correct, small-n throughout. Plus the
contamination row, which means different things in the two harnesses and is
the first thing to look at in both.
Also records that the CG-8 7-repo bucket block no longer re-derives:
bench-readme.sh overwrites /tmp/ab-readme, so the swept logs are gone. The
current logs give a different distribution over the same 62 calls, and the
CG-8-era and current classifiers agree exactly on them — so nothing moved
under the metric, the corpus did. CG-13 re-establishes the baseline.
The three metrics existed but only run-all.sh printed them, one block per
run. ab-new-vs-baseline.sh — the harness that actually isolates a retrieval
change, both arms codegraph-on — grepped its parse output down to `by type`
and `Result`, so occupancy, sufficiency and allocation never reached the
maintainer running the A/B they were built for.
Both harnesses now print the three blocks under every run and end with one
compare-arms.mjs table: median [min–max] per arm across RUNS, sufficiency
pooled (it is per-CALL, so median-of-run-percentages would weight a 1-call
run like a 5-call one), allocation pooled by bytes and per run. The table is
"did it move?"; the per-run blocks stay the "why?" — only they name the query
that fell short and the file nothing cited. It reproduces the recorded CG-22
express result off logs already on disk: baseline 3/6 calls in the
`Read a file we returned` bucket at 82.0%, new 0/5 at 96.9%.
parse-bench-readme.mjs gets the same two metrics as a with-arm table, so the
CG-13 campaign aggregates all three rather than occupancy alone.
Also folds the CLI-block shim into no-cli-shim.sh and gives it to
ab-new-vs-baseline.sh. There it is not a with/without leak but an attribution
one, and it breaks all three metrics at once: output arriving through Bash is
charged to Bash in the occupancy table, and an explore issued through the CLI
is not a tool call at all, so it never reaches the sufficiency classifier or
the allocation parse. The run silently drops calls from every number.
The daemon pre-warm and the model policy are untouched.
Validated on one live gin arm (2 explores, 0 Read, all three blocks + table)
and against the cg22/cg15 and ab-readme logs. Selftest 68/68.
Records the sweep over every A/B log on this machine (103 sessions, 297
explore calls, 0 crashes) and, more usefully, the new-vs-baseline arm
table the metric exists for: express 82% → 100%, cg21/client-go 67% →
95%, two pairs going the other way.
States the caveat in the places it can be misread: the corpus median sits
in the eighties because these are flow questions whose answers name most
of the chain, the metric is byte-weighted, and an agent can use a file
without citing it. It compares two builds on one question; it is not an
absolute waste figure.
The envelope view needed a human to say which files answer the question
(`--answer <glob>`). This reads it off the agent's own final answer and
reports one number per run and per call: bytes returned for files the
answer cited, over all bytes returned. That is the #1500 defect as a
number instead of a hunch.
Attribution has two channels, ranked so the weaker one stays separable:
the answer naming the file (reported alone as the conservative floor),
and the answer citing, in a code span, a symbol only that file DEFINES.
Three guards keep the error from leaning optimistic — the direction a
tuning metric must not lean:
* Only symbols the file defines. Section headers render `name(kind)`
for call sites too, and crediting those marked excalidraw's
dragElements.ts used because the answer named `mutateElement`.
* A definition beats an import alias of the same name (`variable`),
or `lib/application.js` gets credit for `require('./utils')`.
* A name on 3+ returned files identifies none of them.
Bare basenames count as citations (agents write `utils.js:225` in prose)
but only for extensions the envelope shipped, so `res.send` and
`mime.contentType` — the same token shape — do not read as files.
Both the envelope view and this share one parse of the rendered markdown
(parseExploreCall), still not the CG-4 sidecar: the sidecar exists only
on a post-CG-4 build and so cannot measure a baseline arm.
The 14 multi-turn with-arm sessions of the README corpus exercise every bucket
(62 calls), so the sweep is no longer one repo family: 47% explored again, 11%
Read a file we returned, 2% Read a file we did not, 23% Grep/Glob, 18% moved on.
Flagged as a baseline rather than a verdict -- three-turn sessions on hard flow
questions, and "explored again" includes the legitimate second call on a repo
whose budget is 2-3.
The recall bucket's one real instance is worth reading: explore returned
InteractiveCanvas.tsx and named StaticCanvas.tsx without shipping it, and the
agent went and read exactly that.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Records what each bucket means and which fix it points at, the four rules that
keep the classification honest (same-message calls, bookkeeping tools, subagent
threads, earlier-explore files), and the three real transcripts it was
hand-checked against -- including the excalidraw canvasNonce run, where it
independently found the data-flow frontier CLAUDE.md already documents: 0%
sufficient, without being told what to look for.
Also states what it does NOT say: sufficient is not correct, one Read is a vote
rather than a proof, and bucket 1 is ambiguous by construction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The agent's next action after a codegraph_explore is free ground truth about
whether the response was enough, and the harness was discarding it. Every run
now bucketed: explored again (insufficient), Read a file we returned
(allocation -- right file, wrong bytes), Read a file we did not return (recall),
Grep/Glob (recall, weak), or moved on (sufficient). The buckets are chosen so
each one names a distinct fix.
The classifier lives in parse-run.mjs next to the occupancy math and takes raw
events, so parse-session.mjs reuses it for interactive runs -- no new
scripts/agent-eval/*.mjs, which would score into the self-query eval fixture's
corpus.
Four rules, three of them found by validating against real transcripts rather
than reasoned up front:
* A call in the SAME assistant message as the explore predates its response,
so it is not a verdict on it. Stepped over, counted as `concurrent`.
* ToolSearch/TodoWrite carry no signal; the call behind them is the verdict.
* SUBAGENTS ARE A SEPARATE THREAD. Claude Code interleaves a subagent's calls
into the same stream under parent_tool_use_id -- verified on a live
excalidraw run where a delegated search's greps landed between the parent's
own calls. Matching reactions across threads scored the subagent's grep as
the parent's verdict on an explore it never saw. A delegation is judged by
what the subagent did FIRST: before that, the same run reported 33%
sufficient while the subagent was off grepping for the file, which is the
one direction of error a tuning metric must not have. In interactive
sessions the subagent is a separate FILE instead, so parse-session.mjs
stitches the threads back with the toolUseId in agent-*.meta.json.
* A re-read of a file an EARLIER explore shipped is still allocation, not
recall -- filing it as recall aims the fix at the wrong end of the pipeline.
Shell file access counts too (`sed -n 100,200p f` reads, `grep`/`find` search),
since both arms have Bash and counting only the Read tool would score those
explores as sufficient. A heredoc or redirect is writing, not reading.
Validated by hand on cg22/ab-express/run-baseline-1 (explore, explore, Read of
lib/response.js which explore #2 returned -- the #1500 allocation bug as a
bucket instead of a hunch; the new-build arm is 100% sufficient) and on
cg15/ab-express/run-new-2 (four explores, the last returning lib/utils.js which
the agent then read at offset 195). Swept over all 76 A/B logs on this machine:
0 crashes, 176 calls bucketed. --selftest covers every bucket, both thread
rules, delegation, shell reads/searches and errored calls: 46/46.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CG-3 branched from main before CG-1 landed and rewrote parse-run.mjs wholesale
into an exported parseSession(), which dropped CG-1's --envelope/--answer
reporting entirely. That view is the instrument the CG-1/CG-22 allocation gate
measures bar 2 with, and it is in that benchmark's documented reproduce steps,
so it cannot be lost to the merge.
Resolution takes CG-3's rewrite as the structure and ports the envelope feature
into it: parseSession now collects codegraph_explore response text in call
order, formatEnvelope renders the per-file share, and the CLI parses
--envelope/--answer ahead of the positional filter so a glob is never mistaken
for a log path.
The glob sentinel stays written as a \u0000 escape, never a literal NUL byte --
a raw one makes git treat the whole script as binary, exactly as the comment
there warns.
Verified: --selftest 18/18, and a synthetic explore transcript reports the
expected per-file shares and answer-set total.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The hook denies the invocation, so a denied attempt puts no codegraph output in
the window and must not disqualify the run -- only a call that actually returned
content does. Attempts are still reported, since an agent hunting for the CLI is
worth seeing.
An agent denied `codegraph` on PATH ran `find / -maxdepth 4 -iname "*codegraph*"`,
found the binary, and invoked it by ABSOLUTE PATH — 12 times in one without-arm
run. So block the invocation itself with a PreToolUse hook on Bash, written into
the run's output dir as an artifact alongside the MCP configs rather than as a
repo file.
The pattern matches command positions only, so looking is still allowed and only
using is denied: `grep codegraph src/`, `ls .codegraph` and `which codegraph`
pass through, while `codegraph explore`, `/abs/path/codegraph …`, `cd x &&
codegraph …` and `VAR=1 codegraph …` are refused. run-all.sh proves both
directions at startup and refuses to run if either fails. parse-run.mjs's
detector uses the same rule, so prevention and detection cannot drift — and it no
longer false-positives on the corpus path, which contains the word codegraph.
Verified end-to-end: the without-arm now probes with `ls .codegraph; which
codegraph`, finds nothing usable, and falls back to Read/Bash.
The without-arm had no MCP server but still had Bash, and the target repo carries
the .codegraph/ index the with-arm needs. Agents found it: 14 of 15 without-arm
runs in a 7-repo pass ran `codegraph explore` through Bash, one of them via
`ls .codegraph && codegraph explore ...`. That arm was measuring
codegraph-over-CLI, not codegraph-absent, so every without-arm number it produced
was wrong. It bit the with-arm too -- output arriving through Bash is attributed
to Bash, understating what codegraph itself occupies (1 of 15 runs).
Both arms now run on a PATH where the CLI is hidden, so the MCP server is the
only way to reach codegraph and stays the single variable. The binary shares a
directory with tools the run needs -- claude itself sits next to it -- so the
directory is substituted in place by one of symlinks to every entry except
codegraph, preserving PATH order and precedence. The run aborts if claude or node
did not survive the substitution.
Prevention alone would fail silently the next time the CLI lands somewhere new,
so parse-run.mjs flags any Bash command naming codegraph and parse-bench-readme
drops contaminated without-arm runs from the aggregate (CG_INCLUDE_CONTAMINATED=1
keeps them). CG_ARMS re-runs one arm without redoing the other.
A median over 2-3 runs hides swings big enough to flip a repo's sign. On vscode
the without-arm ranged 40k to 67k and the with-arm 59k to 65k across two runs;
the deciding variable is the tool mix, since a with-arm run that reads files ON
TOP of calling explore pays for both.
parse-run.mjs --selftest runs the math over synthetic transcripts with known
answers: attribution, message.id dedupe, compact_boundary, FIFO micro-compaction,
and multi-turn stitching. It found a real bug. A gap where the window also SHED
content has a delta far below what was added, which reads as absurdly dense text
and dragged the whole run's ratio with it -- a shed gap in the fixture pushed
2.5 chars/tok to 4.4 and left the wrong result resident. Shedding can only push a
gap's ratio up, so the calibration now takes the lower median as its centre,
drops gaps well above it, and pools the rest. Runs that never shed are unaffected
(gin and vscode re-measure identically).
Also drafts docs/benchmarks/residual-context-occupancy.md -- method, error bar,
and the limitations this metric does not settle. Baseline numbers to follow.