Resolves the CHANGELOG conflict — main and this branch each prepended a
bullet to [Unreleased] > Fixes; both are kept. Everything else auto-merged,
including src/mcp/tools.ts, which main reworked heavily for the explore
allocation/displacement work (CG-28/31/36/38) while this branch added the
`union` kind to its container sets.
Verified on the merged tree with the native kernel built: 3070 passed,
9 skipped, 0 failed.
Making unions first-class nodes leaves the third loss in #1515 open:
interfaceOverrideEdges enumerates its concrete side as ['class','struct'],
so a union implementor is skipped even though it now has a real node and a
real `implements` edge. "Who implements this trait" then answers wrongly
rather than incompletely — the struct beside it bridges and the union does
not.
Add 'union' to that tuple, plus a regression test that pins the Rust
trait -> union-impl hop (the struct implementor is the control proving the
synthesizer ran). Verified the test fails on the union assertion alone
before this change.
No EXTRACTION_VERSION bump: main is already at 25 against v1.5.0's 24, so
existing indexes are flagged stale for the next release regardless, and
over-bumping is what turns the re-index hint into noise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Recut of #678 against the current instructions — the original predated the
explore-first rewrite and conflicted in both files it touched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`codegraph_explore` never returned `queueMessage` (L1087) or
`flushQueuedMessages` (L1102) from a 1,414-line file, on a symbol bag or a
prose question, even with that file at rank #1 holding 67% of the envelope —
the agent got a same-stem `QueuedMessage` interface at L70 and had to Read the
file for the functions it had named. Pre-existing at every build including
pre-epic (controlled bisect, index held fixed).
Two independent causes:
1. `buildFlowFromNamedSymbols` returns the Flow prose AND the set of node ids
the agent named — and the latter is the whole guarantee, since it injects a
named def into its file's cluster ranges at importance 9. Its bail-outs
returned EMPTY, zeroing the identity whenever there was nothing to PRINT.
Two sibling closures that never call each other produce no chain, no synth
hop and no boundary, so both defs lost importance 9 and the file rendered
from its head. `identityOnly()` now separates the two, gated on
shape-precise tokens so a prose word that exact-matches a callable cannot
promote itself.
2. The ceiling trim filled in SOURCE order, so an over-ceiling render always
dropped the END of a large file first. The shrink HAD kept both symbols
(1022-1121); the trim cut back to 839. `windowToCeiling` now takes the
spine call site plus every importance>=9 member as focus lines, tries the
full ceiling first, and splits the held-back reserve evenly with
carry-forward — greedy-in-source-order reproduced the bug one level down.
The shrink's loose size estimate is left alone deliberately, and the comment
now says why: making it exact was built and measured WORSE (it stops at the
last member that fits whole and the released bytes carry forward to
lower-ranked files, costing payroll-go's `s.store.Upsert`). `bound()` clamps to
the ceiling anyway, so the slack costs no bytes; it just must not pick the
survivors, which is what the trim now handles.
The measurement gap this closes: every existing probe is aggregate — envelope
share, per-file spend, source totals, file counts — and all are green on a
response that returns 25K from the right file and omits the named function.
`probe-named-symbol.mjs` checks the definition LINE against the response's
rendered lines, per symbol.
Suite envelope byte-identical to main on all six repos; probe-allocation 4/4,
no starvation flags; 180 files / 2,997 tests green. Fixture: 7/7 fail on main,
7/7 pass here, deterministic over 4 runs per arm.
The epic record said nothing was open. CG-38 is: agent-named symbols in the
tail of a large file never render, which the epic's probes cannot see because
none of them measures whether the named symbol appeared.
Also corrects a wrong claim made while investigating it. The epic was said to
have regressed its own motivating query; that comparison varied the index as
well as the engine. A controlled bisect holding the index fixed shows the
pre-epic engine rendering 12 lines and CG-36 rendering 463 — the epic strictly
improves the case, and the symbols render at neither.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A file whose top-ranked cluster was trivial kept it, dropped the cluster
carrying the answer WHOLE, and left most of its reservation unspent — because
only the first-chosen cluster could be shrunk. CG-31's carry-forward then
correctly handed that slack down the rank order, so the budget was not merely
unspent but REDIRECTED to weaker files.
django's sql/query.py (score 83, reserved 7,947) went from 1,923 delivered
chars to 10,082, and its envelope share from 7.7% to 40.4%; contrib/admin/
filters.py (score 18) went from 8,057 at 355% of its reservation down to 2,198
at 97%. All 8 starvation flags across the suite clear. Net +1,012 source chars.
The issue named the wrong fix point and the measurement said so: both real
cases lost on maxImportance, NOT on the density tiebreak the issue and its
duplicate (CG-37) suspected. Cluster ranking was left untouched, so the
Session.swift case density-first exists for still works — now pinned by a
dense-header fixture.
Accepted cost: okhttp trades its rank-6 file (score 21, reserved 1,999) for
+7,196 chars in the two files that answer the question, taking it from 6
delivered files to 5. django -159, okhttp -219 and tokio -25 source chars
against the epic tip; gin +1,176, alamofire +187, excalidraw +52. django also
stops cutting its epilogue.
Ships probe-file-spend.mjs, a standing suite-wide probe for reservation vs
spend, so this stays measurable — the original evidence came from ad-hoc
instrumentation that no longer existed and had to be re-derived by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The issue blamed the density tiebreak; both real cases lost on maxImportance,
so ranking was left alone. Full before/after table, the one cost (okhttp's
rank-6 file, squeezed out by reservations that were already structurally
over-subscribed), and what ships to keep it measurable.
A file's ranked clusters were all-or-nothing past the first one: the top-ranked
cluster was taken (shrunk to fit when it had to be) and every cluster below it
was rendered whole, then either fit the remainder or was dropped entirely. On a
file whose top-ranked cluster is TRIVIAL that discards the answer — django's
`db/models/sql/query.py` kept a 22-line glue cluster and dropped the 624-line
`Query` body, spending 1,923 of a 7,947 reservation; okhttp's
`RealInterceptorChain.kt` did the same behind its import header.
The response stayed full, which is why this was invisible: the unspent
reservation carried forward exactly as designed and a file scoring a fifth as
much took the bytes.
Two sites, the same rule — hold the remainder while it is still worth a section
(CG-26's between-FILES lesson, applied between CLUSTERS):
- selection now shrinks a later cluster into what is left of the file's budget,
by the same whole-member rule the first cluster already used;
- the ceiling trim re-renders the weakest cluster into the room that remains
before dropping it. On excalidraw's `typeChecks.ts` the section-cost estimate
missed by 13 chars and a 1,512-char cluster — the file's highest-SCORING one —
was thrown away to pay for it.
Cluster RANKING is untouched: measured, both real cases lost on `maxImportance`,
not on the density tiebreak the issue suspected, and density-first is what keeps
Alamofire's `Session.swift` from burying its methods under the property list.
Suite (6 repos, clean-rebuilt indexes): all 8 starvation flags cleared,
+1,012 source chars net. django's `sql/query.py` 1,923 -> 10,082 of 7,947,
okhttp's `RealInterceptorChain.kt` 1,474 -> 6,038 of 6,058, gin's
`routergroup.go` 3,273 -> 5,632. okhttp trades its rank-6 file (score 21) for
+7,196 chars in the two files that answer the question.
Ships two fixtures pulling in opposite directions (`starved-cluster-ts` and
`dense-header-ts`), a `spendShareAtLeast` gate in probe-allocation, and
probe-file-spend.mjs — a standing per-file reservation-vs-delivered sweep.
Four shipped fixes, one open defect (CG-36), and five issues closed because
measurement contradicted them. The headline is that the reported symptom was
not an explore bug at all — it was a degraded index (CG-33), and the reported
query answers correctly on a clean rebuild with no explore change.
Records the two traps that cost real time and are now guarded in tooling: the
nonexistent .codegraph/graph.db path that sqlite3 silently creates, and
ab-new-vs-baseline.sh swapping src/ mid-run so a commit captures baseline
sources.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both halves of the issue were measured on a hermetic fixture of four
declaration-shaped files varying on banner and depended-on-ness.
CG-25 already handles the motivating file: the Wrangler worker-configuration.d.ts
that opened this issue is demoted by the generated penalty alone, worth 15-46
points of envelope share across four flow queries. No new mechanism for it.
The narrower gap is real. A declaration file with NO banner carried pen 1.00,
took rank #1 and 51% of delivered source on a prose flow query, and displaced
the flow's own entry file out of the response entirely.
The rule is deliberately narrow, and both conditions were derived by survey
rather than guessed. 'Declares no callable and calls nothing' flags 1.1-18.0%
of files across the corpus and catches real source — okhttp's SocketPolicy.kt,
tokio/src/runtime/mod.rs, Alamofire's umbrella file, django's locale format
tables. Requiring every symbol to be type-level drops that to 0-4%. The
'nothing depends on it' condition was added after the broader version demoted a
pure-interface file with 13 inbound imports and broke the CG-31 displacement
gate — a different invariant entirely.
Does NOT stack with the generated penalty: rankPenalty takes Math.min of the
two, so a file that is both takes the stronger, never the product. A query that
NAMES a declaration symbol exempts its file entirely, so asking about a type
still reaches it at full weight.
Six-repo envelope is byte-identical to the pre-change tip — the rule does not
fire on any benchmark repo, consistent with the 0-4% survey.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A file that declares nothing but types and that nothing in the index depends
on — a hand-written ambient `.d.ts` of global shims, vendored typings, module
augmentation — cannot answer a flow question: no bodies, no call edges, no
behaviour, nothing typed by it. But the identifiers it declares are exactly the
generic ones a prose question uses (`Body`, `Message`, `ImageMetadata`,
`ReadableStream`), so on term overlap it out-scored the implementation. Measured
on the new fixture: rank #1 and 51% of delivered source, with the flow's own
entry file pushed out of the response entirely.
Measured first, per the issue: the Wrangler `worker-configuration.d.ts` that
opened this is already handled by CG-25's banner detection, worth 15-46 points
of envelope share across four flow queries. CG-25 credited; only the un-bannered
case needed anything.
`rankPenalty` now multiplies score and graph mass by 0.5 for such files, taken
as the STRONGER of it and the generated penalty rather than multiplied — one
property two signals see must not be charged twice. Detection is structural, not
by extension, and four conditions deep. Two of them were forced by measurement:
requiring every symbol to be type-level takes the corpus flag rate from 1-18%
(which swept in Kotlin sealed classes, Rust mod.rs re-exports and django's
locale tables) down to 0-4%; requiring that nothing depends on the file
separates an ambient shim from a working types module, and without it the rule
demoted displacement-ts's pipeline `types.ts` and broke the CG-31 gate.
A query that NAMES a declared type is exempt, so a question about a type still
reaches its declaration at full weight. Precise tokens only, so "…the file
body…" cannot exempt a `Body` interface it never meant to name; this needs its
own set because `namedSeedIds` is callable-only and a type never becomes one.
Regression evidence in docs/benchmarks/explore-declaration-only-cg28.md:
6-repo envelope sweep byte-identical against a clean baseline build, zero
ambient files reach the candidate set on VS Code across five queries, corpus
flag rate 0-0.74%, both allocation fixtures PASS, full suite 2,978 green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CG-27 proposed adding function/method to ENVELOPE_KINDS so a factory closure
spanning most of its file stops merging every inner symbol into one cluster.
The issue required the ranking claim be measured before any fix. It was, on a
hermetic fixture built to make the pattern maximally visible, and it does not
hold — nothing shipped to src/.
The literal change is a large regression: dropping the enclosing range SPLITS
the file into a trivial cluster (a type alias plus a helper, span 7) and the
answer-bearing one (every closure, span 359). Cluster ranking breaks the equal
maxImportance tie on density, so the trivial cluster wins, is taken first, and
is the only one that may be shrunk; the answer-bearing cluster then does not fit
and is dropped whole. Rank #1 fell from 7,539 delivered chars to 397, and from
7 of 11 inner closures to 0. The enclosing range was holding the file together
as one cluster, inside which shrinkCluster already did the per-symbol ranking
the issue asked for.
A better mechanism reaching the same intent — deferring the envelope member
inside shrinkCluster, leaving clustering untouched — is noise: 69 vs 68 inner
definitions across nine query shapes, one better, one worse, seven unchanged.
The one configuration where the envelope IS selected (the factory as sole
top-tier member) is already absorbed by CG-30, which windows it on whole lines:
a contiguous readable head carrying 6 of 9 closures, bounded and never empty.
Kept: the fixture, the deterministic probe, and the measurement record.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CG-27 asked whether the >50%-of-file envelope drop should cover `function` /
`method`, so a `createFoo()` factory returning an object of closures stops
merging every closure inside it into one cluster. Measured on a hermetic
fixture, it should not, and the issue is closed as obsolete with CG-30 credited.
Two mechanisms already absorb the shape. shrinkCluster orders members by
(importance desc, size ASC) and refuses any member that overruns the cap once
something is kept, so a file-spanning member is only selected when it is the
sole member of the top importance tier — eight of nine query shapes never
selected it at all. When it IS selected, CG-30 windows it on whole lines, so
the file still delivers bounded, readable source (6 of 9 closure definitions
in that configuration).
Dropping the range instead SPLITS the file, and only the first-chosen cluster
may be shrunk: a trivial 7-line cluster won the density tiebreak and the
answer-bearing cluster was dropped whole — rank-#1 file 7,539 chars and 7 of 11
closures to 397 and none. Reaching the same intent more carefully (defer the
envelope MEMBER inside shrinkCluster, leaving clustering untouched) is noise:
69 vs 68 closure definitions across nine query shapes. Nothing shipped.
Adds the fixture, the probe, a standing gate on the outcome, and the record —
including a real defect the measurement exposed on the epic tip: django's
query.py leaves 8,212 of 10,135 unspent and drops a score-290 cluster to keep a
score-14 one. Filed separately.
No behaviour change, so no CHANGELOG entry.
A file whose top-level symbol spans almost all of it — createFoo() returning
an object of closures — is how Svelte 5 rune stores, React custom-hook modules,
IIFE module-pattern JS and Zustand's create((set,get)=>({…})) are all written.
probe-factory-closure.mjs measures what such a file DELIVERS from within: which
inner symbols' definitions reach the agent, not how many bytes did.
A generated Cloudflare Wrangler ambient-types file was not flagged generated, so
it ranked with no penalty and competed with hand-written source on generic token
overlap. The banner shape it uses — "Generated by <tool> by running <command>" —
matched none of the existing content patterns, all of which require DO NOT EDIT,
a standalone @generated, or the "auto(matically) generated by" phrasings.
Precision is held by requiring TWO 'by' clauses: the banner must name a tool and
then say 'by running'. Ordinary prose ("the report is generated by running the
nightly job") has only one and does not match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Lands the three-branch allocation stack. Every admitted file now receives at
least its reservation before any file draws on carry-forward slack, on every
render path — cluster, whole-file GRACE, and whole-file BUY.
CG-30 bounded how far an oversize cluster member may overshoot (windowed on
whole lines past 1.5x rather than emitted whole or dropped). CG-31 gave the
cluster path the `owedBelow` displacement guard the BUY arm always had, holding
back only the prefix of what is owed below that the response can actually pay.
CG-26 closed the three remaining holes: the whole-file arms had no displacement
guard at all, section overhead was charged at a flat 200 against a real 300-500,
and `owedPayableBelow` held all-or-nothing where it should hold partially.
Deterministic across the 6-repo suite, clean-rebuilt indexes, both builds: no
repo truncates, no repo loses a file, okhttp gains one, and every repo lands at
or under the 25,000 hard ceiling.
Accepted trade (maintainer decision): excalidraw -552 and okhttp -164 source
chars against the CG-31 tip, in exchange for the trailing pointer list surviving
instead of being discarded whole. Those bytes existed at the CG-31 tip only
because it over-filled a ceiling it mis-measured and then dropped the entire
epilogue; a pointer the agent can act on beats a few hundred chars on the
last-ranked file.
Two issues opened during this work were closed as invalid rather than fixed:
CG-32 (named-file ordering) and CG-34 (allocator over-reservation). Both were
filed on diagnoses that did not survive measurement — CG-32's symptom was index
drift (CG-33), and CG-34's premise was overturned by CG-31's own results.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Cloudflare Wrangler's `worker-configuration.d.ts` (~12k lines of ambient
types) carried no banner any GENERATED_CONTENT_PATTERNS entry matched:
every existing marker requires `DO NOT EDIT`, a standalone `@generated`,
`<auto-generated>`, or the literal `automatically/auto-generated by`
phrasings. Wrangler emits a bare `Generated by Wrangler by running
`wrangler types``, so the file ranked with pen 1.00 and won 79.4% of an
explore envelope on generic token overlap alone (CG-24).
The discriminator is the reproduction instruction, not the word
"generated": the banner must name a tool AND then say `by running`, i.e.
two separate "by" clauses. That keeps prose out — "the nightly summary is
generated by running the ETL job" has only one — while catching every
CLI-driven emitter that tells you how to regenerate.
Precision swept over 441,856 files across the whole local source tree: 5
hits, all genuine Wrangler output, no false positives.
Isolated before/after on the CG-24 repro (same query, same index, only
the `files.generated` flag differing):
before pen 1.00 score 115.0 share 79.4% 3 files rendered
after pen 0.30 score 35.4 share 21.1% 4 files rendered
The new pattern stays in the existing table position, below the header
window the detector scans, so the module still does not classify itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The suite passed unchanged with `CODEGRAPH_NO_REBIND=1`, so the larger half
of CG-33 — the rebind pass — had no coverage at all.
The cause was the ground truth, not the cases: `rebuildEdgeSet` called
`indexAll()` on the live handle. That is not a rebuild. Every file hashes
identical, so the store writes nothing (`nodesCreated: 0`), no reference is
re-created, and every edge survives — the comparison read the synced index
against itself and could never fail. It now goes through `CodeGraph.recreate`,
which deletes the database file the way the CLI's `index` command does.
With a real rebuild, three existing cases fail under the kill switch. Adds two
more for the rules that carry the risk:
- an edge with no `refName` stamp (older engine) and a synthesized
(`provenance='heuristic'`) edge are never deleted — both planted directly,
and each verified load-bearing by mutation;
- a name over the 500-edge ceiling is declined losslessly rather than
rebound in part, with a rare name in the same sync as the control that
proves the pass ran.
The per-file-vs-batch-wide delta rule is likewise confirmed by mutation: a
batch-wide name set fails its case.
CODEGRAPH_NO_REBIND=1 now fails 4 cases; unset is green; full suite green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A live, auto-synced index did not converge to a clean rebuild of the same
tree — 4.3% of distinct edges wrong in both directions on this repo's own
index, overwhelmingly `calls`, which is what flow queries traverse and what
explore's file ranking weights. Silent: nothing warned, and the symptom read
as "codegraph isn't very good" rather than "this index needs rebuilding."
Two causes, and the fix needed both. Resolution binds a reference to one of
the same-named definitions PROJECT-WIDE, so a definition appearing or
vanishing changes the correct answer for references in files the sync never
touches — and those references resolved successfully once, which deletes
their unresolved_refs row, leaving nothing to revisit them with (#1240's
retry only revisits refs parked as failed). Separately, when nothing
disambiguated the candidates the winner came down to rowid, i.e. the order
files happened to be WRITTEN, which differs between a scan-order full index
and a sync that appends each file as it changes. That second one is why
re-resolution alone could not converge: re-resolving against the identical
graph still picked a different candidate.
So getNodesByName now orders by (file_path, start_line) — a property of the
code, not of the write order — and sync computes a definitionDelta and
re-opens the resolution edges whose answer it may have invalidated,
re-inserting each as the reference that created it for the orphan sweep to
bind against the post-sync graph.
The delta compares `file\0name` pairs per file rather than one name set over
the batch: a commit that adds `collect` to a new file while an unrelated
changed file already defines `collect` cancels out of a batch-wide set, and
that miss was the largest residual class in the first measurement.
Conservative where the failure modes are asymmetric — a wrong deletion is a
permanent edge loss, a missed rebind is only residual drift. Edges without a
refName stamp are never touched (nothing to restore them from), sources the
sync already re-extracted are skipped, and a per-name ceiling declines the
generic names. Edges are deleted before the sweep re-inserts, since
INSERT OR IGNORE against idx_edges_identity would otherwise keep both rows
when a reference rebinds elsewhere.
Replaying real commits of this repo through sync, then diffing against a
rebuild: 16 commits 48 -> 0; 80 commits 1,634 -> 361, with the actively
misleading direction (stale edges the index keeps asserting) 671 -> 2.
Index and sync wall-clock are unchanged; the ORDER BY costs 18% per uncached
name lookup, which never reaches wall-clock because the resolver memoizes it.
The 357-edge residual at 80 commits is one pre-existing class: refs to
generic names (`push`, `join`) parked above #1240's per-name retry ceiling,
which a rebuild resolves into cross-language garbage — a TS test file
"calling" an R method. Converging there would mean manufacturing wrong edges,
so it is left alone. And no drift metric in `codegraph status`: it cannot be
computed without the rebuild it would be recommending, and a proxy would fire
on that residual and train users to ignore it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Deterministic 6-repo table, the three agent A/Bs (django, excalidraw, okhttp,
2 runs/arm, Read 0 in all 12 runs), and an honest read of the two repos that
deliver a few hundred fewer source chars: at the CG-31 tip both were over-filled
by the flat-200 section overhead and paid for it by discarding their epilogue
whole.
Also: CHANGELOG entries for the two user-visible changes, and the memory note
now carries the fourth accounting gap plus the two lessons — hold the REMAINDER
when a full reservation no longer fits, and never skip a file over an accounting
difference.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The invariant this closes: every admitted file receives at least its
reservation before any file draws on carry-forward slack. CG-30 bounded an
oversize cluster member and CG-31 gave the cluster path a displacement guard;
three holes were left, and each one starved a file that had been admitted,
reserved and — in the worst case — rendered.
1. The whole-file arms had no displacement guard. BUY's fit test read
`renderCeiling - totalChars` (everyone's room) while its source-space
sibling refused the same trade, and GRACE was not fit-tested at all.
okhttp's CallServerInterceptor.kt shipped 8,499 chars on a 5,964 funded
ceiling and the rank-6 file below it delivered nothing. Both arms now test
the render they actually produce against `fundedHeadroom`, and a whole
render that does not fit falls through to clustering instead of skipping
the file.
2. Every section was charged a flat 200 chars while a real header runs
300-500. The loop believed it had room it did not have — okhttp allocated
26,601 against a 24,400 ceiling — so the final truncation threw a
fully-rendered section away. Sections are charged their real cost now, the
owed-below arithmetic uses a per-file overhead estimated from the file's own
symbols, and a marginal overrun trims the weakest cluster (or windows the
last one into the room that is left) rather than skipping the file over a
rounding difference.
3. `owedPayableBelow` held all-or-nothing. When the last admitted file's FULL
reservation no longer fit, nothing was held for it: on the precise-query
fixture the rank-5 file took 4,134 chars against a 2,948 reservation while
rank 6 — admitted, reserved 2,539 — was left 4 chars and skipped. It now
holds the remainder while that remainder is still worth a section
(MIN_CHARS).
And the epilogue is budgeted instead of discarded. The flat 600-char margin was
neither the epilogue's size (1,064 gin, 1,788 django, 2,231 excalidraw) nor a
bound on it, so four of six suite repos shipped with no pointer list and no
reminders at all. The loop now reserves the epilogue's FLOOR — the one line
that says an uncovered area exists, plus a pointer for every file whose bytes
were deliberately withheld (CG-12) — and the rest is fitted to the room that
actually remains, in priority order, entry by entry. Sized from the real
strings; no constant was swept against the suite.
Deterministic, same clean-rebuilt indexes, baseline = CG-31 tip:
repo base source new source files ceiling
django 20,791 20,878 6 -> 6 was discarding its epilogue
tokio 21,521 21,607 5 -> 5 was discarding its epilogue
okhttp 19,034 18,870 5 -> 6 +1 file delivered
excalidraw 20,204 19,652 8 -> 8 keeps its pointer list
gin 10,776 10,776 4 -> 4 byte-identical
alamofire 11,662 11,662 2 -> 2 byte-identical
No repo truncates any more and none loses a file. okhttp and excalidraw trade
164 and 552 source chars on their LAST-ranked file for the pointer list naming
what the response could not cover — bytes the CG-31 tip only had because it
over-filled a ceiling it mis-measured and then discarded the epilogue whole.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The '⚠ changed on disk after the last index sync' banner is an honesty claim
about source we DID render — line refs elsewhere in the response may be
shifted — not a note about the response. Drawing the epilogue boundary after
it means the size cut can never be what silences it. Suite numbers unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Deterministic (6 repos, clean rebuilds, both builds): four deliver more source
and one more file each, two are byte-identical, none deliver less. Agent A/B
(django n=3, okhttp n=2, gin n=2, sonnet/effort high, both arms codegraph-on,
0 contamination): the new arm is faster on all three, Read at or below
baseline, occupancy lower.
Also records the two corrections the suite forced on the first cut of the
guard, and the two residuals CG-26 inherits — the render loop's 600-char
epilogue margin (a sweep was run and deliberately NOT shipped) and the BUY
arm's source-space-only guard.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two corrections found by measuring the first cut of the guard against the
6-repo suite. The first version held back the FULL sum of the reservations
below a file. On django that took 2,319 chars off a file the agent receives
and handed them to a section the hard ceiling then threw away — the guard's
own failure mode, one layer down. tokio lost 1,298 the same way.
1. `owedPayableBelow` — hold back only the prefix of what is owed below that
the response can still PAY, in rank order. A promise the ceiling cannot
reach is not a claim on this file's bytes.
2. The final truncation now spends the EPILOGUE before it spends a rendered
file section. It used to cut at the last section header, dropping that
section AND the trailing notes; dropping the notes alone is almost always
enough. A section is source the agent otherwise has to Read; the epilogue
is a pointer list and two reminders, and the note that replaces it carries
the "explore these names" instruction forward.
Also count `flow.text` in `totalChars`. It is prepended to `lines` to make the
final output, so the render loop always spent against a ceiling it was ~2K
under on symbol-bag queries.
Deterministic, same clean-rebuilt indexes, both builds (baseline = CG-30 tip):
repo base source new source files
django 20,033 20,791 5 trunc -> 6
excalidraw 18,776 20,204 7 trunc -> 8
okhttp 15,628 19,034 4 trunc -> 5
tokio 20,340 21,521 4 trunc -> 5
gin 10,776 10,776 4 -> 4 (byte-identical)
alamofire 11,662 11,662 2 -> 2 (byte-identical)
No repo delivers less; four stop truncating. `funded` in the diagnostic now
reports the render CEILING the guard allows, which is what every render path
is actually bounded by.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Carry-forward slack let a file spend what the files ABOVE it left on the
table. Nothing held back what was promised BELOW it. The whole-file BUY arm
has always refused that trade (`owedBelow`); the cluster path read `headroom`
— what is left before the hard ceiling — instead of what is still owed, so
`fileBudget` and `SPINE_CEILING` could pay a 1.5x overshoot out of another
file's reservation.
`fundedHeadroom` is the same inequality in the units the cluster path spends
in: source PLUS the per-section overhead each unreached file will charge.
Floored at the file's own reservation — a kept promise is not a displacement —
and it is <= `headroom` by construction, so it is the only bound the three
render sites need. The skeleton path's `bodyCap` takes it too.
Measured on `__tests__/fixtures/displacement-ts` (a 4-stage pipeline padded
past 500 files, where the 24K envelope genuinely saturates the 24.4K render
ceiling):
before ingest.ts emitted 9,301 on a 6,289 spendable, then lost the whole
section to the final ceiling — 0 delivered. types.ts and sink.ts
skipped `budget-whole-file`. 3 of 6 admitted files delivered.
after ingest.ts bounded to the 4,913 actually free. 6 of 6 delivered,
envelope 14,908 -> 22,066.
The self-query allocation fixture flips back to PASS with it, on a clean full
rebuild of this repo's index (CG-33). Its `afterCG30` verdict blamed an
over-RESERVED incidental file; the reservation was identical in both arms —
the file was over-SPENDING. Recorded honestly in `afterCG31`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Primary evidence is deterministic: on django, query.py rendered 2.12x its budget
on main and 1.49x with the bound, and the freed bytes reach the files below it
(+2,104 chars of source in the same five files). gin is a true control — the two
builds emit byte-identical explore output there, which is what makes its agent-run
deltas variance by construction.
Also records the harness trap that voided the first two batches: ab-new-vs-baseline
swaps src/ to the baseline ref mid-run, so a commit made while it runs captures
baseline sources. Check the "changed:" line before believing any run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
d652c14 was committed while an ab-new-vs-baseline run had the engine checked out
at the BASELINE ref — that harness swaps src/ files mid-run and restores them on
exit — so it captured main's tools.ts and explore-diagnostics.ts and silently
undid 765c06a. Restored from 765c06a, with d652c14's doc tweak re-applied.
The A/B runs started after that commit are void with it (their "changed:" line
lists only explore-diagnostics.ts, i.e. both arms ran the same retrieval code)
and are re-run rather than reported.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The two kernel port checklists record the extractor configs as surveyed
at porting time; their structTypes lines are marked superseded rather
than rewritten, so the surveys stay readable as history.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four cases in extraction.test.ts: a Rust union carrying an
`impl Trait for` edge and owning the impl's method; a named C union
alongside a forward declaration that must NOT mint a node; a
`typedef union` taking the typedef name with no `<anonymous>` twin; a
C++ union with a member function.
Verified they fail without the fix on BOTH extraction paths — the wasm
walker via CODEGRAPH_KERNEL=0 and the kernel with a staged build.
torture.c / torture.rs gain the same shapes. The parity gate compares
the two walkers rather than a snapshot, so the fixtures do not detect
the bug on their own — they pin that the fix stays SYMMETRIC. The
regression tests above are what pin that it is present.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The self-query allocation probe fixture's delivered-share gates now fail. The
cause is not the new bound: allocation is unchanged between arms (parse-run.mjs
32.3% vs 33.9% on main) and tools.ts delivers the same 8,282 chars in both. What
changed is that the incidental file now DELIVERS — on main its whole section was
cut by the hard-ceiling truncation, so the fixture passed on truncation luck.
Every file on this repo obeys the new bound (max 1.40x of spendable).
Recorded as `afterCG30` with that reasoning rather than tuning the bound to
restore the pass. The over-reservation it exposes is epic CG-24's subject.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A `union` declaration produced no symbol at all in any of the four
languages that have one. The type never entered the graph, and neither
did anything attached to it — in Rust every `impl Trait for MyUnion`
lost its edge, and the impl's methods were left with a qualifiedName
pointing at a type the graph did not contain.
`union_specifier` / `union_item` were absent from the extraction layer
entirely: no `<x>Types` list on the TS side, no dispatch branch in
either kernel walker.
They join `structTypes` (kind `struct` — NodeKind has no `union`), which
is the extension point the table-driven extractors already provide. The
body guard in extractStruct is untouched, so a bodiless `union U;` stays
a forward declaration and is still skipped, exactly like `struct U;`.
`resolveTypeAliasKind` accepts `union_specifier` too, so
`typedef union { … } N;` takes the typedef's name the way
`typedef struct { … } N;` already did. Without it the anonymous union
body would mint a second `<anonymous>` node beside the alias.
Both walkers change together so kernel<->wasm parity holds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
shrinkCluster keeps an oversize cluster's highest-importance member WHOLE on
purpose — an empty file section sends the agent to Read, the outcome explore
exists to prevent. What it lacked was a bound, and "never empty" quietly meant
"never bounded": on the reporting repo one file emitted 22,376 chars against a
9,181-char reservation (2.44x), past both the per-file budget and the spine
ceiling. That overshoot is what collapses `headroom` for every file below it.
The same rule has a second face. When the top member is bigger than the whole
response ceiling, the file does not overshoot — it is dropped entirely at the
renderCeiling check, so the agent gets nothing for a file it named.
renderCluster now takes a ceiling (1.5x what the file may spend — the same
multiple SPINE_CEILING already draws, and never below the cap, so a cluster
that fits is untouched). Past it the member is WINDOWED on whole lines rather
than emitted whole or dropped: leading window plus, on a flow cluster, a window
on the spine's call site. A partial window shorter than 12 lines is dropped
instead — a sliver in the session record forces the next call's dedup to shred
the block around it or re-send it — unless nothing else was emitted, where the
never-empty floor wins.
Measured on the new fixture, pre-fix vs post-fix:
monthly.ts 12,391 chars on a 3,334 budget (3.7x) → 4,941 (1.48x)
quarterly.ts dropped, no headroom left → 4,004 delivered
Also: the diagnostic now reports `spendable` (reservation + inherited slack)
alongside `reserved`. Every render bound reads the former, so reporting only
the latter makes an ordinary carry-forward read as a file spending over budget
— and it made the overshoot this issue is about unmeasurable. A windowed file
is now flagged `clipped` too, instead of presenting a window as the whole file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A live, auto-sync-maintained index does not converge to a clean full
rebuild of the identical tree. On codegraph's own repo, 4.3% of distinct
edges are wrong in both directions (751 missing, 476 stale), dominated by
`calls` — the edges flow queries traverse and that feed the RWR mass
explore ranks files by.
Raw edge rows differ by only +0.7%, because the divergence is
bidirectional and nets out; any drift check must compare edge SETS.
Rebuild-vs-rebuild is 0, so the indexer is deterministic and this is not
noise. Node sets are identical and every integrity check is 0 on both
indexes, so this is stale cross-file resolution, not accumulated residue.
`diff-index-drift.mjs` is read-only and takes two index paths — rebuilding
is the caller's job, so the tool can never clobber the artifact it is
measuring. It also refuses a missing path, since node:sqlite creates an
empty database rather than failing and an empty schema reads exactly like
a stale pre-migration index.
Diagnostic captures from the originating incident are deliberately NOT
committed: they contain verbatim source from a private repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Swept every benchmark doc for figures produced off `result.usage` and fixed
the ones that had raw logs to re-derive from.
residual-context-occupancy.md — the sonnet 3-turn throughput table. Re-derived
from the preserved logs: tokens saved 23% -> 56%, and vscode's "98% MORE tokens
with codegraph" was never real, it is 41% fewer. Cost, time and tool calls were
never affected by this field and are unchanged. The occupancy table itself is
measured off the timeline, so every number in it stands -- including the 82%
higher residual, which is the finding the document exists for.
call-sequence-analysis.md — this doc DIAGNOSED the bug and its reproduce block
claimed the aggregator summed per-turn tokens. It did not, until 04c0f8e. Noted,
with the three wrong results the gap produced: the excalidraw cut recorded here,
the sonnet campaign, and the Opus re-measure that invented a token regression.
answer-directly-vs-explore-agent.md — build 0.9.4, 2026-05-24, raw logs gone.
Cannot be re-derived, so flagged rather than silently left or invented: its
token figure is indicative, its turn/read/context findings do not depend on the
broken field and stand.
The remaining benchmark docs (allocation-ab-1500, dedup-cg20, allocation-
efficiency, feedback-metrics) carry no throughput tables — checked, clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Re-measured on the CLI-blocked harness at the README's own stated methodology
(claude-opus-4-8, single question, median of 4, same 7 repos), with tokens
summed per turn per 04c0f8e.
What moved, and what did not:
tool calls 89% -> 88% fewer holds
file reads 0 on all seven holds
tokens 69% -> 62% fewer holds
time 20% -> 53% faster was badly understated
cost 60% -> 44% cheaper genuinely lower
Time was the most wrong figure in the old table, in our favour: every repo is
faster with the graph, including the two that previously carried a footnote
excusing them for being slower. That footnote is gone. Cost is the one real
downgrade, and it varies with how much discovery a question demands rather
than with repo size -- 57-78% where the file-reading arm needed 28-43 tool
calls, near-even on Gin at 7. Said plainly instead of averaged away.
Methodology now records that the CLI is blocked in BOTH arms and why: on an
unblocked harness the control reached codegraph through Bash in 26 of 28 runs,
so it was not a control. All 28 attempted it here; all 28 were blocked.
The context-footprint note from a7db24d is unaffected -- occupancy is measured
off the timeline, not the token field that was wrong.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
"Tokens processed" was read off result.usage. That was correct when the README
figures were measured; in current Claude Code the field reports the LAST turn
only. Nothing here changed — the host did, silently — and the harness kept
reporting the smaller number.
The error is one-sided, which makes it worse than noise: it under-counts
whichever arm takes more turns, and that is always the WITHOUT arm. On the
2026-08-05 campaign it turned a real 62% token saving into 19% and invented a
token REGRESSION on tokio (-41%) and alamofire (-25%) that does not exist. Those
numbers were one push away from the README.
Now summed per assistant request and deduped by message.id, the same rule the
occupancy timeline already used — Claude Code emits one event per content block
carrying identical usage, so summing per event double-counts (~1.7x measured).
CLAUDE.md already warned about this field. The code did not follow; it does now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>