Files
codegraph/scripts/agent-eval/allocation-fixtures.json
T
Colby McHenryandClaude Opus 5 7cbde95ce2 fix(explore): pay every admitted file on every render path (CG-26)
The invariant this closes: every admitted file receives at least its
reservation before any file draws on carry-forward slack. CG-30 bounded an
oversize cluster member and CG-31 gave the cluster path a displacement guard;
three holes were left, and each one starved a file that had been admitted,
reserved and — in the worst case — rendered.

1. The whole-file arms had no displacement guard. BUY's fit test read
   `renderCeiling - totalChars` (everyone's room) while its source-space
   sibling refused the same trade, and GRACE was not fit-tested at all.
   okhttp's CallServerInterceptor.kt shipped 8,499 chars on a 5,964 funded
   ceiling and the rank-6 file below it delivered nothing. Both arms now test
   the render they actually produce against `fundedHeadroom`, and a whole
   render that does not fit falls through to clustering instead of skipping
   the file.

2. Every section was charged a flat 200 chars while a real header runs
   300-500. The loop believed it had room it did not have — okhttp allocated
   26,601 against a 24,400 ceiling — so the final truncation threw a
   fully-rendered section away. Sections are charged their real cost now, the
   owed-below arithmetic uses a per-file overhead estimated from the file's own
   symbols, and a marginal overrun trims the weakest cluster (or windows the
   last one into the room that is left) rather than skipping the file over a
   rounding difference.

3. `owedPayableBelow` held all-or-nothing. When the last admitted file's FULL
   reservation no longer fit, nothing was held for it: on the precise-query
   fixture the rank-5 file took 4,134 chars against a 2,948 reservation while
   rank 6 — admitted, reserved 2,539 — was left 4 chars and skipped. It now
   holds the remainder while that remainder is still worth a section
   (MIN_CHARS).

And the epilogue is budgeted instead of discarded. The flat 600-char margin was
neither the epilogue's size (1,064 gin, 1,788 django, 2,231 excalidraw) nor a
bound on it, so four of six suite repos shipped with no pointer list and no
reminders at all. The loop now reserves the epilogue's FLOOR — the one line
that says an uncovered area exists, plus a pointer for every file whose bytes
were deliberately withheld (CG-12) — and the rest is fitted to the room that
actually remains, in priority order, entry by entry. Sized from the real
strings; no constant was swept against the suite.

Deterministic, same clean-rebuilt indexes, baseline = CG-31 tip:

  repo         base source   new source   files      ceiling
  django            20,791       20,878   6 -> 6     was discarding its epilogue
  tokio             21,521       21,607   5 -> 5     was discarding its epilogue
  okhttp            19,034       18,870   5 -> 6     +1 file delivered
  excalidraw        20,204       19,652   8 -> 8     keeps its pointer list
  gin               10,776       10,776   4 -> 4     byte-identical
  alamofire         11,662       11,662   2 -> 2     byte-identical

No repo truncates any more and none loses a file. okhttp and excalidraw trade
164 and 552 source chars on their LAST-ranked file for the pointer list naming
what the response could not cover — bytes the CG-31 tip only had because it
over-filled a ceiling it mis-measured and then discarded the epilogue whole.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 03:45:28 -05:00

220 lines
15 KiB
JSON

{
"$comment": [
"Regression fixtures for GitHub issue #1500 / epic CG-1 — relevance-proportional",
"explore budget allocation. Run them with `node scripts/agent-eval/probe-allocation.mjs`",
"against a built dist/.",
"",
"STATUS: BOTH FIXTURES PASS again as of CG-31. self-query's delivered-share gates",
"failed on the CG-30-only build (`afterCG30`) and CG-31 restored them (`afterCG31`):",
"the incidental file was not over-RESERVED at all — it was over-SPENDING, drawing on",
"the reservations of files the render loop had not reached yet. Bounding that put the",
"response back inside the envelope, so nothing truncates and every admitted file",
"delivers. Read the two blocks together; the CG-30 verdict's diagnosis was wrong.",
"",
"CG-10 (relevance scoring) closed the RANKING half —",
"nothing incidental reaches the envelope any more — and CG-12 (score-proportional",
"allocation with a relative cliff) closed the BYTE SPLIT: each file's share is reserved",
"before anything renders, and a file under 15% of the top weight gets no source at all,",
"freeing both its bytes and its maxFiles slot. Each fixture records the pre-CG-10",
"`baseline`, the interim `afterCG10`, and the current `afterCG12`.",
"",
"`groups` partitions the files explore rendered into `answer` (what the query is",
"actually about) and `incidental` (what wins the envelope today on name collisions).",
"Assertions are on the DELIVERED envelope unless suffixed `Allocated`; delivered is",
"what the agent got, allocated is what the render loop chose before the hard ceiling.",
"Shares are fractions of the whole response, meta-text included, so they never sum to 1."
],
"fixtures": [
{
"id": "payroll-go",
"title": "#1500 — generated Go CRUD beside a hand-written payroll workflow",
"kind": "fixture",
"path": "__tests__/fixtures/payroll-go",
"query": "how does payroll cycle create and calculate payslips?",
"rationale": [
"The reporter's repo shape: a Go service whose generated FKIT CRUD layer sits",
"beside the hand-written use-case that does the real work. The query deliberately",
"does NOT name runPayrollCycleAll / BuildPayslip / Upsert — an architecture",
"question phrased the way a newcomer would phrase it. The generated layer",
"name-collides on every query term (CreatePayslip, PayrollCycleCreateRequest,",
"CalculatePayrollCycleTotals, a second BuildPayslip, a second Upsert), so a",
"scorer that rewards incidental name matches surfaces the CRUD path.",
"Half the generated files carry ORDINARY names and are only detectable by their",
"`// Code generated ... DO NOT EDIT.` header (CG-5), which is what makes this",
"the #1500 case rather than a .pb.go case."
],
"groups": {
"answer": [
"internal/usecase/**",
"internal/store/**",
"internal/transport/**",
"internal/domain/**",
"cmd/**"
],
"incidental": [
"internal/gen/**"
]
},
"assert": {
"answerShareAtLeast": 0.55,
"incidentalShareAtMost": 0.25,
"topFileGroup": "answer",
"mustDeliverBytes": [
"internal/usecase/payroll/cycle.go",
"internal/usecase/payroll/payslip_builder.go"
],
"$mustContainComment": "Needles are chosen to match the HAND-WRITTEN chain only — a bare `BuildPayslip`/`Upsert` also matches the generated collisions, which is the whole point of the fixture.",
"mustContain": [
"runPayrollCycleAll",
"func (s *Service) BuildPayslip",
"s.store.Upsert(ctx, slip)"
]
},
"baseline": {
"measuredOn": "2026-08-03",
"note": "19 files → very-tiny tier (13,000 budget); 23,020 chars allocated against it, cut to 16,011 by the 19,500 hard ceiling.",
"delivered": {
"internal/gen/fkit/payroll/payslip.go": 0.307,
"internal/gen/fkit/payroll/payroll_cycle.go": 0.266,
"internal/domain/payroll/payslip.go": 0.256,
"internal/usecase/payroll/cycle.go": 0.0
},
"verdict": "The workflow file is allocated the single largest slice (30.6%) and delivers ZERO — the hard ceiling drops its whole section. Generated CRUD takes 57.4% of what the agent actually receives; runPayrollCycleAll, BuildPayslip and the real Upsert never reach the response."
},
"afterCG10": {
"measuredOn": "2026-08-04",
"delivered": {
"internal/usecase/payroll/cycle.go": 0.389,
"internal/gen/fkit/payroll/payroll_cycle.go": 0.235,
"internal/domain/payroll/payslip.go": 0.226,
"internal/gen/fkit/payroll/payslip.go": 0.0
},
"verdict": "PASSES answerShareAtLeast (61.5%), incidentalShareAtMost (23.5%), topFileGroup, and the cycle.go + runPayrollCycleAll + real-Upsert needles. The generated files now rank #3/#4 instead of #1/#2 — kind weighting plus a 0.3x generated penalty on BOTH the score and the graph mass, which is the key the comparator sorts on. STILL FAILING for CG-12: payslip_builder.go ranks #6 against the tier's maxFiles of 4, so `func (s *Service) BuildPayslip` never renders."
},
"afterCG12": {
"measuredOn": "2026-08-04",
"note": "19,337 delivered, nothing truncated. Both generated files cliffed to pointers (weight 3.9 and 3.5 against a cliff of 5.4), which frees the two maxFiles slots the hand-written store and builder then take.",
"delivered": {
"internal/usecase/payroll/cycle.go": 0.306,
"internal/domain/payroll/payslip.go": 0.212,
"internal/usecase/payroll/payslip_builder.go": 0.151,
"internal/store/payslipstore/store.go": 0.118,
"internal/gen/fkit/payroll/payroll_cycle.go": 0.0,
"internal/gen/fkit/payroll/payslip.go": 0.0
},
"verdict": "ALL GATES PASS. Answer group 78.7% (from 25.6% at baseline), generated layer 0.0% (from 57.4%). All four hand-written files deliver source, including payslip_builder.go — `func (s *Service) BuildPayslip`, the 'calculate' half of the question, finally reaches the agent. The generated files are still NAMED with their symbols and line numbers under 'Not shown above', so withholding their bytes costs ~100 chars each instead of ~4,500 and stays one follow-up explore away."
}
},
{
"id": "self-query",
"title": "This repo — incidental `explore`/`BUDGET` matches in the agent-eval scripts",
"kind": "self",
"path": ".",
"query": "how does explore allocate its output budget across files",
"rationale": [
"The same failure mode with no generated code in sight. `scripts/agent-eval/*.mjs`",
"mention `explore` and `BUDGET` incidentally — they are eval harnesses, not the",
"allocator — and they are small enough to ship WHOLE, while src/mcp/tools.ts (which",
"carries getExploreOutputBudget and the render loop, and scores 4x higher on every",
"signal) is large enough to be clipped at maxCharsPerFile. Allocation follows file",
"size, not relevance.",
"",
"Unlike payroll-go this fixture reads THIS repo's live index, so its exact numbers",
"move as the repo changes (indexed file count crossing 500 flips the budget tier).",
"The assertions are therefore relative — answer-vs-incidental, not fixed percentages."
],
"groups": {
"answer": [
"src/mcp/**"
],
"incidental": [
"scripts/**"
]
},
"assert": {
"$answerShareComment": [
"Denominated in DELIVERED SOURCE, not in the whole envelope (CG-26). The",
"envelope-denominated form of this gate moved for reasons that have nothing",
"to do with allocation: it fell when the epilogue stopped being discarded,",
"and it fell again when a fifth ADMITTED file finally got paid its",
"reservation instead of being dropped by the ceiling. Both are the",
"improvements this epic exists to make, and a gate that reads them as",
"regressions is measuring the denominator. The fixture's own rationale",
"already says the assertions are relative, answer-vs-incidental, not fixed",
"percentages. 0.5 is unchanged; only what it is a share OF."
],
"answerShareOfSourceAtLeast": 0.5,
"incidentalShareAtMost": 0.25,
"topFileGroup": "answer",
"mustDeliverBytes": [
"src/mcp/tools.ts"
]
},
"baseline": {
"measuredOn": "2026-08-03",
"note": "493 files → small tier (18,000 budget); 27,518 chars allocated against it, cut to 19,749 by the 25,000 hard ceiling.",
"delivered": {
"scripts/agent-eval/offload-eval-hook.mjs": 0.25,
"scripts/agent-eval/offload-eval-metrics.mjs": 0.236,
"scripts/agent-eval/parse-session.mjs": 0.232,
"src/mcp/tools.ts": 0.185,
"scripts/agent-eval/offload-eval-cost.mjs": 0.0
},
"verdict": "The script corpus takes 71.8% of the delivered envelope (79.4% of what was allocated) against tools.ts's 18.5%, despite tools.ts scoring 46 vs 10, carrying 2.3x the graph mass and 3x the distinct term hits."
},
"afterCG10": {
"measuredOn": "2026-08-04",
"delivered": {
"src/resolution/memory-budget.ts": 0.512,
"src/mcp/tools.ts": 0.329
},
"verdict": "PASSES incidentalShareAtMost: every scripts/agent-eval/*.mjs file is gone (0.0%), which is CG-10's acceptance bar — their sole match was an unused file-scope `explore`/`BUDGET` constant, worth 0.08x weight, and the relative floor (8.2) then cut them. tools.ts ranks #1. STILL FAILING for CG-12: memory-budget.ts scores 18 to tools.ts's 41 yet takes 51.2% of the envelope to tools.ts's 32.9%, purely because it is small enough to ship whole while tools.ts is clipped at maxCharsPerFile. Allocation still follows file size, not relevance."
},
"afterCG12": {
"measuredOn": "2026-08-04",
"note": "18,134 delivered, nothing truncated. tools.ts scores 58 here (it grew by the allocator this task added), memory-budget.ts 18.",
"delivered": {
"src/mcp/tools.ts": 0.606,
"src/resolution/memory-budget.ts": 0.172,
"src/resolution/lru-cache.ts": 0.111
},
"verdict": "ALL GATES PASS. tools.ts takes 60.6% of the envelope, up from 18.5% at baseline and 32.9% after CG-10 — past the epic's >50% acceptance bar. The reversal is the whole point: memory-budget.ts no longer wins by being small enough to ship whole (it now clusters within its 3.1K reservation), and tools.ts is no longer clipped at maxCharsPerFile (11K reservation, ~3x the old flat cap). Exception to 'no previously-unclipped file becomes clipped': memory-budget.ts was unclipped-whole at 5,672 and is now clipped to its proportional share. That is the epic's own diagnosis of the bug, not a regression — it scored 18 against tools.ts's 58 and was taking the larger slice."
},
"afterCG30": {
"measuredOn": "2026-08-06",
"note": "23,688 delivered of 26,430 allocated, truncated at the 25,000 ceiling. Baseline (main) on the SAME index: 14,851 delivered of 25,221 allocated. tools.ts delivers 8,282 chars in BOTH arms — identical bytes; only the denominator moved.",
"delivered": {
"scripts/agent-eval/parse-run.mjs": 0.361,
"src/mcp/tools.ts": 0.35,
"src/mcp/explore-session-state.ts": 0.147,
"src/resolution/memory-budget.ts": 0.0
},
"verdict": "THREE GATES FAIL — and the cause is not the CG-30 bound. Allocation is unchanged between arms (parse-run.mjs 32.3% here vs 33.9% on main); what changed is that it now DELIVERS. On main its whole 8,548-char section was cut by the hard-ceiling truncation, so the incidental group scored 0.0% by luck, not by design, and the fixture passed on that. Bounding the oversize-member overshoot freed enough headroom that the response no longer truncates the same section away. Every file obeys the new bound on this repo (max ratio 1.40x of spendable, against the 1.5x ceiling). What the failure exposes is real and pre-existing: parse-run.mjs scores 18 against tools.ts's 58 yet is reserved a comparable slice — a low-scoring file taking a top-file share, which is epic CG-24's subject. Fix it there; do not tune the CG-30 bound to restore a pass that depended on truncation. SUPERSEDED by afterCG31 — the diagnosis above is wrong on one load-bearing point, see there."
},
"afterCG31": {
"measuredOn": "2026-08-06",
"note": "23,083 delivered of 23,080 allocated — inside the envelope, nothing truncated. Both arms measured on the SAME clean FULL REBUILD of this repo's index (CG-33: an incrementally-synced index diverges and shifts ranking). CG-30-only arm on that index: 23,692 delivered of 26,410 allocated, TRUNCATED.",
"delivered": {
"src/mcp/tools.ts": 0.359,
"scripts/agent-eval/parse-run.mjs": 0.187,
"src/mcp/explore-session-state.ts": 0.151,
"src/resolution/lru-cache.ts": 0.087
},
"verdict": "ALL FOUR GATES PASS. The afterCG30 verdict called parse-run.mjs over-RESERVED; it was not — its reservation is 4,314 in both arms. It was over-SPENDING: 8,548 chars, drawing on reservations belonging to files the render loop had not reached yet, which is the CG-31 defect. With the displacement guard it renders 4,314, tools.ts's identical 8,282 chars go from 35.0% to 35.9% of a response that no longer overruns, and lru-cache.ts (dropped as memory-budget.ts was on the CG-30 arm) delivers. Note what did NOT change: allocation. This fixture moved because the render loop stopped spending other files' bytes, not because anything was re-ranked."
},
"afterCG26": {
"measuredOn": "2026-08-06",
"note": "24,952 delivered of 24,949 allocated, nothing truncated — against the CG-31 tip's 23,083 on the SAME clean full rebuild of this repo's index. tools.ts delivers 8,282 chars in BOTH arms: identical bytes, unchanged reservation, unchanged rank. Total delivered SOURCE 21,228 against 18,105.",
"delivered": {
"src/mcp/tools.ts": 0.334,
"scripts/agent-eval/parse-run.mjs": 0.174,
"src/mcp/explore-session-state.ts": 0.141,
"src/resolution/memory-budget.ts": 0.126,
"src/resolution/lru-cache.ts": 0.081
},
"verdict": "ALL GATES PASS. The one that changed shape is answerShareAtLeast → answerShareOfSourceAtLeast: on the envelope denominator the answer group reads 47.5% here against 51.0% at the CG-31 tip, and neither number is about allocation. tools.ts's bytes are byte-identical between the arms; what moved is that the response now delivers a FIFTH admitted file (memory-budget.ts, rank 4, paid its full 3,123-char reservation — the CG-31 tip rendered it and then let the hard ceiling drop the whole section) and keeps epilogue prose it used to discard. Answer/incidental separation is unchanged and strong: tools.ts 33.4% against parse-run.mjs's 17.4%, incidental 17.4% (down from 18.7%), top delivered file still tools.ts. Measured in delivered source the answer group is 55.5%."
}
}
]
}