feat(extraction): C deferral round 2 — 8 new preParse passes, linux kernel/+mm/ deferral 58.6%→33.9% (#1353)
Census-driven cut of the top-ranked post-R7a lever. All passes TS-side, C-only (preParseCSource), shared by both arms: - parameterized-annotation whole-blank (__free/__printf/__counted_by/ __bpf_md_ptr…; extends through a stranded field `;`) - type-keyword-arg scanner (kzalloc_obj(struct T), list_entry, multi-line continuations behind nested-paren args; bounded hand scanner, head exclusions + call-vs-declaration guard; blanks trailing stars) - static/extern CAPS-macro declaration lines at any scope; the initialized form is REWRITTEN to its expansion (name/tail keep exact offsets) - va_arg qualified-type blank; GNU named-variadic #define dots-only blank (post-restore); sandwiched notrace-family; C23 auto; multi-line iterator-macro spans (hlist_for_each_entry_rcu + lockdep arg) - word list += cacheline family (2- and 4-underscore spellings) + 10 more census-confirmed annotations Gates: five-repo parity sweeps 0 diffs (git deferral 16.1→12.2%, redis 25.3→24.1%, fmt/protobuf unchanged); linux full-tree both arms 2,049,153 nodes / 6,413,518 edges (+858/+6,585 vs R7a) with byte-identical dumps (10,446,478 lines, sha256 6dd1185b); kernel-arm parse-loop 356→306s at 2c; suite 2517 green under CODEGRAPH_KERNEL_EXPECT=1. Honesty note recorded in the docs: error recovery was already salvaging most SYMBOLS on deferred files — the graph win is relationships + phantom cleanup, and the unreleased CHANGELOG entry was rewritten off the sweep-subset framing. Also records §7a.5: post-R7a 8-core cg1212 re-run 16.4min (was 18.3min); 8c parse sits on the single-writer floor, so the <10min-on-8c gap re-ranks to the per-ref resolution path. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
2d72891b59
commit
b9d0f57a64
@@ -66,12 +66,29 @@ them are the ORIGINAL plan and carry expectations that measurement later correct
|
||||
envelope now **19.1min on a substantially RICHER graph** (the new
|
||||
preParse blanks recover previously-error-swallowed code; wasm-arm on
|
||||
the same graph is 22.9min — the 17.6 record was the old smaller graph
|
||||
and isn't directly comparable). The <10min-on-8c target remains open;
|
||||
levers left, ranked: **C/C++ deferral cuts (58% of linux files still
|
||||
defer to wasm — each recovered idiom moves parse toward the native
|
||||
floor)** > backpressure ~120s (checkpoint I/O floor) > E-scan/settle/
|
||||
read-mapping (~70–90s each, approaching honest work) > the 8c re-run
|
||||
formality.
|
||||
and isn't directly comparable). 8c re-run DONE post-R7a (§7a.5):
|
||||
**16.4min** (pre-R7a record 18.3min), EXIT 0, counts == both 2c arms,
|
||||
WAL 1.34GB. The <10min-on-8c target remains open, and the re-run
|
||||
re-ranked the levers honestly: at 8c the parse-loop (202.6s) is
|
||||
already AT the single-writer floor, so **the target gap is ~entirely
|
||||
the core-invariant resolution superphase (715s ≈ 12 of the 16.4min)
|
||||
— the per-ref path is THE 8c lever**. C/C++ deferral round 2 DONE
|
||||
2026-07-18 (full record: checklist doc): eight new C-only preParse
|
||||
passes + word-list extensions took kernel/+mm/ deferral
|
||||
**58.6% → 33.9%** (git 16.1 → 12.2%, redis 25.3 → 24.1%,
|
||||
fmt/protobuf unchanged — cpp-dominant, correct no-op), five-repo
|
||||
sweeps 0-diff, linux full-tree both arms **2,049,153 / 6,413,518**
|
||||
with **byte-identical dumps** (10,446,478 lines, sha `6dd1185b…`);
|
||||
kernel-arm parse-loop **356 → 306s** at 2c, envelope ~17.1min
|
||||
(host-contaminated, indicative). Honesty note: full-graph node
|
||||
deltas are small (+858) — wasm error recovery was already salvaging
|
||||
most SYMBOLS on deferred files; the real win is EDGES (+6,585),
|
||||
phantom cleanup, and native-path coverage. Remaining deferral is
|
||||
policy-skips (CONFIG interleaves, TP_PROTO DSL, module_init-no-semi)
|
||||
+ small buckets — this lever is largely SPENT. Queue now: **per-ref
|
||||
resolution path** (the core-invariant superphase) > backpressure
|
||||
~120s (checkpoint I/O floor) > E-scan/settle/read-mapping (~70–90s
|
||||
each, approaching honest work).
|
||||
- [x] **R7a. C/C++ port** — DONE 2026-07-17, same-day walker+gates after the
|
||||
survey (#1344) and grammar vendoring (#1345). One dual-language walker
|
||||
(`codegraph-kernel/src/ccpp/`), preParse HOISTED to the route point
|
||||
@@ -741,6 +758,26 @@ algorithmic wins only.
|
||||
rock) > backpressure ~120s (checkpoint I/O floor) > E-scan 69–93s (approaching
|
||||
honest regex work over 1.5GB) > settle 88s > read-mapping 57s.
|
||||
|
||||
#### 7a.5 8-core re-run, post-R7a (2026-07-17) — 16.4min; the 8c gap is now all resolution
|
||||
|
||||
Same provisioning as the §7a.2 retry (cg1212 at cpuset 0-7 / 7GB), the deployed
|
||||
R7a build, fresh init of the v7.2-rc2 tree: **EXIT 0, envelope 981s = 16.4min**
|
||||
(pre-R7a 8c record: 18.3min — and that was the smaller pre-blank graph).
|
||||
Counts **2,048,295 / 6,406,933 == both 2c arms**; WAL peak 1.34GB (same
|
||||
contained regime as 1.09–1.57GB records). Phases: parse-loop **202.6s**
|
||||
(pre-R7a all-wasm 8c: 208.7s — both sit ON the single-writer store floor, so
|
||||
8c parse is writer-bound, not extraction-bound) · resolution superphase
|
||||
**715.0s** (was 835.9s) containing callback-synthesis **257.4s** (was 338.7s)
|
||||
and edge-index-recreate 52.0s · maintenance 47.6s. The −1.9min vs the record
|
||||
is the post-#1336 rounds (#1339 countGuard, #1341 cFnPtr, R7a native parse +
|
||||
defer-reuse) landing at 8c for the first time.
|
||||
|
||||
**Consequence for the <10min target:** ~12 of the 16.4 minutes are the
|
||||
core-invariant resolution superphase. Deferral cuts can't materially move the
|
||||
8c envelope (parse is already at the writer lane); they remain queued for
|
||||
graph richness + the 2c/low-core envelope. The 8c target now lives or dies on
|
||||
the per-ref resolution path (§7a.2's lever (a)).
|
||||
|
||||
### 7b. Arc 3 — graph richness (forensics-backed; adopt cbm's real extras, skip inflation)
|
||||
Priority order, each gated by the standard A/B + node-explosion probes:
|
||||
1. **Test→subject edges** (first-class `tests` edges at index time; we compute covering
|
||||
|
||||
Reference in New Issue
Block a user