fix(scale): kernel-scale hardening — OOM-safe pass skipping + watchdog-safe index recreate (#1323)

Two hazards found by running today's full stack against the Linux kernel
(70,129 files) in the cg1212 repro container:

1. The parallel-synthesis fallback retried a worker-failed pass on the MAIN
   thread. At multi-million-node scale a worker failure is usually a memory
   ceiling, so the retry would OOM the process and take the whole index with
   it. Above 1.5M nodes a failed pass is now skipped with a clear stderr
   message (its synthesized edges are absent; the index completes). Below
   that, the main-thread retry stays — small-scale worker crashes are
   transient and the retry keeps coverage.

2. endBulkEdgeLoad rebuilt all four edge indexes in one synchronous span —
   measured 79s at kernel scale, past the #850 liveness watchdog's 60s
   stall window. A daemon-triggered re-index would have been SIGKILLed right
   after doing the work. Now async with an event-loop yield between builds,
   keeping each stall to a single index (~20s at kernel scale).

Validation: full Linux kernel index to completion in the repro container —
2,048,674 nodes / 6,405,964 edges, EXIT 0, zero passes skipped, on a 2-CPU
VM (worst case: pool disabled, sequential resolution + synthesis) in ~27min.
Phase walls: parse 6.0m, resolution 19.5m (incl. synthesis 6.3m, recreate
79s), maintenance 74s. Suite green (2444).

Also adds docs/design/native-extraction-kernel.md — the spike-validated
design for the native extraction kernel (Rust parse+walk over dubbo's Java:
202ms rayon / 1.07s single-thread vs 4.7s for the current wasm pipeline).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Colby Mchenry
2026-07-16 19:09:01 -05:00
committed by GitHub
co-authored by Claude Fable 5
parent 567b4ad4be
commit 4efc6c70e2
5 changed files with 119 additions and 4 deletions
+87
View File
@@ -0,0 +1,87 @@
# Native extraction kernel — design + spike results
**Status:** spike validated 2026-07-16; project approved, not yet started.
**Owner context:** the last structural lever for fresh-index wall clock after the
2026-07-16 arc (#1305, #1320, #1321, #1322) exhausted Node-side scheduling.
## Why
Post-arc, the fresh-index profile on dubbo (4,402 Java files) is:
parse-loop ~4.7s, resolution ~5.5s (persist-bound), synthesis ~0.9s, total ~11.1s
vs codebase-memory-mcp v0.9.0 at 7.1s. Two levers were measured dead:
- **RAM-backed DB** (parse-loop 6.9s on a ramdisk vs 4.64.8s on SSD, n=2
interleaved): fast-init `synchronous=OFF` already writes at page-cache speed.
The parse phase is CPU-bound.
- **TreeCursor rewrite** (earlier arc): web-tree-sitter's traversal is not the
cost; the floor is per-node JS↔WASM **marshaling** — every `node.kind`,
`.childForFieldName`, `.text` crosses the boundary.
The only remaining parse lever is doing the walk on the native side and
crossing the boundary **once per file** instead of once per node.
## Spike (2026-07-16)
Minimal Rust binary (`tree-sitter` 0.25 + `tree-sitter-java`, TreeCursor walk
touching every node's kind/range + `name`-field text, emitting flat
`(kind_id, start, end, name_len)` rows — the extraction access pattern).
Dubbo's 4,048 `.java` files, 17MB, 3.59M AST nodes, Apple M3 Pro:
| | wall |
|---|---|
| Current pipeline parse-loop (7 wasm workers, incl. extraction + store dispatch) | 4,700ms |
| Rust parse+walk, rayon | **202ms** |
| Rust parse+walk, single thread | 1,067ms |
One native thread beats the whole 7-worker wasm pool 4.4×; at equal
parallelism the walk is ~14× faster. Even charging the kernel for the
extraction logic it must still perform, parse-loop 4.7s → ~1.01.5s is
realistic, putting dubbo ≈ 7.58s total (parity with cbm).
## Architecture
- **Crate:** `codegraph-kernel`, napi-rs, links tree-sitter's C library and
vendored grammars natively. Input: `(filePath, content, language)`. Output:
flat typed buffers (nodes, edges, unresolved refs) — one boundary crossing
per file.
- **Per-language logic:** migrate extractors to tree-sitter **query files**
(`.scm`) executed by a generic Rust emitter; bespoke TS logic that queries
can't express (macro salvage, dialect sniffing, content-gated `.h`
detection) stays as TS pre/post passes over the returned buffers.
- **Distribution:** prebuilt `.node` per platform through the existing
release-bundle pipeline (same per-platform packages as the Node runtime).
The wasm path remains as the universal fallback — same crate compiled to
wasm keeps one implementation.
- **Rollout:** per-language, funnel languages first (TS/JS → Java → Python →
Go). A language ships only when its equivalence gate passes.
## Equivalence gate (per language)
Byte-identity against hand-written extractors is NOT expected (bespoke logic
ports approximately). The gate is:
1. Node/edge/ref **counts** within ±0.5% on 3 real repos (small/medium/large),
with every diff category eyeballed.
2. The retrieval invariants hold: explore-flow connects the language's
canonical flows end-to-end (`docs/design/dynamic-dispatch-coverage-playbook.md`),
agent A/B shows no regression per the standard methodology.
3. Fresh-index wall improves on the language's repos; no regression on a
control repo of a non-migrated language.
## Non-goals
- Porting resolution, synthesis, frameworks, MCP, or the installer — they are
pool-parallel and not marshal-bound. The measured native advantage there is
~1.4× CPU, not worth the correctness moat (2,444 tests, byte-identical
determinism, years of invariants).
- A single static binary (distribution polish, orthogonal to speed).
## Risks
- ABI drift between vendored native grammars and the wasm fallback grammars
(keep both built from the same grammar source revs; CI asserts).
- `.scm` expressiveness ceilings — budget for a per-language "escape hatch"
callback in the emitter before declaring a language blocked.
- napi-rs threading vs the parse-pool: the kernel replaces the wasm workers'
parse+extract; the pool orchestration (file-order commit, retry, recycle)
stays in TS and drives the kernel synchronously per file.