Files
codegraph/docs/design/native-extraction-kernel.md
T
4efc6c70e2 fix(scale): kernel-scale hardening — OOM-safe pass skipping + watchdog-safe index recreate (#1323)
Two hazards found by running today's full stack against the Linux kernel
(70,129 files) in the cg1212 repro container:

1. The parallel-synthesis fallback retried a worker-failed pass on the MAIN
   thread. At multi-million-node scale a worker failure is usually a memory
   ceiling, so the retry would OOM the process and take the whole index with
   it. Above 1.5M nodes a failed pass is now skipped with a clear stderr
   message (its synthesized edges are absent; the index completes). Below
   that, the main-thread retry stays — small-scale worker crashes are
   transient and the retry keeps coverage.

2. endBulkEdgeLoad rebuilt all four edge indexes in one synchronous span —
   measured 79s at kernel scale, past the #850 liveness watchdog's 60s
   stall window. A daemon-triggered re-index would have been SIGKILLed right
   after doing the work. Now async with an event-loop yield between builds,
   keeping each stall to a single index (~20s at kernel scale).

Validation: full Linux kernel index to completion in the repro container —
2,048,674 nodes / 6,405,964 edges, EXIT 0, zero passes skipped, on a 2-CPU
VM (worst case: pool disabled, sequential resolution + synthesis) in ~27min.
Phase walls: parse 6.0m, resolution 19.5m (incl. synthesis 6.3m, recreate
79s), maintenance 74s. Suite green (2444).

Also adds docs/design/native-extraction-kernel.md — the spike-validated
design for the native extraction kernel (Rust parse+walk over dubbo's Java:
202ms rayon / 1.07s single-thread vs 4.7s for the current wasm pipeline).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 19:09:01 -05:00

88 lines
4.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Native extraction kernel — design + spike results
**Status:** spike validated 2026-07-16; project approved, not yet started.
**Owner context:** the last structural lever for fresh-index wall clock after the
2026-07-16 arc (#1305, #1320, #1321, #1322) exhausted Node-side scheduling.
## Why
Post-arc, the fresh-index profile on dubbo (4,402 Java files) is:
parse-loop ~4.7s, resolution ~5.5s (persist-bound), synthesis ~0.9s, total ~11.1s
vs codebase-memory-mcp v0.9.0 at 7.1s. Two levers were measured dead:
- **RAM-backed DB** (parse-loop 6.9s on a ramdisk vs 4.64.8s on SSD, n=2
interleaved): fast-init `synchronous=OFF` already writes at page-cache speed.
The parse phase is CPU-bound.
- **TreeCursor rewrite** (earlier arc): web-tree-sitter's traversal is not the
cost; the floor is per-node JS↔WASM **marshaling** — every `node.kind`,
`.childForFieldName`, `.text` crosses the boundary.
The only remaining parse lever is doing the walk on the native side and
crossing the boundary **once per file** instead of once per node.
## Spike (2026-07-16)
Minimal Rust binary (`tree-sitter` 0.25 + `tree-sitter-java`, TreeCursor walk
touching every node's kind/range + `name`-field text, emitting flat
`(kind_id, start, end, name_len)` rows — the extraction access pattern).
Dubbo's 4,048 `.java` files, 17MB, 3.59M AST nodes, Apple M3 Pro:
| | wall |
|---|---|
| Current pipeline parse-loop (7 wasm workers, incl. extraction + store dispatch) | 4,700ms |
| Rust parse+walk, rayon | **202ms** |
| Rust parse+walk, single thread | 1,067ms |
One native thread beats the whole 7-worker wasm pool 4.4×; at equal
parallelism the walk is ~14× faster. Even charging the kernel for the
extraction logic it must still perform, parse-loop 4.7s → ~1.01.5s is
realistic, putting dubbo ≈ 7.58s total (parity with cbm).
## Architecture
- **Crate:** `codegraph-kernel`, napi-rs, links tree-sitter's C library and
vendored grammars natively. Input: `(filePath, content, language)`. Output:
flat typed buffers (nodes, edges, unresolved refs) — one boundary crossing
per file.
- **Per-language logic:** migrate extractors to tree-sitter **query files**
(`.scm`) executed by a generic Rust emitter; bespoke TS logic that queries
can't express (macro salvage, dialect sniffing, content-gated `.h`
detection) stays as TS pre/post passes over the returned buffers.
- **Distribution:** prebuilt `.node` per platform through the existing
release-bundle pipeline (same per-platform packages as the Node runtime).
The wasm path remains as the universal fallback — same crate compiled to
wasm keeps one implementation.
- **Rollout:** per-language, funnel languages first (TS/JS → Java → Python →
Go). A language ships only when its equivalence gate passes.
## Equivalence gate (per language)
Byte-identity against hand-written extractors is NOT expected (bespoke logic
ports approximately). The gate is:
1. Node/edge/ref **counts** within ±0.5% on 3 real repos (small/medium/large),
with every diff category eyeballed.
2. The retrieval invariants hold: explore-flow connects the language's
canonical flows end-to-end (`docs/design/dynamic-dispatch-coverage-playbook.md`),
agent A/B shows no regression per the standard methodology.
3. Fresh-index wall improves on the language's repos; no regression on a
control repo of a non-migrated language.
## Non-goals
- Porting resolution, synthesis, frameworks, MCP, or the installer — they are
pool-parallel and not marshal-bound. The measured native advantage there is
~1.4× CPU, not worth the correctness moat (2,444 tests, byte-identical
determinism, years of invariants).
- A single static binary (distribution polish, orthogonal to speed).
## Risks
- ABI drift between vendored native grammars and the wasm fallback grammars
(keep both built from the same grammar source revs; CI asserts).
- `.scm` expressiveness ceilings — budget for a per-language "escape hatch"
callback in the emitter before declaring a language blocked.
- napi-rs threading vs the parse-pool: the kernel replaces the wasm workers'
parse+extract; the pool orchestration (file-order commit, retry, recycle)
stays in TS and drives the kernel synchronously per file.