fix(indexing): bounded-memory yielding pipeline tail + daemon session fixes (#1212) (#1226)

Large-codebase indexing died at the end of "Resolving refs" two ways:
watchdog kills of healthy work (24k-file Java on Windows, #1212 — third
iteration of the #1091/#1122 class) and hard OOMs (Linux kernel scale,
where v1.3.0 could not complete at any watchdog setting). Root causes:
~31 of 37 dynamic-edge synthesis passes ran start-to-finish with no
yield points, several materialized whole-graph snapshots (kotlin
expect/actual opened with getAllNodes() — 2M nodes in one array; the
C fn-pointer pass retained every C file's contents twice plus every
function node), and the post-index WAL checkpoint ran minutes of
synchronous IO on the main thread, killing even a successful index at
the finish line.

The pipeline tail now follows the same discipline as the rest: never
hold O(graph) in the heap, yield everywhere.

- All synthesis passes stream node-kind scans (cursors, not arrays) and
  yield on time-budgeted checkpoints; language gates skip passes whose
  filters a project's file languages provably can't satisfy.
- kotlin expect/actual filters SQL-side; c-fnptr caches are LRU-bounded,
  units stream one file at a time, and the all-functions array +
  write-only id map are gone; spring reads each .java once, not twice.
- runMaintenance moved to a worker thread (own SQLite connection);
  per-file store commits chunk with yields behind a serialized flush
  chain (preserving #1015 file-order determinism); resolver warm-up
  streams the DISTINCT name set; resolution batch-tail and merged-edge
  inserts run in bounded sub-transactions.
- Daemon: fixed a socket-handoff race that could leave a fresh MCP
  session permanently silent (client-hello tail unshifted into a
  flowing stream with zero listeners — the long-standing #662 test
  flake was this real bug); first tool call no longer queues behind
  the query pool's cold start (pool.ready gate).

Validation: Linux kernel (70,129 files, 2.05M nodes, 6.4M edges) fully
indexes in 27m8s on a 2-core/6GB container at default heap + default
watchdog; llvm-project (180k files) completes under 1GB RSS including
kill-and-sync recovery; synthesized-edge and full-graph parity are
byte-identical vs baseline on elasticsearch/redis/vim; the ex-flaky
daemon test passed 25/25 under load. Env-gated diagnostics kept:
CODEGRAPH_SYNTH_TIMINGS pass/phase timings, CODEGRAPH_MCP_DEBUG hop
tracing. Design record: docs/design/main-thread-stall-followup.md.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Colby Mchenry
2026-07-08 23:18:23 -05:00
committed by GitHub
co-authored by Claude Fable 5
parent 58b6bf5c60
commit a3f90089e8
21 changed files with 1036 additions and 264 deletions
+44
View File
@@ -880,6 +880,37 @@ export class QueryBuilder {
return rows.map(rowToNode);
}
/**
* Stream nodes of one language whose `decorators` JSON array contains
* `decorator`. The LIKE on the JSON text is a cheap index-free pre-filter
* (a decorator name can appear as a substring of another), so callers must
* still exact-check `node.decorators.includes(decorator)`. Exists so the
* kotlin expect/actual synthesizer never materializes the whole node table
* the way `getAllNodes().filter(...)` did — that array alone OOM'd Node's
* default heap on a 2M-node graph (#1212).
*/
*iterateNodesByLanguageWithDecorator(language: Language, decorator: string): IterableIterator<Node> {
// Fresh statement per call — an iterator holds an open cursor (see
// iterateNodesByKind).
const stmt = this.db.prepare(
"SELECT * FROM nodes WHERE language = ? AND decorators LIKE '%' || ? || '%'"
);
for (const row of stmt.iterate(language, `"${decorator}"`)) {
yield rowToNode(row as NodeRow);
}
}
/**
* Distinct languages present in the files table. One indexed aggregate —
* lets the dynamic-edge synthesizers skip passes for languages the project
* doesn't contain at all (a Kotlin pass has no work on a pure-C repo), so
* their cost is zero rather than a full-graph scan that finds nothing (#1212).
*/
getDistinctFileLanguages(): Set<string> {
const rows = this.db.prepare('SELECT DISTINCT language FROM files').all() as Array<{ language: string }>;
return new Set(rows.map((r) => r.language));
}
/**
* Get nodes by exact name match (uses idx_nodes_name index)
*/
@@ -1853,6 +1884,19 @@ export class QueryBuilder {
return rows.map((r) => r.name);
}
/**
* Stream the distinct node names one row at a time — the incremental
* counterpart to {@link getAllNodeNames} for callers that need to yield
* to the event loop mid-scan (resolver cache warm-up on multi-million-node
* indexes). Fresh statement per call: the iterator holds an open cursor.
*/
*iterateNodeNames(): IterableIterator<string> {
const stmt = this.db.prepare('SELECT DISTINCT name FROM nodes');
for (const row of stmt.iterate()) {
yield (row as { name: string }).name;
}
}
/**
* Get unresolved references scoped to specific file paths.
* Uses the idx_unresolved_file_path index for efficient lookup.