Large-codebase indexing died at the end of "Resolving refs" two ways: watchdog kills of healthy work (24k-file Java on Windows, #1212 — third iteration of the #1091/#1122 class) and hard OOMs (Linux kernel scale, where v1.3.0 could not complete at any watchdog setting). Root causes: ~31 of 37 dynamic-edge synthesis passes ran start-to-finish with no yield points, several materialized whole-graph snapshots (kotlin expect/actual opened with getAllNodes() — 2M nodes in one array; the C fn-pointer pass retained every C file's contents twice plus every function node), and the post-index WAL checkpoint ran minutes of synchronous IO on the main thread, killing even a successful index at the finish line. The pipeline tail now follows the same discipline as the rest: never hold O(graph) in the heap, yield everywhere. - All synthesis passes stream node-kind scans (cursors, not arrays) and yield on time-budgeted checkpoints; language gates skip passes whose filters a project's file languages provably can't satisfy. - kotlin expect/actual filters SQL-side; c-fnptr caches are LRU-bounded, units stream one file at a time, and the all-functions array + write-only id map are gone; spring reads each .java once, not twice. - runMaintenance moved to a worker thread (own SQLite connection); per-file store commits chunk with yields behind a serialized flush chain (preserving #1015 file-order determinism); resolver warm-up streams the DISTINCT name set; resolution batch-tail and merged-edge inserts run in bounded sub-transactions. - Daemon: fixed a socket-handoff race that could leave a fresh MCP session permanently silent (client-hello tail unshifted into a flowing stream with zero listeners — the long-standing #662 test flake was this real bug); first tool call no longer queues behind the query pool's cold start (pool.ready gate). Validation: Linux kernel (70,129 files, 2.05M nodes, 6.4M edges) fully indexes in 27m8s on a 2-core/6GB container at default heap + default watchdog; llvm-project (180k files) completes under 1GB RSS including kill-and-sync recovery; synthesized-edge and full-graph parity are byte-identical vs baseline on elasticsearch/redis/vim; the ex-flaky daemon test passed 25/25 under load. Env-gated diagnostics kept: CODEGRAPH_SYNTH_TIMINGS pass/phase timings, CODEGRAPH_MCP_DEBUG hop tracing. Design record: docs/design/main-thread-stall-followup.md. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
58b6bf5c60
commit
a3f90089e8
@@ -233,21 +233,22 @@ describe('runMaintenance', () => {
|
||||
if (fs.existsSync(dir)) fs.rmSync(dir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it('runs without throwing on a fresh database', () => {
|
||||
expect(() => db.runMaintenance()).not.toThrow();
|
||||
it('runs without throwing on a fresh database', async () => {
|
||||
await expect(db.runMaintenance()).resolves.toBeUndefined();
|
||||
});
|
||||
|
||||
it('runs without throwing after writes', () => {
|
||||
it('runs without throwing after writes', async () => {
|
||||
const q = new QueryBuilder(db.getDb());
|
||||
q.insertNodes([makeNode('n1'), makeNode('n2')]);
|
||||
expect(() => db.runMaintenance()).not.toThrow();
|
||||
await expect(db.runMaintenance()).resolves.toBeUndefined();
|
||||
});
|
||||
|
||||
it('swallows failures rather than propagating (best-effort)', () => {
|
||||
it('swallows failures rather than propagating (best-effort)', async () => {
|
||||
// Close the DB so the underlying handle would normally throw on any
|
||||
// exec(). runMaintenance must still not propagate.
|
||||
// exec(). runMaintenance (worker on its own connection, or the in-line
|
||||
// fallback) must still not propagate.
|
||||
db.close();
|
||||
expect(() => db.runMaintenance()).not.toThrow();
|
||||
await expect(db.runMaintenance()).resolves.toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
Reference in New Issue
Block a user