perf(index): faster fresh indexing + parallel reference resolution, byte-identical graphs (#1305)
* perf(index): ~34% faster fresh indexing, byte-identical graphs Profiling a fresh init on a medium TS repo (excalidraw, 657 files) showed the main thread as the critical path: per-row SQLite statement calls, repeated import-resolution walks, and per-row FTS trigger firings, with the parse workers ~75% idle behind it. This lands the semantics-preserving tranche of fixes: - Multi-row batched INSERTs (nodes/edges/unresolved refs/name segments) behind cached per-batch-size prepared statements; row order preserved, so rowid-based resolution determinism (#1015) is unchanged. - storeFileBundle: one transaction per file instead of four; nested transaction() calls now flatten (BEGIN-in-BEGIN previously threw, so no caller depended on nested rollback). - Dedicated store-writer thread for the fresh-DB bulk path (bundles applied in file order on a single writer connection; main thread does no DB work during the parse loop). Kill switch: CODEGRAPH_NO_STORE_WORKER=1. - Bulk FTS mode: drop the nodes_fts sync triggers during the bulk load, rebuild once at the end; crash inside the window self-heals on the next open. - Per-context memos for resolveImportPath/findExportedSymbol + a per-file exported-symbol index, invalidated exactly where clearCaches() already resets the resolver's own caches. - Fast-init on completely fresh DBs (journal in memory, no fsync until the index completes; interrupted init re-runs from scratch). Kill switch: CODEGRAPH_NO_FAST_INIT=1. - MaybeYield returns undefined on the not-due path so per-ref yield checks stop paying a promise + microtask hop each. - Parse pool prewarm for bulk indexing; compile-cache enabled at CLI and worker entry points. Excalidraw fresh init: 5.11s -> 3.36s median (n=5, warm cache, M-series). Graph dumps byte-identical across init, re-index, and sync paths; full suite green (2403 passed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * perf(resolution): parallel reference resolution with canonical admission Fan resolution batches across a pool of read-only worker threads, each hosting a full ReferenceResolver over its own SQLite connection; results are admitted on the main thread in chunk order, so edge insertion order, row cleanup, failure parking, and deferred post-pass queues are exactly the sequence the single-threaded loop produces. Per-ref inputs match the baseline because the sequential path already resolves each batch against the state committed BEFORE that batch. Validated byte-identical on excalidraw (pool forced on) and apache/dubbo (4,048 Java files): dubbo full index 39s -> 19s (2.05x) with identical graph dumps (91,495 nodes / 223,953 edges). The pool only engages when total pending refs clear a threshold (default 150k, CODEGRAPH_PARALLEL_RESOLVE_MIN to tune, CODEGRAPH_NO_PARALLEL_RESOLVE=1 to disable): measured on a ~58k-ref repo the workers' boot CPU contends with resolution on the same cores and makes indexing slower, so small repos keep the sequential path. When fast-init left the DB in memory-journal mode, WAL is restored before resolution only when the pool will run (readers + rollback-journal writers don't mix). Also: sqlite adapter readOnly open support. TreeCursor spine rewrite of the body walker was built, measured neutral on real repos and equal in a 20k-child microbench (web-tree-sitter's namedChild(i) is not quadratic in this binding), and rejected — per-node JS<->WASM marshaling is the floor, which a traversal swap cannot remove. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
246aee8373
commit
5736e24bb6
@@ -112,9 +112,78 @@ export class DatabaseConnection {
|
||||
runMigrations(db, currentVersion);
|
||||
}
|
||||
|
||||
// Self-heal a bulk-load window that never closed (crash between
|
||||
// beginBulkNodeLoad and endBulkNodeLoad): the FTS triggers are missing and
|
||||
// nodes_fts is stale. Rebuild + recreate so search stays in sync.
|
||||
conn.healBulkNodeLoad();
|
||||
|
||||
return conn;
|
||||
}
|
||||
|
||||
/**
|
||||
* FTS maintenance triggers dropped/recreated around a bulk load.
|
||||
* Names must match schema.sql.
|
||||
*/
|
||||
private static readonly FTS_TRIGGER_NAMES = ['nodes_ai', 'nodes_ad', 'nodes_au'] as const;
|
||||
|
||||
/**
|
||||
* Enter bulk-load mode: drop the per-row FTS sync triggers so mass node
|
||||
* inserts skip per-row tokenization. MUST be paired with endBulkNodeLoad()
|
||||
* (use try/finally); a crash inside the window is healed on the next open().
|
||||
* The window is DB-wide (triggers are schema objects), which is safe because
|
||||
* endBulkNodeLoad() rebuilds nodes_fts from the nodes table wholesale — any
|
||||
* row written by anyone during the window is captured by the rebuild.
|
||||
*/
|
||||
beginBulkNodeLoad(): void {
|
||||
for (const t of DatabaseConnection.FTS_TRIGGER_NAMES) {
|
||||
this.db.exec(`DROP TRIGGER IF EXISTS ${t}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Leave bulk-load mode: rebuild the whole FTS index from the nodes table in
|
||||
* one pass (far cheaper than per-row trigger firings), then recreate the
|
||||
* triggers by re-running schema.sql (idempotent — everything in it is
|
||||
* IF NOT EXISTS).
|
||||
*/
|
||||
endBulkNodeLoad(): void {
|
||||
this.db.exec(`INSERT INTO nodes_fts(nodes_fts) VALUES('rebuild')`);
|
||||
this.recreateFtsTriggers();
|
||||
}
|
||||
|
||||
/** Recreate the FTS triggers + rebuild if a bulk-load window never closed. */
|
||||
private healBulkNodeLoad(): void {
|
||||
const row = this.db
|
||||
.prepare(
|
||||
`SELECT count(*) AS c FROM sqlite_master WHERE type = 'trigger' AND name IN ('nodes_ai','nodes_ad','nodes_au')`
|
||||
)
|
||||
.get() as { c: number } | undefined;
|
||||
if ((row?.c ?? 0) >= DatabaseConnection.FTS_TRIGGER_NAMES.length) return;
|
||||
this.endBulkNodeLoad();
|
||||
}
|
||||
|
||||
/**
|
||||
* Recreate the FTS sync triggers from schema.sql — extracted from the file
|
||||
* rather than duplicated here so the DDL cannot drift from the schema.
|
||||
* (Re-execing the whole schema is not an option: it contains data INSERTs
|
||||
* that are not idempotent, e.g. schema_versions.)
|
||||
*/
|
||||
private recreateFtsTriggers(): void {
|
||||
const schemaPath = path.join(__dirname, 'schema.sql');
|
||||
const schema = fs.readFileSync(schemaPath, 'utf-8');
|
||||
const triggerDdls = schema.match(
|
||||
/CREATE TRIGGER IF NOT EXISTS nodes_a[idu]\b[\s\S]*?END;/g
|
||||
);
|
||||
if (!triggerDdls || triggerDdls.length !== DatabaseConnection.FTS_TRIGGER_NAMES.length) {
|
||||
throw new Error(
|
||||
`schema.sql: expected ${DatabaseConnection.FTS_TRIGGER_NAMES.length} nodes FTS triggers, found ${triggerDdls?.length ?? 0}`
|
||||
);
|
||||
}
|
||||
for (const ddl of triggerDdls) {
|
||||
this.db.exec(ddl);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Get the underlying database instance
|
||||
*/
|
||||
|
||||
Reference in New Issue
Block a user