CG-33: converge incremental sync with a full rebuild

A live, auto-synced index did not converge to a clean rebuild of the same
tree — 4.3% of distinct edges wrong in both directions on this repo's own
index, overwhelmingly `calls`, which is what flow queries traverse and what
explore's file ranking weights. Silent: nothing warned, and the symptom read
as "codegraph isn't very good" rather than "this index needs rebuilding."

Two causes, and the fix needed both. Resolution binds a reference to one of
the same-named definitions PROJECT-WIDE, so a definition appearing or
vanishing changes the correct answer for references in files the sync never
touches — and those references resolved successfully once, which deletes
their unresolved_refs row, leaving nothing to revisit them with (#1240's
retry only revisits refs parked as failed). Separately, when nothing
disambiguated the candidates the winner came down to rowid, i.e. the order
files happened to be WRITTEN, which differs between a scan-order full index
and a sync that appends each file as it changes. That second one is why
re-resolution alone could not converge: re-resolving against the identical
graph still picked a different candidate.

So getNodesByName now orders by (file_path, start_line) — a property of the
code, not of the write order — and sync computes a definitionDelta and
re-opens the resolution edges whose answer it may have invalidated,
re-inserting each as the reference that created it for the orphan sweep to
bind against the post-sync graph.

The delta compares `file\0name` pairs per file rather than one name set over
the batch: a commit that adds `collect` to a new file while an unrelated
changed file already defines `collect` cancels out of a batch-wide set, and
that miss was the largest residual class in the first measurement.

Conservative where the failure modes are asymmetric — a wrong deletion is a
permanent edge loss, a missed rebind is only residual drift. Edges without a
refName stamp are never touched (nothing to restore them from), sources the
sync already re-extracted are skipped, and a per-name ceiling declines the
generic names. Edges are deleted before the sweep re-inserts, since
INSERT OR IGNORE against idx_edges_identity would otherwise keep both rows
when a reference rebinds elsewhere.

Replaying real commits of this repo through sync, then diffing against a
rebuild: 16 commits 48 -> 0; 80 commits 1,634 -> 361, with the actively
misleading direction (stale edges the index keeps asserting) 671 -> 2.
Index and sync wall-clock are unchanged; the ORDER BY costs 18% per uncached
name lookup, which never reaches wall-clock because the resolver memoizes it.

The 357-edge residual at 80 commits is one pre-existing class: refs to
generic names (`push`, `join`) parked above #1240's per-name retry ceiling,
which a rebuild resolves into cross-language garbage — a TS test file
"calling" an R method. Converging there would mean manufacturing wrong edges,
so it is left alone. And no drift metric in `codegraph status`: it cannot be
computed without the rebuild it would be recommending, and a proxy would fire
on that residual and train users to ignore it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Colby McHenry
2026-08-06 04:32:40 -05:00
co-authored by Claude Opus 5
parent 2cf63fd114
commit 03893b0ab9
6 changed files with 640 additions and 12 deletions
+139 -2
View File
@@ -1113,11 +1113,28 @@ export class QueryBuilder {
}
/**
* Get nodes by exact name match (uses idx_nodes_name index)
* Get nodes by exact name match (uses idx_nodes_name index).
*
* This is resolution's candidate list, and the ORDER BY is load-bearing for
* index correctness, not cosmetic (CG-33). When a reference names a symbol
* that several files define and nothing disambiguates them, resolution binds
* to the first candidate — so without an ORDER BY the winner was decided by
* rowid, i.e. by the order files happened to be WRITTEN. A full index writes
* them in scan order; an incremental sync appends each file as it changes, so
* the same tree resolved to different edges depending on how the index was
* built, and a long-lived synced index drifted away from a rebuild of itself
* (measured at 4.3% of distinct edges, mostly `calls`).
*
* `(file_path, start_line)` is a property of the CODE, so both paths now pick
* the same candidate. The sort is paid once per distinct name per resolution
* run — ReferenceResolver memoizes this in its nameCache — and the population
* is capped by AMBIGUOUS_NAME_CEILING (#999).
*/
getNodesByName(name: string): Node[] {
if (!this.stmts.getNodesByName) {
this.stmts.getNodesByName = this.db.prepare('SELECT * FROM nodes WHERE name = ?');
this.stmts.getNodesByName = this.db.prepare(
'SELECT * FROM nodes WHERE name = ? ORDER BY file_path, start_line'
);
}
const rows = this.stmts.getNodesByName.all(name) as NodeRow[];
return rows.map(rowToNode);
@@ -2445,6 +2462,99 @@ export class QueryBuilder {
}));
}
/**
* Resolution edges whose TARGET symbol is named one of `names` — the edges a
* sync must re-resolve after `names` gained or lost a definition (CG-33).
*
* Resolution binds a reference to a node whose name matches the reference's
* tail, and it picks among ALL same-named definitions project-wide. So adding
* or removing one definition of `pct` changes the answer for every `pct(...)`
* reference in the repo — including references in files this sync never
* touches, whose edges nothing else revisits. Those edges' current target is,
* by that same rule, a node named `pct`, which is why the target's name is a
* sufficient (and index-backed, via idx_nodes_name) way to find them without
* a schema change or a scan of edge metadata.
*
* Returns the source file/language alongside each edge so the caller can
* resurrect it as its original reference. Excludes `provenance='heuristic'`
* (synthesized dispatch edges are not resolution output and carry no refName
* stamp to resurrect from — deleting one would be a permanent loss).
*
* Names matching more than `perNameCeiling` edges are skipped entirely, same
* rationale and same default as {@link getRetryableFailedReferences}: at that
* population the name is generic (`get`, `clear`, …), one definition changing
* won't flip most of them, and rebinding an arbitrary subset is both wasted
* work and incoherent coverage.
*/
getResolutionEdgesByTargetName(
names: string[],
perNameCeiling: number = 500
): Array<Edge & { edgeId: number; sourceFilePath: string; sourceLanguage: Language }> {
if (names.length === 0) return [];
// Pass 1: per-name edge counts, chunked under the SQLite parameter limit.
const keep: string[] = [];
for (let i = 0; i < names.length; i += SQLITE_PARAM_CHUNK_SIZE) {
const chunk = names.slice(i, i + SQLITE_PARAM_CHUNK_SIZE);
const placeholders = chunk.map(() => '?').join(',');
const counts = this.db
.prepare(
`SELECT tgt.name AS name, COUNT(*) AS count
FROM edges e
JOIN nodes tgt ON tgt.id = e.target
WHERE tgt.name IN (${placeholders})
AND (e.provenance IS NULL OR e.provenance != 'heuristic')
GROUP BY tgt.name`
)
.all(...chunk) as Array<{ name: string; count: number }>;
for (const row of counts) {
if (row.count <= perNameCeiling) keep.push(row.name);
}
}
if (keep.length === 0) return [];
// Pass 2: load the surviving edges with the source file context a
// resurrection needs.
const out: Array<Edge & { edgeId: number; sourceFilePath: string; sourceLanguage: Language }> = [];
for (let i = 0; i < keep.length; i += SQLITE_PARAM_CHUNK_SIZE) {
const chunk = keep.slice(i, i + SQLITE_PARAM_CHUNK_SIZE);
const placeholders = chunk.map(() => '?').join(',');
const rows = this.db
.prepare(
`SELECT e.*, src.file_path AS source_file_path, src.language AS source_language
FROM edges e
JOIN nodes tgt ON tgt.id = e.target
JOIN nodes src ON src.id = e.source
WHERE tgt.name IN (${placeholders})
AND (e.provenance IS NULL OR e.provenance != 'heuristic')`
)
.all(...chunk) as Array<EdgeRow & { source_file_path: string; source_language: Language }>;
for (const row of rows) {
out.push({
...rowToEdge(row),
edgeId: row.id,
sourceFilePath: row.source_file_path,
sourceLanguage: row.source_language,
});
}
}
return out;
}
/** Delete edges by primary key — the rebind pass's half of a re-resolution. */
deleteEdgesByIds(edgeIds: number[]): number {
if (edgeIds.length === 0) return 0;
let changed = 0;
this.db.transaction(() => {
for (let i = 0; i < edgeIds.length; i += SQLITE_PARAM_CHUNK_SIZE) {
const chunk = edgeIds.slice(i, i + SQLITE_PARAM_CHUNK_SIZE);
const placeholders = chunk.map(() => '?').join(',');
changed += this.db.prepare(`DELETE FROM edges WHERE id IN (${placeholders})`).run(...chunk).changes;
}
})();
return changed;
}
/**
* Distinct node names present in the given files — the symbol names a sync
* pass uses to look up retryable failed refs after those files changed.
@@ -2463,6 +2573,33 @@ export class QueryBuilder {
return [...names];
}
/**
* Distinct `file\0name` pairs defined by the given files — the shape sync's
* definition delta needs (CG-33).
*
* Deliberately NOT `getNodeNamesByFiles`: a bare name set is taken over the
* WHOLE changed batch, so a name that moves between two files in one commit
* (or exists in one changed file and is newly added to another) appears on
* both sides and cancels out of the symmetric difference — even though a
* definition genuinely appeared or vanished and every reference to that name
* repo-wide may now bind elsewhere. Keying by file makes each definition its
* own fact, so the move is seen as one removal plus one addition.
*/
getNodeNamePairsByFiles(filePaths: string[]): Set<string> {
const pairs = new Set<string>();
if (filePaths.length === 0) return pairs;
for (let i = 0; i < filePaths.length; i += SQLITE_PARAM_CHUNK_SIZE) {
const chunk = filePaths.slice(i, i + SQLITE_PARAM_CHUNK_SIZE);
const placeholders = chunk.map(() => '?').join(',');
const rows = this.db
.prepare(`SELECT DISTINCT file_path, name FROM nodes WHERE file_path IN (${placeholders})`)
.all(...chunk) as Array<{ file_path: string; name: string }>;
// NUL-joined: a path or a symbol name can contain a space, never a NUL.
for (const row of rows) pairs.add(`${row.file_path}\0${row.name}`);
}
return pairs;
}
// ===========================================================================
// Statistics
// ===========================================================================
+107
View File
@@ -116,6 +116,20 @@ export interface SyncResult {
nodesUpdated: number;
durationMs: number;
changedFilePaths?: string[];
/**
* Symbol names whose set of definitions this sync CHANGED — names the synced
* files gained or lost, as the symmetric difference of their `file\0name`
* definition pairs before and after the store phase (per file, so a name
* moving between two changed files does not cancel itself out).
* Resolution picks among all same-named definitions project-wide,
* so these are exactly the names whose already-resolved edges — in files this
* sync never touched — may now bind elsewhere and must be re-resolved for the
* index to stay convergent with a full rebuild (CG-33).
*
* A body-only edit leaves this empty, which is the common case and costs
* nothing downstream.
*/
definitionDelta?: string[];
}
/**
@@ -2491,6 +2505,64 @@ export class ExtractionOrchestrator {
}
}
/**
* Re-open, for re-resolution, every resolution edge whose answer this sync
* may have changed — the fix for index drift (CG-33).
*
* Incremental sync re-resolves only the references IN the changed files, but
* resolution's answer is a function of the WHOLE graph: a reference binds to
* one of the same-named definitions project-wide, so adding or removing a
* definition of `pct` can change which `pct` every other file's `pct(...)`
* should bind to. Those other files are never revisited, and their references
* resolved successfully once and were deleted from `unresolved_refs`, so
* nothing existed to revisit them with — the index kept an answer that was
* correct against an older graph. Measured on codegraph's own long-lived
* index: 4.3% of distinct edges differed from a clean rebuild, in BOTH
* directions, overwhelmingly `calls`. See docs/benchmarks/index-drift-cg33.md.
*
* This deletes each affected edge and re-inserts it as the reference that
* created it (the refName/refKind stamp), status='pending', for the sync's
* resolution sweep to bind against the post-sync graph — the same input a
* full rebuild resolves from, which is what makes the two converge.
*
* Deliberately conservative in three ways, because a wrong deletion is a
* permanent edge loss while a missed rebind is only residual drift:
* - an edge with no refName stamp (synthesized, or built by an engine older
* than the stamp) is left ALONE rather than reconstructed from the target's
* plain name, same rule as `resurrectRefFromDroppedEdge`;
* - edges whose source is in a file this sync already re-extracted are
* skipped — their references were re-resolved from scratch moments ago;
* - very common names are skipped by the per-name ceiling in
* `getResolutionEdgesByTargetName`.
*
* Returns the number of references resurrected.
*/
resurrectStaleResolutionEdges(definitionDelta: string[], changedFilePaths: string[]): number {
if (definitionDelta.length === 0) return 0;
const alreadyFresh = new Set(changedFilePaths);
const candidates = this.queries.getResolutionEdgesByTargetName(definitionDelta);
const edgeIds: number[] = [];
const refs: UnresolvedReference[] = [];
for (const e of candidates) {
if (alreadyFresh.has(e.sourceFilePath)) continue;
const ref = resurrectRefFromDroppedEdge(e);
if (!ref) continue; // no stamp — never delete what we cannot restore
edgeIds.push(e.edgeId);
refs.push(ref);
}
if (refs.length === 0) return 0;
// Delete first. The sweep re-inserts whichever edge resolution now picks,
// and `insertEdges` is INSERT OR IGNORE against idx_edges_identity — so a
// rebind to the same target is a clean no-op, but leaving the old row in
// place for a rebind ELSEWHERE would keep both, turning drift into
// duplication.
this.queries.deleteEdgesByIds(edgeIds);
this.queries.insertUnresolvedRefsBatch(refs);
return refs.length;
}
/**
* Sync the index with the current file state.
*
@@ -2520,6 +2592,10 @@ export class ExtractionOrchestrator {
let filesRemoved = 0;
let nodesUpdated = 0;
const changedFilePaths: string[] = [];
// `file\0name` definition pairs for the files this sync touches, sampled
// BEFORE their nodes are replaced/deleted. Compared against the post-store
// pairs below to derive `definitionDelta` (CG-33).
const pairsBefore = new Set<string>();
onProgress?.({
phase: 'scanning',
@@ -2585,6 +2661,9 @@ export class ExtractionOrchestrator {
// failed until the symbol reappears somewhere. (A deleted file whose
// CALLERS are also being deleted is fine: their nodes cascade later
// in this loop and take the resurrected rows with them.)
// Every name this file defined is about to stop existing here, which
// narrows the candidate set for that name repo-wide (CG-33).
for (const pair of this.queries.getNodeNamePairsByFiles([tracked.path])) pairsBefore.add(pair);
const incoming = this.queries.getCrossFileIncomingEdgesWithTarget(tracked.path);
if (incoming.length > 0) {
const resurrected = incoming
@@ -2651,6 +2730,14 @@ export class ExtractionOrchestrator {
}
}
// Sampled here — after the add/modify classification, before any file is
// re-extracted — because `storeExtractionResult` deletes a file's nodes
// before inserting the new ones, so this is the last point the pre-edit
// definition set is readable (CG-33).
if (filesToIndex.length > 0) {
for (const pair of this.queries.getNodeNamePairsByFiles(filesToIndex)) pairsBefore.add(pair);
}
// Load only grammars needed for changed files
if (filesToIndex.length > 0) {
const overrides = loadExtensionOverrides(this.rootDir);
@@ -2677,6 +2764,25 @@ export class ExtractionOrchestrator {
nodesUpdated += result.nodes.length;
}
// Names whose definition set this sync changed: a `file\0name` pair present
// before but not after (removed/renamed away) or after but not before
// (added). A pair on both sides is untouched as far as resolution's
// candidate set is concerned — only its node id moved, which
// reattachCrossFileEdges already follows — so an edit that only changes
// bodies yields an empty delta and no downstream rebind work (CG-33).
//
// Compared per FILE, not as one name set over the whole batch: a commit
// that adds `collect` to a new file while an unrelated changed file already
// defined `collect` must still flag the name, and a bare name set cancels
// exactly that case out. That miss left the largest residual class in the
// first measurement of this fix.
const pairsAfter = this.queries.getNodeNamePairsByFiles(filesToIndex);
const deltaNames = new Set<string>();
const nameOf = (pair: string) => pair.slice(pair.indexOf('\0') + 1);
for (const pair of pairsBefore) if (!pairsAfter.has(pair)) deltaNames.add(nameOf(pair));
for (const pair of pairsAfter) if (!pairsBefore.has(pair)) deltaNames.add(nameOf(pair));
const definitionDelta = [...deltaNames];
return {
filesChecked,
filesAdded,
@@ -2685,6 +2791,7 @@ export class ExtractionOrchestrator {
nodesUpdated,
durationMs: Date.now() - startTime,
changedFilePaths: changedFilePaths.length > 0 ? changedFilePaths : undefined,
definitionDelta: definitionDelta.length > 0 ? definitionDelta : undefined,
};
}
+26
View File
@@ -883,6 +883,32 @@ export class CodeGraph {
}
}
// Re-open resolution edges this sync may have invalidated ELSEWHERE in
// the repo (CG-33). Everything above re-resolves references in the
// changed files; this covers the opposite direction — references in
// files the sync never touched whose answer depended on a definition
// that just appeared or disappeared. Without it a synced index never
// converges to a full rebuild: measured at 4.3% of distinct edges wrong
// on codegraph's own index, in both directions, mostly `calls`. The
// resurrected refs are pending rows, so the orphan sweep immediately
// below is what resolves them — batched, yielding, multi-pass, exactly
// as a full index resolves.
//
// `definitionDelta` is empty for a body-only edit, so the overwhelmingly
// common sync pays one branch. CODEGRAPH_NO_REBIND=1 disables it.
if (result.definitionDelta && process.env.CODEGRAPH_NO_REBIND !== '1') {
const tRebind = Date.now();
const rebound = this.orchestrator.resurrectStaleResolutionEdges(
result.definitionDelta,
result.changedFilePaths ?? []
);
if (process.env.CODEGRAPH_SYNTH_TIMINGS) {
console.error(
`[phase-timing] sync-rebind: ${Date.now() - tRebind}ms (${result.definitionDelta.length} changed names, ${rebound} edges re-opened)`
);
}
}
// Orphan sweep (#1187). A resolution pass that dies mid-run — the #850
// daemon liveness watchdog's SIGKILL (#1122), Ctrl-C, a crash — leaves
// the refs it never reached in unresolved_refs, and the git-scoped fast