fix(resolution): sweep orphaned unresolved refs so an interrupted index heals on sync (#1187) (#1191)
An indexing run killed mid-"Resolving refs" (crash, Ctrl-C, the #1122 watchdog kill) left the refs it never reached parked in unresolved_refs. The git-scoped sync fast path only re-resolves changed files' refs, so those files' call edges were missing permanently — a too-small blast radius clustering by package/module (the #1187 field report: 3 of 10 caller files for a Spring @Resource-injected method) — until a full re-index. - sync() now sweeps leftover unresolved refs with the batched resolver after its scoped pass, including on no-change syncs, so a bare `codegraph sync` recovers a wedged index (and heals pre-fix indexes on the first post-upgrade sync) - the scoped pass deletes unresolvable rows too (parity with the batched path), making "rows at rest" a sound orphan signal - drop the batched loop's early break that abandoned all later batches when one batch was all-unresolvable (its rows WERE consumed — that early stop could orphan the rest of the table at init) - surface the state: `codegraph status` warns, `status --json` gains index.pendingRefs, and MCP codegraph_status tells agents the blast radius is incomplete until the next sync Verified end-to-end on a 2,414-file synthetic Spring repo: SIGKILL mid-resolution reproduces the reporter's exact 3-of-10-callers state; a bare sync now heals it to 10/10 with the edge count converging to the clean-init total; a healthy-index sync stays a no-op. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
7f325134e0
commit
4c15f84aa4
+24
-5
@@ -968,6 +968,23 @@ export class ReferenceResolver {
|
||||
);
|
||||
}
|
||||
|
||||
// Delete unresolvable refs too — parity with resolveAndPersistBatched.
|
||||
// Keeping them bought nothing: a ref is only ever retried when its file
|
||||
// is re-extracted, which cascade-deletes and re-inserts its rows anyway.
|
||||
// And it broke the #1187 orphan sweep's invariant — after a COMPLETED
|
||||
// pass the table must hold nothing that pass processed, so that any row
|
||||
// still present belongs to an interrupted run and the sweep can key off
|
||||
// a bare row count.
|
||||
if (result.unresolved.length > 0) {
|
||||
this.queries.deleteSpecificResolvedReferences(
|
||||
result.unresolved.map((r) => ({
|
||||
fromNodeId: r.fromNodeId,
|
||||
referenceName: r.referenceName,
|
||||
referenceKind: r.referenceKind,
|
||||
}))
|
||||
);
|
||||
}
|
||||
|
||||
return result;
|
||||
}
|
||||
|
||||
@@ -1152,11 +1169,13 @@ export class ReferenceResolver {
|
||||
// Yield so progress UI can render between batches
|
||||
await new Promise(resolve => setImmediate(resolve));
|
||||
|
||||
// If nothing was resolved or removed in this batch, we'd loop forever
|
||||
// on the same rows. Break to avoid infinite loop.
|
||||
if (result.resolved.length === 0 && result.unresolved.length === batch.length) {
|
||||
break;
|
||||
}
|
||||
// NOTE: there used to be an extra early break here when a batch resolved
|
||||
// nothing (`result.unresolved.length === batch.length`). That was wrong:
|
||||
// an all-unresolvable batch still DELETES its rows (progress), yet the
|
||||
// break abandoned every batch after it in the same run — on a repo whose
|
||||
// first 5000 refs are all external/stdlib calls, resolution stopped at
|
||||
// batch one and left the rest of the table as permanent orphans (#1187).
|
||||
// The count-based guard below catches the true no-progress case.
|
||||
|
||||
// Non-progress guard (defense-in-depth). Because we re-read from offset 0
|
||||
// each pass, the unresolved_refs table MUST shrink every iteration — both
|
||||
|
||||
Reference in New Issue
Block a user