On a very large repo (the report is a ~93k-file / 5.7GB-DB Java monorepo) the
first MCP `tools/call` after a fresh `serve --mcp` could hang for 10+ minutes
with zero output, and with the liveness watchdog on, the daemon was SIGKILLed
mid-query instead. Root cause: the post-open catch-up reconcile that the first
tool call is gated on does ~2*N synchronous `fs.existsSync`/`fs.statSync` calls
plus a load-all-files query in two non-yielding loops. On a huge repo that wedges
the event loop for minutes, which (a) trips the 60s watchdog (it SIGKILLs a
process whose loop stops turning) and (b) blocks the first call the whole time.
Two complementary fixes:
- Make the reconcile yield. `ExtractionOrchestrator.sync()` now uses the
yielding `scanDirectoryAsync`, and both O(files) reconcile loops
`await setImmediate` every SYNC_RECONCILE_YIELD_INTERVAL (1000) files. The loop
can no longer wedge the main thread, so the watchdog stays fed and the socket /
any concurrent read stays responsive while a big reconcile runs. Results are
unchanged — only yield points are added.
- Time-box the catch-up gate. The first `tools/call` now waits on the reconcile
for at most CODEGRAPH_CATCHUP_GATE_TIMEOUT_MS (default 3000ms), then serves and
lets the reconcile finish in the background (which now yields, so the served
call runs concurrently). `=0` restores the old unbounded wait. On a normal repo
the reconcile finishes well under the budget, so behavior is unchanged.
Tests: adds two time-box cases to mcp-catchup-gate (serves promptly when the
reconcile runs long; `=0` restores the unbounded wait). Full suite green
(1655 passed). Validated end-to-end through the real daemon: first call returns
at the ~3s time-box instead of waiting an injected 8s reconcile; no-delay control
unchanged; `=0` opt-out waits the full reconcile.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>