feat(mcp): worker-thread liveness watchdog to self-kill a wedged main thread (#856)
Belt-and-suspenders follow-up to #855. Any non-yielding sync loop on the main thread wedges the event loop, and nothing running on that loop (timers, signal handlers, PPID watchdog) can recover it — only another thread can. A tiny worker thread (in the detached daemon + direct modes) watches a shared-memory heartbeat the main thread bumps each event-loop turn; if it stops advancing across enough consecutive checks (~CODEGRAPH_WATCHDOG_TIMEOUT_MS, default 60s) the worker SIGKILLs the process so a fresh daemon starts on the next connection. Counts consecutive stale checks (not wall-clock) so it's immune to clock jumps / sleep; tuned never to fire on real work; opt out with CODEGRAPH_NO_WATCHDOG=1.
This commit is contained in:
@@ -9,6 +9,10 @@ and adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### New Features
|
||||
|
||||
- The CodeGraph MCP server now self-heals if its main thread ever locks up. A lightweight watchdog notices when the process has stopped responding and stops it so a fresh one starts on your next request — it can no longer sit pinned at 100% CPU with no way to recover. Tune the detection window with `CODEGRAPH_WATCHDOG_TIMEOUT_MS`, or turn it off entirely with `CODEGRAPH_NO_WATCHDOG=1`. (#850)
|
||||
|
||||
### Fixes
|
||||
|
||||
- The CodeGraph MCP server no longer risks getting stuck at 100% CPU after an unexpected internal error. Previously such an error was logged but the process was left running in a broken state, where it could spin a CPU core indefinitely and had to be killed by hand. The server now logs the error and exits cleanly, so a fresh one starts on the next request. Thanks @songhlc. (#850)
|
||||
|
||||
Reference in New Issue
Block a user