Files
codegraph/TELEMETRY.md
T
0a91d0f512 perf(resolution): fix O(K²) import-node blowup in "Resolving refs" (#915) (#965)
* perf(resolution): resolve imports to definitions, not sibling import nodes (#915)

"Resolving refs" crawled (tens of minutes) on large projects — most painfully
ones mixing a big front-end and back-end. An external package or module imported
across hundreds/thousands of files (react, a shared UI package, Python
logging/typing) is re-declared as an `import` node in every importing file, so
its unresolved import ref fell through to the exact-name matcher, which scored
all K same-named import nodes via findBestMatch — K refs x K candidates = O(K^2)
per package, producing only meaningless import->import edges.

Fix: exclude `import`-kind nodes as name-match targets (they're statements, not
definitions; real import->definition resolution is the import resolver's job).
Plus two safe constant-factor wins in findBestMatch: hoist the per-candidate
ref.filePath split, and skip cross-language candidates when a same-language one
exists (provably the same winner — same-language scores >=50, cross-language
maxes at 35).

Measured: superset (Py+TS) candidates scored 7.5M -> 833K (9x), non-import edges
preserved (+1618 now resolve to real defs), ~22K useless import->import edges
removed; kubernetes (Go) computePathProximity 37.2s -> 5.0s; synthetic 8k-file
mixed repo (K=4000) resolution 16.0s -> 1.7s. Full suite green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: correct stale better-sqlite3/wasm references to node:sqlite

The SQLite backend has been Node's built-in node:sqlite (real SQLite, WAL + FTS5,
from the bundled runtime) for a while — there is no native build step and no
node-sqlite3-wasm fallback. README and the docs site were already updated; this
catches the stragglers:

- CLAUDE.md: the src/db/ backend description and the sqlite-backend test note.
- src/db/index.ts, src/mcp/tools.ts: two code comments that still blamed "the
  wasm backend" for non-WAL behavior (reworded to "when WAL isn't in effect").

Leaves tree-sitter grammar wasm (web-tree-sitter / --liftoff-only) untouched —
that's a different, still-current use of wasm.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(telemetry): drop the dead sqlite_backend field (schema v2)

node:sqlite is now the only backend, so the `index` event's `sqlite_backend`
field was a constant ("native") carrying no signal — and the `install` event
never actually sent it. Remove the field and the backendKind() helper, bump the
telemetry SCHEMA_VERSION 1 -> 2, and update TELEMETRY.md + docs/design/telemetry.md.

The ingest worker is deliberately left tolerant: `index` doesn't require the
field and schema_version validates as nonNegInt(99), so v2 events ingest fine and
old clients still sending v1 + sqlite_backend keep validating too. Added a legacy
comment there explaining it's safe to drop once old-client share is negligible.

telemetry.test.ts: the assertion pinning schema_version and a stale-claim fixture
line updated 1 -> 2. All telemetry tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 11:26:23 -05:00

4.3 KiB

Telemetry

CodeGraph collects a small set of anonymous usage statistics — which commands and tools get used, which languages get indexed, which agents drive usage — so we can tell which of the 20+ languages and 8 agent integrations deserve the most work. This page is the complete list of what is collected. If a field isn't on this page, it isn't collected; the ingest endpoint enforces this list as an allowlist and is itself public, auditable code in this repository.

Turning it off

Any of these works, permanently:

codegraph telemetry off        # stores your choice (and deletes any unsent data)
export CODEGRAPH_TELEMETRY=0   # per-shell / per-CI override
export DO_NOT_TRACK=1          # the cross-tool standard — always honored

codegraph telemetry status shows the current state, what decided it, and your machine ID. The interactive installer (codegraph install) asks up front with a visible default-on toggle and never re-asks. If you never saw the installer (e.g. npx straight into init), a one-line notice is printed to stderr before the first time anything is sent.

Off means off: when disabled, CodeGraph records nothing, opens no connection to the telemetry endpoint, and sends no "opted out" ping.

What is collected

Every payload carries this envelope:

field example notes
machine_id b3a8c1… random UUID minted on first send — derived from nothing
codegraph_version 0.9.9
os / arch darwin / arm64 platform identifiers only
node_major 22 major version only
ci false whether the CI env var was set
schema_version 2 bumped when this page changes (v2 dropped the index event's sqlite_backend field)

And one of four events:

  • install — when codegraph install configures agents: which agents (["claude","cursor",…]), global vs project-local, and whether it was a fresh install, an upgrade, or a re-run.
  • index — when a full index completes: the language names present (e.g. ["typescript","go"]), the file count as a coarse bucket (<100, 100-1k, 1k-10k, 10k+), and the duration as a bucket (<10s, 10-60s, 1-5m, 5m+).
  • usage_rollup — one line per day per tool: the tool or CLI command name (e.g. codegraph_explore, init), how many times it ran, how many errored, and — for MCP tools — the connecting agent's name and version from the MCP handshake (e.g. Claude Code 2.1).
  • uninstall — when codegraph uninstall/uninit runs: which agents were removed.

Usage is aggregated locally into daily totals before anything is sent — there is no per-call event stream, and nothing is sent in real time.

What is never collected

  • No source code. No file paths, file names, directory names, repository names or URLs, symbol names, search queries, or anything else derived from the contents of an indexed project.
  • No IP addresses. The ingest endpoint never reads, logs, or forwards the client IP, and IP discarding is enabled at the analytics backend on top of that. No geolocation.
  • No fingerprinting. The machine ID is a random UUID stored in ~/.codegraph/telemetry.json — delete that file (or run codegraph telemetry off, then on) and the old ID is gone forever, with no way to reconnect it.
  • No personal data. No usernames, hostnames, emails, or environment variables.

How it travels

Events POST to telemetry.getcodegraph.com — a first-party endpoint whose complete source lives in telemetry-worker/ in this repository. It validates every event and property against the allowlist above (anything else is dropped), strips IPs, rate-limits, and forwards to a managed analytics store (PostHog, US region) as anonymous events. Sends are fire-and-forget with a short timeout: offline or air-gapped machines buffer a bounded local file (256 KB cap) and never retry-loop, log errors, or slow a command down. Telemetry never adds latency to MCP tool calls — recording is an in-memory counter.

The engineering contract behind all of this — including the rule that schema changes must update this page, the client, and the public endpoint in one PR — is in docs/design/telemetry.md.