* perf(resolution): resolve imports to definitions, not sibling import nodes (#915) "Resolving refs" crawled (tens of minutes) on large projects — most painfully ones mixing a big front-end and back-end. An external package or module imported across hundreds/thousands of files (react, a shared UI package, Python logging/typing) is re-declared as an `import` node in every importing file, so its unresolved import ref fell through to the exact-name matcher, which scored all K same-named import nodes via findBestMatch — K refs x K candidates = O(K^2) per package, producing only meaningless import->import edges. Fix: exclude `import`-kind nodes as name-match targets (they're statements, not definitions; real import->definition resolution is the import resolver's job). Plus two safe constant-factor wins in findBestMatch: hoist the per-candidate ref.filePath split, and skip cross-language candidates when a same-language one exists (provably the same winner — same-language scores >=50, cross-language maxes at 35). Measured: superset (Py+TS) candidates scored 7.5M -> 833K (9x), non-import edges preserved (+1618 now resolve to real defs), ~22K useless import->import edges removed; kubernetes (Go) computePathProximity 37.2s -> 5.0s; synthetic 8k-file mixed repo (K=4000) resolution 16.0s -> 1.7s. Full suite green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct stale better-sqlite3/wasm references to node:sqlite The SQLite backend has been Node's built-in node:sqlite (real SQLite, WAL + FTS5, from the bundled runtime) for a while — there is no native build step and no node-sqlite3-wasm fallback. README and the docs site were already updated; this catches the stragglers: - CLAUDE.md: the src/db/ backend description and the sqlite-backend test note. - src/db/index.ts, src/mcp/tools.ts: two code comments that still blamed "the wasm backend" for non-WAL behavior (reworded to "when WAL isn't in effect"). Leaves tree-sitter grammar wasm (web-tree-sitter / --liftoff-only) untouched — that's a different, still-current use of wasm. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(telemetry): drop the dead sqlite_backend field (schema v2) node:sqlite is now the only backend, so the `index` event's `sqlite_backend` field was a constant ("native") carrying no signal — and the `install` event never actually sent it. Remove the field and the backendKind() helper, bump the telemetry SCHEMA_VERSION 1 -> 2, and update TELEMETRY.md + docs/design/telemetry.md. The ingest worker is deliberately left tolerant: `index` doesn't require the field and schema_version validates as nonNegInt(99), so v2 events ingest fine and old clients still sending v1 + sqlite_backend keep validating too. Added a legacy comment there explaining it's safe to drop once old-client share is negligible. telemetry.test.ts: the assertion pinning schema_version and a stale-claim fixture line updated 1 -> 2. All telemetry tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
87 lines
4.3 KiB
Markdown
87 lines
4.3 KiB
Markdown
# Telemetry
|
|
|
|
CodeGraph collects a small set of **anonymous usage statistics** — which commands and
|
|
tools get used, which languages get indexed, which agents drive usage — so we can tell
|
|
which of the 20+ languages and 8 agent integrations deserve the most work. This page is
|
|
the complete list of what is collected. If a field isn't on this page, it isn't collected;
|
|
the ingest endpoint enforces this list as an allowlist and is itself
|
|
[public, auditable code](telemetry-worker/) in this repository.
|
|
|
|
## Turning it off
|
|
|
|
Any of these works, permanently:
|
|
|
|
```bash
|
|
codegraph telemetry off # stores your choice (and deletes any unsent data)
|
|
```
|
|
|
|
```bash
|
|
export CODEGRAPH_TELEMETRY=0 # per-shell / per-CI override
|
|
export DO_NOT_TRACK=1 # the cross-tool standard — always honored
|
|
```
|
|
|
|
`codegraph telemetry status` shows the current state, what decided it, and your machine ID.
|
|
The interactive installer (`codegraph install`) asks up front with a visible default-on
|
|
toggle and never re-asks. If you never saw the installer (e.g. `npx` straight into `init`),
|
|
a one-line notice is printed to stderr before the first time anything is sent.
|
|
|
|
Off means off: when disabled, CodeGraph records nothing, opens no connection to the
|
|
telemetry endpoint, and sends no "opted out" ping.
|
|
|
|
## What is collected
|
|
|
|
Every payload carries this envelope:
|
|
|
|
| field | example | notes |
|
|
|---|---|---|
|
|
| `machine_id` | `b3a8c1…` | random UUID minted on first send — derived from nothing |
|
|
| `codegraph_version` | `0.9.9` | |
|
|
| `os` / `arch` | `darwin` / `arm64` | platform identifiers only |
|
|
| `node_major` | `22` | major version only |
|
|
| `ci` | `false` | whether the `CI` env var was set |
|
|
| `schema_version` | `2` | bumped when this page changes (v2 dropped the `index` event's `sqlite_backend` field) |
|
|
|
|
And one of four events:
|
|
|
|
- **`install`** — when `codegraph install` configures agents: which agents
|
|
(`["claude","cursor",…]`), global vs project-local, and whether it was a fresh install,
|
|
an upgrade, or a re-run.
|
|
- **`index`** — when a full index completes: the **language names** present (e.g.
|
|
`["typescript","go"]`), the file count as a **coarse bucket** (`<100`, `100-1k`,
|
|
`1k-10k`, `10k+`), and the duration as a bucket (`<10s`, `10-60s`, `1-5m`, `5m+`).
|
|
- **`usage_rollup`** — one line per day per tool: the tool or CLI command **name** (e.g.
|
|
`codegraph_explore`, `init`), how many times it ran, how many errored, and — for MCP
|
|
tools — the connecting agent's name and version from the MCP handshake (e.g.
|
|
`Claude Code 2.1`).
|
|
- **`uninstall`** — when `codegraph uninstall`/`uninit` runs: which agents were removed.
|
|
|
|
Usage is **aggregated locally into daily totals** before anything is sent — there is no
|
|
per-call event stream, and nothing is sent in real time.
|
|
|
|
## What is never collected
|
|
|
|
- **No source code.** No file paths, file names, directory names, repository names or
|
|
URLs, symbol names, search queries, or anything else derived from the contents of an
|
|
indexed project.
|
|
- **No IP addresses.** The ingest endpoint never reads, logs, or forwards the client IP,
|
|
and IP discarding is enabled at the analytics backend on top of that. No geolocation.
|
|
- **No fingerprinting.** The machine ID is a random UUID stored in
|
|
`~/.codegraph/telemetry.json` — delete that file (or run `codegraph telemetry off`,
|
|
then `on`) and the old ID is gone forever, with no way to reconnect it.
|
|
- **No personal data.** No usernames, hostnames, emails, or environment variables.
|
|
|
|
## How it travels
|
|
|
|
Events POST to `telemetry.getcodegraph.com` — a first-party endpoint whose complete
|
|
source lives in [`telemetry-worker/`](telemetry-worker/) in this repository. It validates
|
|
every event and property against the allowlist above (anything else is dropped), strips
|
|
IPs, rate-limits, and forwards to a managed analytics store (PostHog, US region) as
|
|
anonymous events. Sends are fire-and-forget with a short timeout: offline or air-gapped
|
|
machines buffer a bounded local file (256 KB cap) and never retry-loop, log errors, or
|
|
slow a command down. Telemetry never adds latency to MCP tool calls — recording is an
|
|
in-memory counter.
|
|
|
|
The engineering contract behind all of this — including the rule that schema changes must
|
|
update this page, the client, and the public endpoint in one PR — is in
|
|
[`docs/design/telemetry.md`](docs/design/telemetry.md).
|