Commit Graph
369 Commits
Author SHA1 Message Date
Colby McHenry 3a44d5c4d1 fix: Improve CLI progress display and prevent tree-sitter WASM memory crashes
Replaces fixed-width padding with terminal escape sequences for proper progress line clearing across different terminal widths. Adds periodic parser reset every 5000 parses per language to prevent WASM heap fragmentation that causes "memory access out of bounds" crashes in large repositories. Includes filename truncation to fit available terminal width.
2026-04-04 10:28:07 -05:00
Colby McHenry e4908e1270 feat: Add database schema v3 with optimized node lookups and improved error handling
Adds expression index on lower(name) for memory-efficient case-insensitive searches, replacing in-memory caches that caused OOM on large codebases. Includes batched reference resolution, enhanced error reporting with detailed breakdown by error type, and improved CLI progress display for scanning phases.
2026-04-04 10:19:08 -05:00
Colby MchenryandGitHub 9cd5ef9870 Merge pull request #78 from colbymchenry/fix/explore-depth-and-qualified-symbol-lookup
fix: Improve explore depth and qualified symbol lookups
2026-04-03 19:58:34 -05:00
Colby McHenryandClaude Opus 4.6 b986b78fa9 fix: Increase explore traversal depth and support qualified symbol lookups
Two fixes discovered while benchmarking Swift (Alamofire):

1. codegraph_explore traversalDepth 2→3: Deep call chains (e.g., Alamofire's
   9-step Session.request()→URLSession flow) couldn't be followed in a single
   explore call, forcing agents to fall back to file reads.

2. findSymbol/findAllSymbols now support "Parent.child" notation (e.g.,
   "Session.request") by matching against qualified names (::Parent::child).
   Previously only checked node.name === symbol, which never matched qualified
   queries since node names are unqualified.

Also adds Alamofire Swift benchmark data to README (91% fewer tool calls,
78% faster with CodeGraph).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 19:56:58 -05:00
Colby McHenry 0d6f460b15 fix: Reduce codegraph_explore output limit to stay under MCP client token limits
Decreases maximum output from 50,000 to 35,000 characters to prevent exceeding ~10k token limits in MCP clients that could cause truncation or errors when processing exploration results.
2026-04-03 19:34:56 -05:00
Colby MchenryandGitHub eefa622965 Merge pull request #77 from colbymchenry/refactor/extract-language-configs
refactor: Extract per-language configs from tree-sitter.ts
2026-04-03 19:26:32 -05:00
Colby McHenryandClaude Opus 4.6 c8407ad007 refactor: Extract per-language configs and standalone extractors from tree-sitter.ts
Splits the monolithic tree-sitter.ts (3,358 lines) into modular files:
- 14 language config files under src/extraction/languages/
- 3 standalone extractors (Liquid, Svelte, DFM)
- Shared helpers and types modules to avoid circular imports

Also fixes a bug where Java's extractImport hook incorrectly set
handledRefs: true, preventing unresolved reference creation and
degrading codegraph_explore results for Java codebases.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 19:25:59 -05:00
Colby McHenry 0d63166f9d docs: Update benchmark results with comprehensive multi-codebase testing data
Replaces limited 3-test benchmark with results from 4 real-world codebases (VS Code, Excalidraw, Claude Code) showing 94% fewer tool calls and 77% faster exploration. Updates performance claims and adds detailed breakdown of tool usage patterns with and without CodeGraph.
2026-04-03 17:23:21 -05:00
Colby McHenry 2edc939245 fix: Handle gitignored project directories in git file detection
When a project directory is gitignored by a parent git repository, `git ls-files` returns no results even though files exist. Added detection for this scenario using `git rev-parse` and `git check-ignore` to fall back to filesystem walking when the project directory is ignored by an ancestor repo.
2026-04-03 17:17:31 -05:00
Colby McHenry 4d65d60ded feat: Update Claude instructions to discourage direct codegraph tool usage in main session
Changes guidance to recommend spawning Explore agents for exploration questions instead of using codegraph_explore/codegraph_context directly in main session to avoid filling up context with large amounts of source code. Adds completeness signal to codegraph_explore output so agents know not to re-read files that already have source code included.
2026-04-03 17:03:03 -05:00
Colby McHenry b927492bc0 feat: Add codegraph_explore tool for comprehensive single-call code exploration
Introduces a new MCP tool that performs deep code exploration in a single call, returning comprehensive context with full source code sections grouped by file and relationship mapping. Designed to replace multiple codegraph_node + file read operations for thorough understanding of code topics. Updates documentation to position explore as the primary tool for deep exploration questions.
2026-04-03 16:35:27 -05:00
Colby McHenry 4af51f565b feat: Extract type references from annotations and improve symbol query matching
Adds type annotation parsing to create references edges for parameter types, return types, and variable type annotations in TypeScript and other typed languages. Expands symbol extraction from queries to capture lowercase identifiers and filters out more common English words. Removes obsolete search utility tests.
2026-04-03 16:15:44 -05:00
Colby McHenry d575c945d9 chore: Add test_frameworks to .gitignore 2026-04-03 15:03:01 -05:00
Colby MchenryandGitHub f98dadc2c2 Merge pull request #76 from colbymchenry/fix/liquid-callers-and-context-relevance
fix: Fix Liquid template callers and context relevance
2026-04-03 13:31:55 -05:00
Colby McHenryandClaude Opus 4.6 68ec482bf4 fix: Fix callers/callees for Liquid templates and improve context relevance
Three issues discovered testing CodeGraph against a Shopify Liquid theme:

1. Callers/callees only traversed 'calls' edges, missing 'references' and
   'imports' edges that Liquid extraction creates for {% render %} and
   {% section %} tags. Expanded edge filter in getCallers/getCallees.

2. Context builder only ran text search as a fallback when semantic search
   returned nothing. For template-heavy codebases, semantic search returns
   irrelevant results (e.g., "Toast" for a header navigation query) while
   text/path-based matching would find the right files. Now always runs
   text search alongside semantic search with multi-term boosting.

3. MCP findAllSymbols only matched nodes by exact name, missing file nodes
   whose basename (without extension) matched the symbol. This caused
   callers to find zero results even with correct edges, since references
   edges point to file nodes (e.g., "product-card.liquid") not component
   nodes (e.g., "product-card").

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 13:31:14 -05:00
Colby MchenryandGitHub 47ed47bf04 Merge pull request #75 from colbymchenry/fix/python-resolution-and-context-relevance
feat: Improve Java/Python resolution, context relevance, and multi-symbol aggregation
2026-04-03 12:31:37 -05:00
Colby McHenryandClaude Opus 4.6 d04a911309 feat: Improve Java extraction, multi-symbol aggregation, and impact traversal
Java extraction:
- Handle Java method_invocation AST (receiver.method pattern)
- Support Java extends_interfaces and super_interfaces with type_list
- Create unresolved references for Java imports for cross-file resolution
- Extract interface inheritance via extractInheritance

MCP tools:
- Aggregate callers/callees/impact across ALL matching symbols (e.g. multiple
  overloads or same-named methods in different classes)
- New findAllSymbols() helper for multi-symbol lookup

Graph traversal:
- Impact analysis now traverses into container children (class → methods)
  so that callers of methods appear in the impact radius of their class

Other:
- Add deleteSpecificResolvedReferences() for precise cleanup after resolution
- Add 'instance-method' to resolvedBy union type
- Version bump to 0.6.8

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 12:30:28 -05:00
Colby McHenryandClaude Opus 4.6 8b541be894 fix: Improve Python resolution accuracy and context relevance
Eliminate cross-language false positives in name resolution and deprioritize
test files in context building. Benchmarked on a Python+Rust codebase where
37% of edges were false positives from Python built-in methods resolving to
Rust functions (e.g., list.extend → Rust extend).

Resolution fixes (index-time):
- Filter Python built-in type method calls (list.extend, dict.update, etc.)
- Filter bare Python built-in method names (append, extend, pop, keys, etc.)
- Add language boundary checks to matchMethodCall strategies 1, 2, and 3
- Penalize cross-language matches: -80 points in findBestMatch (was 0)
- Reduce confidence for single cross-language exact matches (0.5 vs 0.9)
- Prefer same-language candidates in matchFuzzy

Context relevance fixes (query-time):
- Add isTestFile() utility detecting test files across Python/JS/TS/Go/Rust/Java
- Deprioritize test files in scorePathRelevance (-15 penalty)
- Reduce test file scores to 30% in context builder result merging
- Both skip deprioritization when query mentions "test" or "spec"

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 12:29:21 -05:00
Colby MchenryandGitHub 5d5e715dec Merge pull request #74 from colbymchenry/feat/always-enable-embeddings
feat: Always enable embeddings, remove config toggle
2026-04-03 10:05:12 -05:00
Colby McHenryandClaude Opus 4.6 7e6914ddff feat: Always enable embeddings, remove enableEmbeddings config option
Testing showed semantic search produces significantly better results for
natural language queries that Claude writes. FTS alone often ranks
properties above their parent classes and misses conceptual matches.
Embeddings are now always on — the vector manager is created eagerly,
with model download and embedding generation still happening lazily.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 10:03:41 -05:00
Colby McHenry 8b5622d346 Remove Sentry telemetry system from CodeGraph
Eliminates anonymous error reporting functionality that was collecting stack traces and error context via Sentry. Removes all telemetry-related code, configuration options, and documentation references.
2026-04-03 09:54:07 -05:00
Colby McHenry 9bb8e96bd4 bump version to 0.6.6 2026-04-01 14:37:38 -05:00
Colby MchenryandGitHub 2de4d1fd4c Merge pull request #73 from colbymchenry/fix/cross-module-resolution
fix: Prevent false cross-module edges in name-based resolution
2026-04-01 14:30:21 -05:00
Colby McHenryandClaude Opus 4.6 584bd94ecc fix: Prevent false cross-module edges in name-based resolution
Name matching was creating false `calls` edges between unrelated modules
in monorepos because `findBestMatch()` had no concept of directory
proximity — functions with common names (e.g. `navigate`) in different
apps scored identically and resolved to whichever came first.

Adds path proximity scoring (shared directory segments) so same-module
candidates strongly win over cross-boundary ones, and lowers confidence
for distant matches so import-based resolution takes precedence.

Closes #67

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 14:23:13 -05:00
Colby MchenryandGitHub 960bed2253 Merge pull request #72 from colbymchenry/fix/telemetry-opt-out
fix: Add telemetry opt-out for Sentry error reporting
2026-04-01 14:15:08 -05:00
Colby McHenryandClaude Opus 4.6 5a185eb736 fix: Add telemetry opt-out via installer prompt and env var
Sentry error reporting can now be disabled by:
1. Declining during the interactive installer (sets CODEGRAPH_TELEMETRY=off
   in the MCP server config env)
2. Setting CODEGRAPH_TELEMETRY=off in your shell environment

README updated with a Telemetry section documenting what is collected
and how to opt out.

Closes #68

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 14:14:33 -05:00
Colby McHenry c401a96d17 Merge branch 'main' into fix/telemetry-opt-out 2026-04-01 14:12:28 -05:00
Colby MchenryandGitHub 61a9961bb8 Merge pull request #71 from colbymchenry/fix/installer-global-install-prompt
fix: Prompt before global npm install during installer
2026-04-01 14:12:08 -05:00
Colby McHenryandClaude Opus 4.6 12b0414900 fix: Prompt before global npm install during installer
The installer previously ran `npm install -g` silently without user
consent. Now it asks for confirmation first, explains why the global
install is needed (hooks & MCP server), and gracefully skips if declined.
README updated to document this step.

Closes #69

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 14:10:01 -05:00
Colby MchenryandGitHub 6647d1c827 Merge pull request #70 from colbymchenry/feat/improved-search-tokenization
feat: Improve search tokenization with camelCase splitting
2026-04-01 13:50:55 -05:00
Colby McHenryandClaude Opus 4.6 0756636bde feat: Improve search tokenization with camelCase splitting and code-aware stop words
extractSearchTerms now splits camelCase, PascalCase, snake_case, and
dot.notation into individual tokens (e.g. "getUserName" → ["user", "name"]).
Stop words expanded with code-specific noise words (code, file, function,
method, class, type, etc.) to improve search precision.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 13:50:12 -05:00
Colby McHenryandClaude Opus 4.6 8f5f88b813 refactor: Remove AI entry point guessing, use search-driven flow
AI picking the entry point was unreliable. Now:
- Type a symbol name → dropdown shows matches
- Click a result (or Enter to pick first) → traces its call graph depth 3
- The user picks the starting point, the graph does the rest deterministically

No more AI guesswork. Search + graph traversal = reliable flows.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 17:22:35 -05:00
Colby McHenryandClaude Opus 4.6 dcd4fa397e refactor: Simplify to entry-point + call graph tracing
Completely reworked the explore approach:
- Claude (or keyword search) finds ONE entry point, not a list of symbols
- getCallGraph(entry, depth=3) traces the actual call chain deterministically
- No more AI-guessed symbol lists, bridge passes, or relevance filtering
- Search result clicks also trace the full call graph from that point

The graph data was always accurate — the problem was AI trying to guess
the whole flow. Now AI just finds the starting point, graph does the rest.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 17:19:23 -05:00
Colby McHenryandClaude Opus 4.6 3d2b38918e refine: Tighten Claude prompt for focused 5-8 node execution paths
Stricter prompt rules: only symbols directly in the execution path,
every symbol must call or be called by the next, no tangential features.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 16:55:52 -05:00
Colby McHenryandClaude Opus 4.6 2278d3fd6f feat: Flow-oriented exploration with entry point identification
Update Claude prompt to identify the entry point and return symbols in
execution order. The graph now centers on the entry point and auto-opens
its detail panel, giving users a clear starting point to trace the flow.

- Claude returns {entry, flow} instead of flat array
- Entry point is auto-selected and centered on load
- Detail panel opens immediately for the entry point
- Prompt asks for max 8-10 symbols in execution order

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 16:52:10 -05:00
Colby McHenry 61d43fa26c Merge remote-tracking branch 'origin/main' 2026-03-22 16:50:10 -05:00
Colby McHenryandClaude Opus 4.6 c433e7d1a0 refactor: Trust Claude's seed picks, only add bridge nodes
Instead of expanding all callers/callees of seeds (which pulls in noise
from hub nodes like getSession), now:
1. Find direct edges between Claude's seeds
2. Only add non-seed nodes if they bridge 2+ isolated seeds
3. Cross-connection pass discovers hidden edges between result nodes
4. No more unrelated callers of hub nodes polluting the graph

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 16:48:55 -05:00
Colby McHenryandClaude Opus 4.6 ba30c74461 feat: Improve visualizer graph quality and UI
- Bridge pass: connect isolated seed nodes that share callees
- Kind labels on nodes (fn, class, comp, etc.) for quick identification
- Quick action buttons in detail panel (Expand Callees, Callers, Call Graph, Impact)
- Wider detail panel (460px) for better code readability
- Multiline node labels showing name + kind

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 16:46:09 -05:00
Colby McHenryandClaude Opus 4.6 43ea0a40ba feat: Add interactive graph visualization with Claude-powered exploration
Adds `codegraph visualize` command that launches a localhost web UI for
visually exploring code relationships. Users can ask natural language
questions like "how does authentication work?" and see the relevant code
flow rendered as an interactive graph.

Key components:
- Visualizer HTTP server (src/visualizer/server.ts) with REST API
- Single-page frontend with Cytoscape.js graph + highlight.js code preview
- Claude CLI integration for intelligent query interpretation
- Dark theme, right-click context menus, keyboard shortcuts
- Detail panel with source code, callers, callees, hierarchy

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 16:43:57 -05:00
Colby MchenryandGitHub 3729132b0d Merge pull request #49 from markhu/main
exclude .pio/ folder for Platform.io IoT libs
2026-03-18 22:55:54 -05:00
Colby McHenry 1f715cd6d9 chore: Bump version to 0.6.5 2026-03-18 17:10:59 -05:00
Colby McHenry 5334a0f023 feat: Add codegraph affected command to find test files impacted by changes
Traverses dependency graph to identify which test files depend on changed source files. Supports stdin input for git integration, custom test file patterns, and configurable traversal depth. Useful for targeted test execution in CI/CD pipelines.
2026-03-18 16:31:54 -05:00
Colby MchenryandGitHub d20021a322 Merge pull request #63 from colbymchenry/fix/stale-lock-mcp-retry
fix: Stale lock recovery and MCP init retry
2026-03-18 14:52:59 -05:00
Colby McHenry 3ec4cde542 chore: Reduce stale lock timeout from 10 minutes to 2 minutes 2026-03-18 14:52:20 -05:00
Colby McHenry b964d5909a fix: Stale lock recovery and MCP init retry
Fixes #47 — "database is locked" after crash and MCP "not initialized"
when project IS initialized.

- FileLock: treat locks older than 10 minutes as stale regardless of PID
  status, covering cases where PID was reused or kill signal check fails
- MCP server: log errors from tryInitializeDefault() to stderr instead of
  silently swallowing, so transient open failures are diagnosable
- MCP server: retryInitIfNeeded() properly cleans up failed instances
  before retrying, preventing resource leaks
- CLI: add 'codegraph unlock' command for manual lock file removal
2026-03-18 14:50:25 -05:00
Colby MchenryandGitHub bdbe59b457 Merge pull request #62 from colbymchenry/fix/fk-constraint-empty-names
fix: Prevent FK constraint failure from nodes with empty names
2026-03-18 14:36:16 -05:00
Colby MchenryandGitHub 3ffcac781e Merge pull request #61 from colbymchenry/fix/wasm-oom-lazy-grammars
fix: Lazy grammar loading to prevent V8 WASM OOM on large codebases
2026-03-18 14:36:04 -05:00
Colby McHenry 92631d53d3 fix: Prevent FK constraint failure from nodes with empty names
Fixes #42 — tree-sitter can produce nodes with empty names (e.g. from
complex C/C++ declarators in header files). These nodes were silently
skipped at DB insert time, but their containment edges were still
inserted, causing a FOREIGN KEY constraint violation that crashed
indexing.

Two-layer fix:
- createNode() now returns null for empty names, preventing the node
  and its edges from ever being created (Option A)
- storeExtractionResult() filters edges and unresolved refs to only
  reference nodes that passed validation, as a safety net (Option B)
2026-03-18 14:34:25 -05:00
Colby McHenry 15b5e56322 fix: Lazy grammar loading and quantized embeddings to prevent V8 WASM OOM
Fixes #54 — `codegraph init -i` crashes with "Fatal process out of memory: Zone"
on large codebases because all 16 tree-sitter WASM grammar modules were compiled
upfront by V8, exhausting the WASM Zone allocator.

Changes:
- initGrammars() now only initializes the tree-sitter WASM runtime (Parser.init()),
  no longer eagerly loads all grammar files
- New loadGrammarsForLanguages() loads only grammars for languages actually present
  in the project (e.g. a Dart project loads ~2-3 grammars instead of 16)
- Orchestrator detects needed languages after file scan, before parsing begins
- Embedding pipeline now uses quantized model (~67MB vs ~270MB) to further reduce
  WASM memory pressure when embeddings are enabled
2026-03-18 14:21:54 -05:00
Mark HandGitHub eb25a166d1 exclude .pio/ folder for Platform.io IoT libs 2026-02-24 19:11:10 -08:00