fix(prompt-hook): fire front-load hook for non-English prompts (#994) (#1004)

The UserPromptSubmit hook's structural-prompt gate was English-only, so a
structural question written in Chinese — or any non-Latin script — silently
injected nothing: JS `\b` is ASCII-only and never matches between Han
characters, so the keyword regex couldn't fire (and couldn't be extended in
place). To the user the hook looked unwired, with no error to explain why.

Make the gate language-aware, split into tested helpers in directory.ts:
- hasStructuralKeyword: English (\b-guarded) + CJK structural keywords.
- extractCodeTokens: identifier-shaped tokens (camelCase / snake_case /
  name() / a.b) in any language — verified against the index via
  getNodesByName before firing, so a tech brand like `JavaScript` that looks
  like a symbol but isn't one here doesn't inject ~16KB of spurious context.
- isStructuralPrompt: the cheap candidate gate (keyword OR code-token).

Adds 21 unit tests for the gate (previously untested) covering the reporter's
verification table plus the false-positive guards.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Colby Mchenry
2026-06-26 18:39:34 -05:00
committed by GitHub
co-authored by Claude Opus 4.8
parent 703629edc3
commit b45f309a1b
4 changed files with 174 additions and 7 deletions
+16 -6
View File
@@ -26,7 +26,7 @@
import { Command } from 'commander';
import * as path from 'path';
import * as fs from 'fs';
import { getCodeGraphDir, isInitialized, unsafeIndexRootReason, findNearestCodeGraphRoot, planFrontload } from '../directory';
import { getCodeGraphDir, isInitialized, unsafeIndexRootReason, findNearestCodeGraphRoot, planFrontload, hasStructuralKeyword, extractCodeTokens } from '../directory';
import { detectWorktreeIndexMismatch, worktreeMismatchWarning } from '../sync/worktree';
import { createShimmerProgress } from '../ui/shimmer-progress';
import { getGlyphs } from '../ui/glyphs';
@@ -1053,11 +1053,15 @@ program
try { input = JSON.parse(raw); } catch { return; }
const prompt = String(input.prompt || '');
// Gate: only structural / flow / impact / where-how prompts get context.
// A cheap regex keeps every other prompt ("fix this typo") a zero-cost
// no-op so we never add latency where there's no structural answer to give.
const STRUCTURAL = /\b(how|where|trace|flow|path|reach(?:es|ed)?|call(?:s|ed|er|ers|ee)?|depend|impact|affect|wired?|connect|implement|architect|structure|breaks?|what calls|why does)\b/i;
if (!prompt || !STRUCTURAL.test(prompt)) return;
// Gate: only structural / flow / impact / where-how prompts get context, so
// every other prompt ("fix this typo") stays a zero-cost no-op. Language-aware
// (English + CJK keywords, plus code-shaped tokens) so it fires for non-English
// prompts too (issue #994). A keyword fires on its own; a code-token is only a
// CANDIDATE — verified against the graph below, so a tech brand ("JavaScript")
// that looks like a symbol but isn't one here doesn't inject spurious context.
const keyworded = hasStructuralKeyword(prompt);
const codeTokens = keyworded ? [] : extractCodeTokens(prompt);
if (!keyworded && codeTokens.length === 0) return;
// Decide what to inject, shaped by WHERE the index(es) are: the nearest
// indexed ancestor of cwd, or — when cwd is an un-indexed workspace root
@@ -1079,6 +1083,12 @@ program
const { default: CodeGraph } = await loadCodeGraph();
const cg = await CodeGraph.open(plan.exploreRoot);
try {
// Code-token-only prompt: require that at least one token is a REAL symbol
// in THIS index before front-loading. Without it, a brand name or common
// word that merely looks like code ("JavaScript", "GitHub") would run
// explore and inject ~16KB of low-relevance context (issue #994 follow-up).
// A keyword-bearing prompt skips this — the keyword is signal enough.
if (!keyworded && !codeTokens.some((t) => cg.getNodesByName(t).length > 0)) return;
const { ToolHandler } = await import('../mcp/tools');
const handler = new ToolHandler(cg);
const result = await handler.execute('codegraph_explore', { query: prompt });
+80
View File
@@ -233,6 +233,86 @@ export function findIndexedSubprojectRoots(
return out;
}
/**
* English structural keywords, matched with `\b` word boundaries so a keyword
* inside a longer word doesn't false-positive ("flow" in "flower").
*/
const STRUCTURAL_EN = /\b(how|where|trace|flow|path|reach(?:es|ed)?|call(?:s|ed|er|ers|ee)?|depend|impact|affect|wired?|connect|implement|architect|structure|breaks?|what calls|why does)\b/i;
/**
* Non-English (CJK) structural keywords, matched WITHOUT `\b`. JS's `\b` is
* ASCII-only — it only fires at `[A-Za-z0-9_]` boundaries, never between Han
* characters — so a Chinese keyword wrapped in `\b…\b` could never match. That
* was issue #994: the English-only gate silently no-op'd every Chinese prompt,
* so non-English users got no front-load nudge and no error to explain why. The
* set mirrors the English intent (如何=how, 在哪/哪里=where, 流程/流向=flow,
* 路径=path, 调用=call, 依赖=depend, 影响=impact/affect, 实现=implement,
* 架构=architect, 结构=structure, 追踪/跟踪=trace) plus structural-overview words
* with no single clean English equivalent (介绍/解析/分析/原理/机制).
*/
const STRUCTURAL_CJK = /如何|怎么|在哪|哪里|追踪|跟踪|流程|流向|路径|调用|依赖|影响|实现|架构|结构|介绍|解析|分析|原理|机制/;
/** Doc/data/asset file extensions — a `name.ext` of this kind is a file
* reference, not a code symbol, so it must not trip the member-access signal. */
const DOC_DATA_EXT = /\.(md|markdown|txt|rst|json|ya?ml|toml|lock|csv|tsv|log|ini|cfg|conf|env|xml|html?|png|jpe?g|gif|svg|pdf)$/i;
/**
* Does `prompt` contain an explicit structural keyword (English or CJK)? A
* keyword is a strong, self-contained signal, so the front-load hook fires on it
* directly — no graph check needed. (A *code-token* match, by contrast, is only
* a candidate the hook verifies against the graph first; see {@link extractCodeTokens}.)
*/
export function hasStructuralKeyword(prompt: string): boolean {
return !!prompt && (STRUCTURAL_EN.test(prompt) || STRUCTURAL_CJK.test(prompt));
}
/**
* Identifier-shaped tokens in `prompt` — camelCase / PascalCase-with-inner-cap,
* snake_case, a `name(` call, or the two sides of an `a.b` member access. Naming
* a symbol is a code question whatever the surrounding human language, and these
* shapes almost never occur in ordinary prose, so they catch the common
* "<symbol> 的调用链?" / "where is <symbol> 定義" prompts no keyword list would.
*
* These are *candidates*, not a verdict: a tech brand like `JavaScript` or
* `GitHub` is identifier-shaped too, so the front-load hook checks each token
* against the actual index ({@link getNodesByName}) and only fires when one is a
* real symbol here — otherwise a brand-name prompt would inject ~16KB of
* low-relevance context (issue #994 follow-up). A doc/data filename ("README.md")
* is excluded from the member-access form since it's a file reference, not a symbol.
*/
export function extractCodeTokens(prompt: string): string[] {
if (!prompt) return [];
const out = new Set<string>();
// camelCase / PascalCase-with-inner-cap (getUserId, parseToken, UserService) or
// snake_case (article_publish, get_user) — a whole identifier run that has an
// inner lower→upper transition or an underscore flanked by alphanumerics.
for (const m of prompt.matchAll(/[A-Za-z_$][\w$]*/g)) {
const w = m[0];
if (/[a-z][A-Z]/.test(w) || /[A-Za-z0-9]_[A-Za-z0-9]/.test(w)) out.add(w);
}
// call form: an identifier directly before '(' — parseToken(, render(). No
// whitespace before '(' so prose like "the function (entry point)" doesn't trip it.
for (const m of prompt.matchAll(/([A-Za-z_$][\w$]*)\(/g)) out.add(m[1]!);
// member access on identifiers (user.login) — but not a doc/data filename.
for (const m of prompt.matchAll(/([A-Za-z_$][\w$]*)\.([A-Za-z_$][\w$]*)/g)) {
if (!DOC_DATA_EXT.test(m[0])) { out.add(m[1]!); out.add(m[2]!); }
}
return [...out];
}
/**
* Cheap, graph-free candidate gate for the front-load hook: could `prompt` be a
* structural / flow / impact / "where-how" question worth front-loading context
* for? True on an explicit keyword (English or CJK, issue #994) OR an
* identifier-shaped token. A keyword is sufficient to fire on its own; a
* token-only match is only a candidate the hook then verifies against the graph
* (a brand name like `JavaScript` is token-shaped but isn't a symbol). Every
* non-candidate prompt ("fix this typo", in any language) stays a zero-cost no-op.
*/
export function isStructuralPrompt(prompt: string): boolean {
return hasStructuralKeyword(prompt) || extractCodeTokens(prompt).length > 0;
}
/**
* What the front-load hook should do for a prompt issued from a directory.
*/