Files
codegraph/__tests__/identifier-segments.test.ts
T
e699ee9686 feat(prompt-hook): graph-derived gate tier + confidence-tiered injection + gate telemetry (#1136)
The keyword gate (#1126) can never know a repo's domain nouns. This adds
the graph-derived tier the design discussion converged on: symbol names
are split into prose segments at index time (name_segment_vocab, riding
the insertNode write path), and the hook verifies a prompt's plain words
against them — "the state machine des commandes" → OrderStateMachine, in
any language whose technical nouns are Latin script.

Confidence now decides HOW MUCH to inject, not just whether:
- HIGH (keyword, or index-verified code token): full explore injection,
  unchanged — the validated adoption lever.
- MEDIUM (segment matches only): a ~500-byte pointer naming the matching
  symbols; the AGENT writes the explore query. Never runs explore, so a
  fuzzy match can't inject 16KB of wrong-feature context.
- Silent otherwise, as before.

Precision is derived from the repo's own naming statistics plus measured
FP fixes: co-occurrence (≥2 words on one name) always qualifies; a single
word must be ≥5 chars, cluster across 2–25 names (singletons are prose
coincidence: "deploy to production" → matchesNonProductionDir), match a
multi-segment name, and not be an English function/filler word (the one
place a word list is honest: identifiers are English, so only English
prose collides). Every candidate is re-verified against nodes before
being surfaced — vocab rows are proposals, deletions leave orphans by
design, a full index rebuilds from scratch, and sync heals pre-upgrade
databases (batched + yielding; emptiness captured at sync ENTRY so the
sync's own writes can't mask the backfill).

Schema v7 migration is DDL-only (instant; none of the #1067 row-churn
hazards). Gate outcomes roll up as anonymous usage counters
(prompt-hook-gate-<outcome>, names only, never content) through the
existing telemetry pipeline — recall becomes measurable, and the counters
are the agreed kill-criterion data for ever revisiting a local classifier.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 14:35:38 -05:00

82 lines
3.7 KiB
TypeScript

import { describe, it, expect } from 'vitest';
import {
splitIdentifierSegments,
extractProseCandidates,
normalizeProseWord,
segmentLookupVariants,
} from '../src/search/identifier-segments';
describe('splitIdentifierSegments — symbol names → prose words', () => {
it('splits camelCase / PascalCase at humps', () => {
expect(splitIdentifierSegments('OrderStateMachine')).toEqual(['order', 'state', 'machine']);
expect(splitIdentifierSegments('userId')).toEqual(['user', 'id']);
});
it('handles acronym runs — HTML stays one segment', () => {
expect(splitIdentifierSegments('parseHTMLDocument')).toEqual(['parse', 'html', 'document']);
expect(splitIdentifierSegments('HTMLParser')).toEqual(['html', 'parser']);
});
it('keeps digits glued to their word', () => {
expect(splitIdentifierSegments('base64Encode')).toEqual(['base64', 'encode']);
expect(splitIdentifierSegments('parseHTML5Doc')).toEqual(['parse', 'html5', 'doc']);
});
it('splits snake_case, kebab-case, and dotted file names', () => {
expect(splitIdentifierSegments('snake_case_name')).toEqual(['snake', 'case', 'name']);
expect(splitIdentifierSegments('MAX_RETRY_COUNT')).toEqual(['max', 'retry', 'count']);
expect(splitIdentifierSegments('checkout.service.ts')).toEqual(['checkout', 'service', 'ts']);
expect(splitIdentifierSegments('state-machine')).toEqual(['state', 'machine']);
});
it('drops sub-minimum and digit-only fragments, dedupes', () => {
expect(splitIdentifierSegments('x')).toEqual([]);
expect(splitIdentifierSegments('42')).toEqual([]);
expect(splitIdentifierSegments('getData_getData')).toEqual(['get', 'data']);
});
});
describe('extractProseCandidates — prompt prose → lookup words', () => {
it('keeps content words, drops short function words, in any Latin language', () => {
expect(extractProseCandidates('comment marche la state machine des commandes ?')).toEqual([
'comment', 'marche', 'state', 'machine', 'commandes',
]);
});
it('strips diacritics so loanwords meet ASCII identifier segments', () => {
expect(extractProseCandidates('la résolution des références')).toEqual(['resolution', 'references']);
expect(normalizeProseWord('Übersicht')).toBe('ubersicht');
});
it("splits on apostrophes — l'architecture keeps the noun", () => {
expect(extractProseCandidates("explique l'architecture du module de stock")).toEqual([
'explique', 'architecture', 'module', 'stock',
]);
});
it('caps candidates and skips unsegmented-script sentence runs', () => {
const many = Array.from({ length: 25 }, (_, i) => `distinctword${String.fromCharCode(97 + i)}`).join(' ');
expect(extractProseCandidates(many)).toHaveLength(16);
// A no-spaces CJK sentence is one giant run — over the length ceiling, skipped.
expect(extractProseCandidates('請解釋一下這個訂單狀態機的整體運作流程與架構設計方式')).toEqual([]);
// Short CJK runs pass through as candidates — no script filter; the graph
// verification tier rejects them (identifiers are almost never CJK).
expect(extractProseCandidates('修复这个拼写错误')).toEqual(['修复这个拼写错误']);
});
it('drops digit-only and sub-4-char words', () => {
expect(extractProseCandidates('fix the bug in v2 at 1234')).toEqual([]);
});
});
describe('segmentLookupVariants — light plural folding', () => {
it('folds trailing s/es so plurals hit singular segments', () => {
expect(segmentLookupVariants('services')).toContain('service');
expect(segmentLookupVariants('machines')).toContain('machine');
});
it('never strips a word below the minimum', () => {
expect(segmentLookupVariants('bus')).toEqual(['bus']);
});
});