feat(search): field-qualified queries (kind:/lang:/path:/name:) + fuzzy typo fallback (#131)
* feat(search): field-qualified queries (kind:/lang:/path:/name:) + fuzzy typo fallback
Two UX improvements that turn a free-text search into something a
real user can drive precisely.
1) Field-qualified queries.
A new query parser (src/search/query-parser.ts) splits the raw query
into structured filters and a free-text remainder:
kind:function name:auth path:src/api authenticate
becomes
{ kinds: ['function'], nameFilters: ['auth'],
pathFilters: ['src/api'], text: 'authenticate' }
Filters compose with the SearchOptions arg (intersection). Unknown
prefixes pass through as plain text so `query "TODO:"` keeps working.
Quoted values (`path:"my dir"`) handle whitespace. When the user
specifies only filters with no text, the search uses a filter-only
candidate scan instead of bailing out.
Recognised today:
kind: any NodeKind value
lang: any Language value (alias: language:)
path: case-insensitive substring of file_path
name: case-insensitive substring of node.name
2) Fuzzy fallback.
When BOTH FTS and LIKE return nothing AND the text is at least 3
chars, the resolver scans the distinct-name set with a bounded
Damerau-Levenshtein-style edit distance (≤2 for ≥5 chars, ≤1 for
4-char queries, off for shorter). Bounded edit-distance early-exits
once the row min exceeds maxDist, so this stays O(distinct-names *
avg-name-length) with a very low constant.
Verified live against ollama/ollama@v0.22.0:
query "kind:function auth" → only function-kind hits
query "lang:go path:server route" → Go files under server/
query "getUssr" (typo) → finds getUser, SetUser
query "confg" (typo) → finds Config
Full test suite: 380 passed.
* fix(search): address reviewer findings — tokenizer mid-token quotes, fuzzy fan-out cap, larger filter-only over-fetch, unit tests
Five fixes from independent review:
- parseQuery tokenizer: quotes that appear MID-token (path:"my dir/
file") were not being recognised — only quotes at the start of a
token were treated as quoted spans. The fixture path:"my dir"
parsed as ['path:"my', 'dir"'] instead of ['path:"my dir"'].
Tokeniser is now a single state machine that scans into a token
until whitespace OR a quote, and recognises quotes anywhere within
the token (skips to the matching close quote).
- searchNodesFuzzy: cap the per-name follow-up SQL queries at
Math.max(limit*2, 50) AFTER edit-distance filtering. Without
this, a project with many similar names (getUser1, getUser2...)
could fan out far beyond limit queries before the inner-loop
break kicks in.
- searchAllByFilters (filter-only no-text path): bumped over-fetch
multiplier from 2× to 5× so a selective post-filter (e.g.
path:src/very/specific/file.ts) doesn't return fewer than limit
results despite the DB having matches.
- 23 new unit tests in __tests__/search-query-parser.test.ts:
parseQuery covers known-field filter, lang/language alias,
multiple kind: ORs, quoted spans (incl. mid-token), URL
passthrough, empty-value passthrough, unknown prefix passthrough,
unknown value passthrough, all-filters-no-text, empty input,
20k-char input. boundedEditDistance covers identity, single
insertion/deletion/substitution, length-difference shortcut,
empty inputs, case-sensitivity, early-exit correctness.
Full test suite: 853 passed (up from 830).
* refactor(search): derive parser kind/lang sets from types.ts as const
Convert NodeKind and Language to runtime-iterable as const arrays
(NODE_KINDS, LANGUAGES) so the query parser imports the canonical
list instead of duplicating it. Also fix the path: JSDoc to say
substring (matches the .includes() impl).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Colby McHenry <me@colbymchenry.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
Colby McHenry
parent
38d155618f
commit
56f6b3b485
+157
-7
@@ -19,6 +19,7 @@ import {
|
||||
} from '../types';
|
||||
import { safeJsonParse } from '../utils';
|
||||
import { kindBonus, nameMatchBonus, scorePathRelevance } from '../search/query-utils';
|
||||
import { parseQuery, boundedEditDistance } from '../search/query-parser';
|
||||
|
||||
/**
|
||||
* Database row types (snake_case from SQLite)
|
||||
@@ -478,14 +479,51 @@ export class QueryBuilder {
|
||||
* 3. Score results based on match quality
|
||||
*/
|
||||
searchNodes(query: string, options: SearchOptions = {}): SearchResult[] {
|
||||
const { kinds, languages, limit = 100, offset = 0 } = options;
|
||||
const { limit = 100, offset = 0 } = options;
|
||||
|
||||
// Parse field-qualified bits out of the raw query (kind:, lang:,
|
||||
// path:, name:). Anything not recognised stays in `text` and goes
|
||||
// to FTS unchanged. Filters compose with the SearchOptions arg —
|
||||
// both are applied (intersection-style).
|
||||
const parsed = parseQuery(query);
|
||||
const mergedKinds =
|
||||
parsed.kinds.length > 0
|
||||
? Array.from(new Set([...(options.kinds ?? []), ...parsed.kinds]))
|
||||
: options.kinds;
|
||||
const mergedLanguages =
|
||||
parsed.languages.length > 0
|
||||
? Array.from(new Set([...(options.languages ?? []), ...parsed.languages]))
|
||||
: options.languages;
|
||||
const pathFilters = parsed.pathFilters;
|
||||
const nameFilters = parsed.nameFilters;
|
||||
// The text portion drives FTS/LIKE; if all the user typed was
|
||||
// filters (`kind:function`), we still need *some* candidate set,
|
||||
// so synthesise an empty-text path that returns everything matching
|
||||
// the filters.
|
||||
const text = parsed.text;
|
||||
const kinds = mergedKinds;
|
||||
const languages = mergedLanguages;
|
||||
|
||||
// First try FTS5 with prefix matching
|
||||
let results = this.searchNodesFTS(query, { kinds, languages, limit, offset });
|
||||
let results = text
|
||||
? this.searchNodesFTS(text, { kinds, languages, limit, offset })
|
||||
// Over-fetch by 5× when running filter-only (no text). The
|
||||
// post-scoring path: + name: filters can be very selective, so
|
||||
// a smaller multiplier risks returning fewer than `limit`
|
||||
// results despite the DB having plenty of matches.
|
||||
: this.searchAllByFilters({ kinds, languages, limit: limit * 5 });
|
||||
|
||||
// If no FTS results, try LIKE-based substring search
|
||||
if (results.length === 0 && query.length >= 2) {
|
||||
results = this.searchNodesLike(query, { kinds, languages, limit, offset });
|
||||
if (results.length === 0 && text.length >= 2) {
|
||||
results = this.searchNodesLike(text, { kinds, languages, limit, offset });
|
||||
}
|
||||
|
||||
// Final fuzzy fallback: scan all known names and keep those within
|
||||
// a tight Levenshtein distance. Only fires when both FTS and LIKE
|
||||
// returned nothing AND there's a text portion long enough to be
|
||||
// worth fuzzing (1-char queries would match too much).
|
||||
if (results.length === 0 && text.length >= 3) {
|
||||
results = this.searchNodesFuzzy(text, { kinds, languages, limit });
|
||||
}
|
||||
|
||||
// Supplement: ensure exact name matches are always candidates.
|
||||
@@ -521,13 +559,14 @@ export class QueryBuilder {
|
||||
}
|
||||
|
||||
// Apply multi-signal scoring
|
||||
if (results.length > 0 && query) {
|
||||
if (results.length > 0 && (text || query)) {
|
||||
const scoringQuery = text || query;
|
||||
results = results.map(r => ({
|
||||
...r,
|
||||
score: r.score
|
||||
+ kindBonus(r.node.kind)
|
||||
+ scorePathRelevance(r.node.filePath, query)
|
||||
+ nameMatchBonus(r.node.name, query),
|
||||
+ scorePathRelevance(r.node.filePath, scoringQuery)
|
||||
+ nameMatchBonus(r.node.name, scoringQuery),
|
||||
}));
|
||||
results.sort((a, b) => b.score - a.score);
|
||||
// Trim to requested limit after rescoring
|
||||
@@ -536,6 +575,117 @@ export class QueryBuilder {
|
||||
}
|
||||
}
|
||||
|
||||
// Apply path: + name: filters AFTER scoring. Scoring already uses
|
||||
// path/name as a soft signal; the explicit filters here are a hard
|
||||
// gate. Done last so the FTS limit fetched plenty of candidates to
|
||||
// narrow from.
|
||||
if (pathFilters.length > 0) {
|
||||
const lowered = pathFilters.map((p) => p.toLowerCase());
|
||||
results = results.filter((r) => {
|
||||
const fp = r.node.filePath.toLowerCase();
|
||||
return lowered.some((p) => fp.includes(p));
|
||||
});
|
||||
}
|
||||
if (nameFilters.length > 0) {
|
||||
const lowered = nameFilters.map((n) => n.toLowerCase());
|
||||
results = results.filter((r) => {
|
||||
const nm = r.node.name.toLowerCase();
|
||||
return lowered.some((n) => nm.includes(n));
|
||||
});
|
||||
}
|
||||
|
||||
return results;
|
||||
}
|
||||
|
||||
/**
|
||||
* Match-everything path used when the user supplied only field
|
||||
* filters (`kind:function lang:typescript`) with no text. Returns
|
||||
* candidates ordered by name; the caller's filter pass narrows to
|
||||
* what was asked for.
|
||||
*/
|
||||
private searchAllByFilters(options: {
|
||||
kinds?: NodeKind[];
|
||||
languages?: Language[];
|
||||
limit: number;
|
||||
}): SearchResult[] {
|
||||
const { kinds, languages, limit } = options;
|
||||
let sql = 'SELECT * FROM nodes WHERE 1=1';
|
||||
const params: (string | number)[] = [];
|
||||
if (kinds && kinds.length > 0) {
|
||||
sql += ` AND kind IN (${kinds.map(() => '?').join(',')})`;
|
||||
params.push(...kinds);
|
||||
}
|
||||
if (languages && languages.length > 0) {
|
||||
sql += ` AND language IN (${languages.map(() => '?').join(',')})`;
|
||||
params.push(...languages);
|
||||
}
|
||||
sql += ' ORDER BY name LIMIT ?';
|
||||
params.push(limit);
|
||||
const rows = this.db.prepare(sql).all(...params) as NodeRow[];
|
||||
return rows.map((row) => ({ node: rowToNode(row), score: 1 }));
|
||||
}
|
||||
|
||||
/**
|
||||
* Fuzzy fallback: when zero FTS/LIKE hits, try an edit-distance
|
||||
* sweep over the distinct symbol-name set. Caps `maxDist` at 2 so
|
||||
* `getUssr` finds `getUser` but `process` doesn't match `prosody`.
|
||||
* Bounded edit distance keeps each comparison cheap; the per-query
|
||||
* scan is O(distinct-name-count) which is far smaller than total
|
||||
* node count on any real codebase.
|
||||
*/
|
||||
private searchNodesFuzzy(
|
||||
text: string,
|
||||
options: { kinds?: NodeKind[]; languages?: Language[]; limit: number }
|
||||
): SearchResult[] {
|
||||
const { kinds, languages, limit } = options;
|
||||
const lowered = text.toLowerCase();
|
||||
const maxDist = lowered.length <= 4 ? 1 : 2;
|
||||
|
||||
// Pull the distinct name list once. The set is cached on QueryBuilder
|
||||
// by getAllNodeNames(); even on a 200k-node project the distinct
|
||||
// name set is typically O(10k) because most names repeat. The
|
||||
// candidate-cap below bounds memory regardless.
|
||||
const allNames = this.getAllNodeNames();
|
||||
const candidates: Array<{ name: string; dist: number }> = [];
|
||||
for (const name of allNames) {
|
||||
const dist = boundedEditDistance(name.toLowerCase(), lowered, maxDist);
|
||||
if (dist <= maxDist) candidates.push({ name, dist });
|
||||
}
|
||||
candidates.sort((a, b) => a.dist - b.dist);
|
||||
|
||||
// Cap the per-name follow-up queries. Each survivor triggers a
|
||||
// separate `SELECT * FROM nodes WHERE name = ?`; without this cap
|
||||
// a project with many similar names (`getUser1`, `getUser2`...)
|
||||
// could fan out far beyond `limit` queries before the inner-loop
|
||||
// limit kicks in.
|
||||
const FUZZY_FOLLOWUP_CAP = Math.max(limit * 2, 50);
|
||||
const cappedCandidates = candidates.slice(0, FUZZY_FOLLOWUP_CAP);
|
||||
|
||||
const results: SearchResult[] = [];
|
||||
const seen = new Set<string>();
|
||||
for (const c of cappedCandidates) {
|
||||
if (results.length >= limit) break;
|
||||
let sql = 'SELECT * FROM nodes WHERE name = ?';
|
||||
const params: (string | number)[] = [c.name];
|
||||
if (kinds && kinds.length > 0) {
|
||||
sql += ` AND kind IN (${kinds.map(() => '?').join(',')})`;
|
||||
params.push(...kinds);
|
||||
}
|
||||
if (languages && languages.length > 0) {
|
||||
sql += ` AND language IN (${languages.map(() => '?').join(',')})`;
|
||||
params.push(...languages);
|
||||
}
|
||||
sql += ' LIMIT 5';
|
||||
const rows = this.db.prepare(sql).all(...params) as NodeRow[];
|
||||
for (const row of rows) {
|
||||
if (seen.has(row.id)) continue;
|
||||
seen.add(row.id);
|
||||
// Lower the score for each edit step away from the query so
|
||||
// exact-match fallbacks (dist 0) outrank dist-2 typos.
|
||||
results.push({ node: rowToNode(row), score: 1 / (1 + c.dist) });
|
||||
if (results.length >= limit) break;
|
||||
}
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user