Files
codegraph/scripts/kernel-parity.mjs
T
45a53eb5b5 feat(kernel): R7b Kotlin walker — kotlin module, vendored-grammar-C build, kotlin default-routed (#1382)
Sixth R7b port — the T1½ batch finale. Checklist-first recipe
(docs/design/kotlin-kernel-port-checklist.md, 1,121 lines, dist-extractor
ground truth); parity passed FIRST RUN on all three repos.

THE NOVEL MECHANISM — vendored-grammar-C (the §4 tracker's prescription,
first use): the crates.io tree-sitter-kotlin 0.3.8 pins `tree-sitter >= 0.21,
< 0.23` (the kernel links 0.25) and tree-sitter-kotlin-ng is a DIFFERENT
grammar (8 fields vs 0, renamed kinds — extractor-breaking), so no crate dep
is possible. The fwcd 0.3.8 tag's sha-matched parser.c + scanner.c are
vendored into codegraph-kernel/grammars/kotlin and compiled by build.rs (cc),
exposed via tree-sitter-language::LanguageFn. The wasm re-vendor is
behavior-NEUTRAL (0 CST/error disagreements across 1,984 gate-repo files;
old-vs-new full-init dumps byte-identical ×3) — a reproducibility re-vendor,
ABI stays 14.

Walker firsts: extension-function receivers (getReceiverType →
`WidgetK::extend` QN OVERRIDE with no package prefix, the qualified-receiver
`com::qext` first-segment bug, and the owner-contains fallback that excludes
`interface` kinds and is source-order dependent) and extractModifiers
(expect/actual platform modifiers → the node DECORATORS wire field on every
created node — the KMP synthesizer's feed, incl. `actual typealias`).
Preserved bug-for-bug: the FIELD_COUNT-0 dead cluster (no signatures, ZERO
type-annotation refs), hook-consumed property initializers emitting nothing
(incl. `by lazy {}`), the bodiless-vs-bodied class header asymmetry, enum-
entry bodies being invisible, KDoc never a docstring AND chain-breaking,
comment-gluing into import/package extents, `@Anno(args)` emitting nothing
while `@Marker` decorates, zero instantiates refs, the paren-then-lambda
`trailing()` garbage callee, text-includes visibility/suspend false
positives, and the packaged-file value-ref target drop. The fun-interface
misparse-recovery hook is DEFER-SHIELDED (every such file has_error) and
deliberately not ported. The swift-sweep lesson pre-applied: the shared
`assignment` shadow-prune case is implemented alongside the
property_declaration case.

Gates: sweeps 0-diff okio 299/322, okhttp 531/580, kotlinx.coroutines
1031/1082 (deferrals exactly the predicted 23/49/51 — both-arm grammar
reality incl. PHANTOM hasError files with complete CSTs; the kernel trusts
the flag); full-init dumps byte-identical ×3 (46.5k/108.9k/92.3k lines); KMP
expect/actual synthesis IDENTICAL across arms (412 edges on
kotlinx.coroutines — the tracker's KMP validation); kernel-kotlin-parity
suite (torture reflowed off the phantom shapes + .kts script + CRLF variants
+ fun-interface and phantom defer pins) + kotlin grammar-parity row (the
C-build ↔ wasm table identity proof); full suite 2,633 green ×2 under
CODEGRAPH_KERNEL_EXPECT=1. DEFAULT_ROUTED += kotlin (15 langs).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:33:18 -05:00

278 lines
11 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
/**
* Kernel↔wasm extraction parity harness (R2/R3 of the kernel migration).
*
* Runs BOTH extraction paths over the given files/directories and diffs the
* per-file ExtractionResults as sets (nodes/edges/refs, canonicalized), so a
* behavioral gap in the native kernel shows up as a categorized diff instead
* of a graph-dump surprise later. This is the fast inner loop; the §5 gate's
* full-repo dump-diff still runs before any default-on.
*
* Usage:
* node scripts/kernel-parity.mjs <file-or-dir>... [--lang typescript,tsx]
* [--max-samples N] [--list-files] [--max-deferral 0.1]
*
* --max-deferral: the broken-kernel backstop (default 0.1). For C/C++ pass
* 0.5: macro-heavy C/C++ trees genuinely parse with errors at 1040% file
* rates even after the preParse blanking family (git 19%, protobuf 26%, fmt
* 42% — measured 2026-07-17), and every erroring file defers BY POLICY, so
* the 10% bar calibrated on the 00.4% incidence of ts/java/py/go would fail
* healthy sweeps. A broken walker still trips 0.5 (it defers ~everything).
*
* Requires: npm run build (dist/) and a staged kernel (npm run build:kernel).
* Exit code: 0 = parity, 1 = diffs found, 2 = setup error.
*/
import * as fs from 'node:fs';
import * as path from 'node:path';
import { fileURLToPath } from 'node:url';
const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
const dist = (p) => path.join(ROOT, 'dist', p);
const args = process.argv.slice(2);
const paths = [];
let langFilter = null;
let maxSamples = 5;
let listFiles = false;
let maxDeferral = 0.1;
for (let i = 0; i < args.length; i++) {
if (args[i] === '--lang') langFilter = new Set(args[++i].split(','));
else if (args[i] === '--max-samples') maxSamples = Number(args[++i]);
else if (args[i] === '--list-files') listFiles = true;
else if (args[i] === '--max-deferral') maxDeferral = Number(args[++i]);
else paths.push(args[i]);
}
if (paths.length === 0) {
console.error('usage: kernel-parity.mjs <file-or-dir>... [--lang ts,tsx] [--max-samples N]');
process.exit(2);
}
const KERNEL_LANGS = new Set(['typescript', 'tsx', 'javascript', 'jsx', 'java', 'python', 'go', 'c', 'cpp', 'rust', 'csharp', 'ruby', 'php', 'swift', 'kotlin']);
const EXTS = new Map([
['.ts', 'typescript'], ['.mts', 'typescript'], ['.cts', 'typescript'],
['.tsx', 'tsx'], ['.js', 'javascript'], ['.mjs', 'javascript'],
['.cjs', 'javascript'], ['.jsx', 'jsx'], ['.java', 'java'], ['.py', 'python'], ['.pyw', 'python'], ['.go', 'go'],
// C/C++ (R7a). `.h` needs CONTENT sniffing (C vs C++) — resolved per file
// in the run loop via detectLanguage, matching the real indexer's routing.
['.c', 'c'], ['.h', 'detect'],
['.cpp', 'cpp'], ['.cc', 'cpp'], ['.cxx', 'cpp'], ['.hpp', 'cpp'], ['.hxx', 'cpp'],
['.metal', 'cpp'], ['.cu', 'cpp'], ['.cuh', 'cpp'],
['.rs', 'rust'], // R7b
['.cs', 'csharp'], // R7b
['.rb', 'ruby'], ['.rake', 'ruby'], // R7b
['.php', 'php'], ['.module', 'php'], ['.install', 'php'], ['.theme', 'php'], ['.inc', 'php'], // R7b
['.swift', 'swift'], // R7b
['.kt', 'kotlin'], ['.kts', 'kotlin'], // R7b
]);
/** Collect candidate files. */
function collect(p, out) {
let st;
try {
st = fs.statSync(p); // dangling symlinks (Linux-tree dtc fixtures) throw
} catch {
return;
}
if (st.isDirectory()) {
const base = path.basename(p);
if (base === 'node_modules' || base === '.git' || base === 'dist' || base === '.codegraph') return;
for (const e of fs.readdirSync(p)) collect(path.join(p, e), out);
} else if (EXTS.has(path.extname(p))) {
const lang = EXTS.get(path.extname(p));
// 'detect' (.h) resolves per file in the run loop; under --lang it rides
// along whenever either C-family language is requested.
const passes =
!langFilter ||
(lang === 'detect' ? langFilter.has('c') || langFilter.has('cpp') : langFilter.has(lang));
if (passes) out.push({ file: p, lang });
}
}
const files = [];
for (const p of paths) collect(path.resolve(p), files);
if (files.length === 0) {
console.error('no matching files');
process.exit(2);
}
// --- load the built engine ---------------------------------------------------
const { extractFromSource } = await import(dist('extraction/tree-sitter.js'));
const { initGrammars, loadGrammarsForLanguages, detectLanguage } = await import(dist('extraction/grammars.js'));
const kernel = await import(dist('extraction/kernel/index.js'));
await initGrammars();
await loadGrammarsForLanguages([...KERNEL_LANGS]);
if (!kernel.getKernel()) {
console.error('kernel .node not found — run: npm run build:kernel');
process.exit(2);
}
// --- canonicalization ---------------------------------------------------------
/**
* Node identity for cross-referencing edges/refs: the node id itself (both
* paths compute the same deterministic ids, and id embeds kind+name+line).
*/
function canonNode(n) {
const out = {
id: n.id, kind: n.kind, name: n.name, qualifiedName: n.qualifiedName,
filePath: n.filePath, language: n.language,
startLine: n.startLine, endLine: n.endLine,
startColumn: n.startColumn, endColumn: n.endColumn,
};
for (const k of ['docstring', 'signature', 'visibility', 'isExported', 'isAsync', 'isStatic', 'isAbstract', 'returnType']) {
if (n[k] !== undefined) out[k] = n[k];
}
if (n.decorators !== undefined) out.decorators = n.decorators;
if (n.typeParameters !== undefined) out.typeParameters = n.typeParameters;
return JSON.stringify(out);
}
function canonEdge(e) {
const out = { source: e.source, target: e.target, kind: e.kind };
if (e.line !== undefined) out.line = e.line;
if (e.column !== undefined) out.column = e.column;
if (e.provenance !== undefined) out.provenance = e.provenance;
if (e.metadata !== undefined) out.metadata = e.metadata;
return JSON.stringify(out);
}
function canonRef(r) {
// FULL object — a field only one path sets is a parity bug (the vitest
// parity suite caught decode.ts pre-filling filePath/language this way).
const out = {
from: r.fromNodeId, name: r.referenceName, kind: r.referenceKind,
line: r.line, column: r.column,
};
for (const k of ['filePath', 'language', 'candidates', 'rowId']) {
if (r[k] !== undefined) out[k] = r[k];
}
return JSON.stringify(out);
}
function diffSets(aList, bList) {
const a = new Map(); // canon -> count (multiset — duplicates matter)
const b = new Map();
for (const x of aList) a.set(x, (a.get(x) ?? 0) + 1);
for (const x of bList) b.set(x, (b.get(x) ?? 0) + 1);
const onlyA = [];
const onlyB = [];
for (const [k, c] of a) {
const d = c - (b.get(k) ?? 0);
for (let i = 0; i < d; i++) onlyA.push(k);
}
for (const [k, c] of b) {
const d = c - (a.get(k) ?? 0);
for (let i = 0; i < d; i++) onlyB.push(k);
}
return { onlyA, onlyB };
}
// --- run ----------------------------------------------------------------------
const buckets = new Map(); // category -> {count, samples[]}
function report(category, sample) {
let b = buckets.get(category);
if (!b) buckets.set(category, (b = { count: 0, samples: [] }));
b.count++;
if (b.samples.length < maxSamples) b.samples.push(sample);
}
let filesWithDiffs = 0;
let filesOk = 0;
let deferred = 0;
let processed = 0; // collected files minus content-detect skips
let totals = { nodes: 0, edges: 0, refs: 0 };
process.env.CODEGRAPH_KERNEL_LANGS = 'all';
for (const { file, lang: extLang } of files) {
const source = fs.readFileSync(file, 'utf8');
const rel = path.relative(ROOT, file);
// `.h` resolves C vs C++ by content — the same call the indexer makes.
const lang = extLang === 'detect' ? detectLanguage(rel, source) : extLang;
if (!KERNEL_LANGS.has(lang)) continue;
if (langFilter && !langFilter.has(lang)) continue;
processed++;
delete process.env.CODEGRAPH_KERNEL; // kernel path on
const kres = kernel.tryKernelExtract(rel, source, lang);
if (!kres) {
// Expected: files with parse errors defer to wasm (parity by
// construction — both arms run the same extractor). Counted, and
// guarded below so a broken kernel can't silently defer everything.
deferred++;
report('kernel-deferred', rel);
continue;
}
process.env.CODEGRAPH_KERNEL = '0'; // wasm path
const wres = extractFromSource(rel, source, lang);
delete process.env.CODEGRAPH_KERNEL;
totals.nodes += wres.nodes.length;
totals.edges += wres.edges.length;
totals.refs += wres.unresolvedReferences.length;
let fileHasDiff = false;
const tables = [
['node', wres.nodes.map(canonNode), kres.nodes.map(canonNode)],
['edge', wres.edges.map(canonEdge), kres.edges.map(canonEdge)],
['ref', wres.unresolvedReferences.map(canonRef), kres.unresolvedReferences.map(canonRef)],
];
for (const [table, wasm, kern] of tables) {
const { onlyA, onlyB } = diffSets(wasm, kern);
for (const x of onlyA) {
fileHasDiff = true;
const o = JSON.parse(x);
report(`${table}:missing-in-kernel:${o.kind ?? ''}`, `${rel}: ${x}`);
}
for (const x of onlyB) {
fileHasDiff = true;
const o = JSON.parse(x);
report(`${table}:extra-in-kernel:${o.kind ?? ''}`, `${rel}: ${x}`);
}
// ORDER matters too: identical multisets in a different emission order
// change DB rowids, and resolution iterates refs in rowid order — the
// full-index dump-diff would surface it as a downstream mystery. Catch it
// here instead.
if (onlyA.length === 0 && onlyB.length === 0) {
for (let i = 0; i < wasm.length; i++) {
if (wasm[i] !== kern[i]) {
fileHasDiff = true;
report(`${table}:order-mismatch`, `${rel}: index ${i}: wasm=${wasm[i]} kernel=${kern[i]}`);
break;
}
}
}
}
if (fileHasDiff) {
filesWithDiffs++;
if (listFiles) console.log(`DIFF ${rel}`);
} else {
filesOk++;
}
}
console.log(`\n=== kernel parity: ${filesOk}/${processed} files byte-parity` +
` (${filesWithDiffs} with diffs, ${deferred} deferred-to-wasm)` +
` | wasm totals: ${totals.nodes} nodes / ${totals.edges} edges / ${totals.refs} refs ===\n`);
const sorted = [...buckets.entries()].sort((a, b) => b[1].count - a[1].count);
for (const [cat, { count, samples }] of sorted) {
console.log(`--- ${cat}: ${count}`);
for (const s of samples) console.log(` ${s.length > 400 ? s.slice(0, 400) + '…' : s}`);
}
// Deferrals are per-file parse-error routing (expected; rare for most
// languages, routine for macro-heavy C/C++ — see --max-deferral above). A
// rate past the threshold means the kernel is broken and hiding behind the
// fallback — fail loudly.
const deferralRate = deferred / Math.max(processed, 1);
if (deferralRate > maxDeferral) {
console.error(
`deferral rate ${(deferralRate * 100).toFixed(1)}% exceeds ${(maxDeferral * 100).toFixed(0)}% — kernel likely broken`
);
process.exit(1);
}
process.exit(filesWithDiffs > 0 ? 1 : 0);