perf(resolution): generation-tagged supertype memo + method owner index — Swift compiler 185→98s, byte-identical (#1395)
The swiftc =2/nm:mc-* attribution located the wall: the getSupertypes conformance walk ran 971,200 times (565s of combined worker time, 581µs each) — every resolveMethodOnType miss re-queried implements/extends edges for every same-named type node, recursing depth-4 through Swift's protocol landscape with no memoization, and post-inference resolveMethodOnType averaged 1,912µs per call. Fix 1 — generation-tagged getSupertypes memo. Supertype edges GROW during the resolution loop (batch k persists its edges BEFORE batch k+1 fans out — the #1320 ordering), so a plain cache would freeze an early batch's emptier answer. Within a batch the edge state is fixed by that same ordering, so memo entries carry a generation that advances at every batch entry point (resolveBatchYielding / resolveListForAdmission — covering the sequential loop, pool workers, sync admission, and the conformance pass); a stale-gen entry recomputes. Behavior-identical to no memo at every point in time; walk invocation counts match the unmemoized run exactly (971,200 / 24,336 / 76,415). Fix 2 — per-(language, method-name) owner index in getMethodMatches: candidates bucket once by their qualifiedName's last two segments (exactly the span the match predicate tests), so a (type, method) query is a map lookup instead of an O(candidates) scan per methodMatchCache miss. ObjC selectors and multi-segment typeNames keep the legacy linear path. Also ships nm:mc-rmot / nm:rmot-supers =2 attribution rows. swiftc: settle 100.2→31.3s, resolveMethodOnType 1,912→202µs, wall 183.5→97.8s (was 185s at the head-to-head; cbm's same-box number is 119.1s). Gates: swiftc old-vs-new dump byte-identical (1,837,235 rows), swiftc pooled-vs-CODEGRAPH_NO_PARALLEL_RESOLVE=1 identical (the generation-semantics risk surface), dubbo old-vs-new identical (49k Java instance-method hits share both paths), Alamofire identical; suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
974e6c8b95
commit
157c8e735d
@@ -24,6 +24,7 @@ and adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
- Indexing very large projects on multi-core machines got faster again: the parallel-resolution workers now periodically refresh their read-only database connections, which lets database housekeeping advance instead of silently building up a backlog behind long-lived readers — a backlog that was taxing the indexer's own writes. Graphs remain byte-for-byte identical; the win is largest at Linux-kernel scale on many-core machines.
|
||||
- Indexing on macOS now uses the machine's real memory headroom when sizing its parallel-resolution workers. macOS deliberately keeps RAM filled with reclaimable cache, so the previous free-memory reading came back tiny (~1GB on an otherwise idle machine) and silently halved the worker pool — a medium Java project's fresh index ran about 15–20% slower than the hardware allowed. Graphs remain byte-for-byte identical; the same fix also lets a memory-driven analysis cache engage fully on macOS for large C codebases.
|
||||
- Fresh indexing got a sizeable across-the-board speedup: during the initial build, the database's secondary lookup indexes are set aside and rebuilt once after parsing instead of being maintained row by row — the same proven trick the later linking phase already used, now applied to the whole parse lane — and the reference-resolution loop likewise stops maintaining lookup indexes it never reads, rebuilding them at the end when almost nothing is left in the table. A medium Java project's parse phase runs about 58% faster and its full fresh index about 19% faster end-to-end; a Linux-kernel-scale index that took ~15 minutes on an 8-core machine now completes in about 11, with the resolution phase alone dropping by a third. Graphs remain byte-for-byte identical, and incremental syncs are unaffected.
|
||||
- Indexing Swift and other protocol/interface-heavy codebases got dramatically faster: the conformance walk that checks whether a method lives on a receiver's supertypes (protocols, base classes, extensions) now remembers its answers for the duration of each resolution batch instead of re-querying the graph for every call site — on the Swift compiler repository (27k files) that walk ran nearly a million times per index. A fresh index of that repo drops from about 185 seconds to under 100, with the graph byte-for-byte identical. Method-candidate lookup also gained a per-name owner index, so overload-heavy names (`init` in Swift, `execute` in Java) no longer pay a full candidate scan per receiver type.
|
||||
- Resolving method calls through local variables (`recv.method()`, Lua's `recv:method()`, R's `recv$method()`) got much cheaper on repos where the same receiver is called over and over: the declaration scan that types the receiver now remembers what it has already scanned per scope instead of re-reading the same source lines for every call site, and the regex patterns it scans with are compiled once per receiver instead of per call. Kong's fresh index drops another 8% on top of the require-resolution fix (23% cumulative), with graphs byte-for-byte identical everywhere — including Java projects, where this same scan successfully types tens of thousands of receivers.
|
||||
- Indexing Lua and Luau projects got a sizeable speedup: resolving each `require(...)` no longer rescans the project's entire file list four times — a per-project filename index answers the same lookup instantly, cutting per-require resolution from about a millisecond to microseconds. A fresh index of Kong (1,870 Lua files) runs about 16% faster end-to-end, with the graph byte-for-byte identical. The same housekeeping also closes a latent staleness edge where COBOL copybook lookups could keep serving a cached file list after files changed.
|
||||
- Parallel reference resolution now engages adaptively instead of by a fixed project-size cutoff: the indexer measures the actual per-reference resolution rate on the first batch and spins up the worker pool mid-run whenever the remaining work justifies it. Languages whose references are expensive to resolve benefit most — Rust especially: a fresh index of tokio runs about 23% faster, with the graph byte-for-byte identical. Small projects and low-core machines (2-core CI runners) keep the single-threaded path exactly as before.
|
||||
|
||||
Reference in New Issue
Block a user