Skip to content

Optimization catalog

Every performance decision that shipped, in one place: what it is, where it lives, why it’s safe, and the invariant that keeps it valid. If you change code near one of these, the invariant column is the contract you must re-verify. Measured numbers come from benchmarks.md and the CLAUDE.md decision log.

The headline result: on the 100-project synthetic workspace, vx’s all-hits path is ~3.9× faster than Turbo and ~5.4× faster than Nx, and the no-cache run lands within 4% of the bare-shell floor.

#WhatWhereWhy / effectInvariant to preserve
1xxHash3 for every cache-key site (was SHA-256)util/hash.ts, cache/cache.ts:key, execute-task.ts, fingerprint.ts, project-loader.ts~5× faster derivation on the warm path; 16-hex keys match Turbo’s xxh64 widthNot cryptographic — fine for content addressing, never use for auth/integrity vs. an adversary
2Seed-chained key folding (xxh3(part, prevDigest)) instead of a streaming hashercache/cache.ts:keyBun has no streaming xxh3; chaining avoids concatenating a big key bufferEvery variable-length part must be length-prefixed or \0-delimited so part boundaries stay unambiguous
3Workspace fingerprint computed once per run, folded into every task keyworkspace/fingerprint.tsOne lockfile read/hash instead of NMust cover every supported lockfile + pnpm-workspace.yaml; fixed file order
4Config module-cache busting by content hash, not mtimeworkspace/project-loader.tsSame content → Bun module-cache hit; changed content → fresh eval. Hashing a <10 KB config is ~50 µsThe hash must cover the full file bytes; mtime is not reliable (Bun mtimeNs undefined; ms granularity misses rapid edits)
5Per-run hashCache for repeated file/config hashesorchestrator/prepare.tsexecute-task.tsShared files (presets) hash once per runCache is per-run only; nothing may persist across runs without entering the key itself
#WhatWhereWhy / effectInvariant to preserve
6git ls-files -s --others --exclude-standard instead of an FS walk + ignore parsingcache/inputs.tsTurbo/Nx parity; git’s C-speed ignore handling; correct nested-.gitignore anchoring (fixed a real v13 bug); -s also harvests index OIDs (row 31)vx hard-requires git — no fallback walker. Absent git → UserError, never silent degradation
7One workspace-root git snapshot per run, partitioned per project (gitFilesCache)cache/inputs.ts, invalidated in execute-task.tsOne git ls-files spawn instead of P; zero re-spawns in src→dist layoutsStaleness rule: after a SAVE and after a RESTORE the exact changed (declared-output) paths are recorded via GitFilesCache.markOutputsChanged; downstream tasks re-spawn git only when their input globs can match a changed path. (The save-path used to DROP the whole entry — replaced 2026-06, −28% cold; a task writing UNDECLARED files a same-project downstream reads is undeclared behavior — declare your outputs)
8O(P log P) nested-project-boundary computation (sort + contiguous prefix scan; was O(P²))workspace/nested-dirs.tsNegligible at 10 projects, real at 1000Prefix match must include the trailing path.sep so pkg/a is not treated as parent of pkg/ab
#WhatWhereWhy / effectInvariant to preserve
9O(N + E) scheduler tick: per-node dep counters + ready priority queue (was O(N²) rescan per completion)graph/scheduler.tsTick cost independent of graph sizePriority contract: higher transitive-reverse-dep count first; ties break in graph-insertion order (binary-search insert respects existing equals)
9bBitset transitive-dependent closure in reverse-topo order (was memoized DFS over string Sets)graph/scheduler.ts:computeReverseDepCountSet closures were O(N²) entries: 8.5 s of a 10 s warm run on a 1090-package / 100-layer repo; bitsets = O(E·N/32), single-digit msCounts must stay EXACT (popcount of the closure), not a summed approximation — diamonds double-count under naive summing
10Iterative cycle detection with Uint8Array color array (was recursion + Map)graph/task-graph.ts:detectCycleNo V8 stack ceiling on deep dependsOn chains; no per-node Map costMust still report the cycle path in the error
11Group tasks execute with zero I/O — hash rolled up from upstream outcomesexecute-task.ts:computeGroupHashUmbrella tasks (install, ci) cost microsecondsGroup hash must fold every upstream id:hash pair, sorted, so it stays order-independent
#WhatWhereWhy / effectInvariant to preserve
12Hand-rolled in-process tar parse/extract (kept over Bun.Archive)cache/tar.tsBenchmarked: Bun.Archive is 15–400× slower for our artifact shape (KB–MB, flat trees) — fixed JS-bridge overhead dominates small archives. Also: no tar subprocess fork per restoreSecurity checks are part of the parser: reject absolute paths, .., hardlinks; unlink symlinks before write; re-verify resolved target stays under dest
13Restore skip via output_files rows + stat check (isOutputsCurrent)cache/cache.tsWarm-warm hit = N stats, zero writes, zero decompressStored size/mode/mtime fingerprint must match what extractOutputs produces (mtime compared at floor-to-second — tar headers carry seconds)
14SQLite metadata index, WAL, busy_timeout = 5000, one handle per runcache/cache.tsIndexed lookups; concurrent vx run invocations don’t crashEntry metadata lives in SQL only (v17 artifacts carry just stdout + outputs) — never reintroduce a meta.json that can drift from the rows
15Atomic artifact publish: unique tmp name → rename (no pre-rm)cache/cache.ts:writeArtifactAndIndexConcurrent saves of the same hash are either-or; readers never see partial bytesPOSIX rename replaces atomically; the pre-rm variant reintroduces a delete-after-rename race
16Single-transaction batch writes: recordRuns, prune deletes (+ parallel artifact rm)cache/cache.tsOne fsync instead of NPrune’s IN-list binding must stay under SQLite’s 999-placeholder limit per statement
17v17 artifact = exactly stdout + outputs/<rel>; identical bytes local and remotecache/cache.ts, layered-cache.tsNo stage-dir repack for upload — save re-reads the just-written artifact and PUTs it verbatim; remote hit writes the body straight to diskLocal and remote layers must keep transporting the same byte format; metadata travels out-of-band (SQL row / HTTP headers)
17bAsync remote prefetch: derive stable keys up front, fire remote GETs in the background, overlapping network with executionorchestrator/remote-prefetch.ts, cache/layered-cache.ts (prefetch + inflight map)A remote-served warm run no longer pays remote-GET latency on each task’s critical path; the GETs race alongside execution and land in local before execute-task needs themRemote-only (gated on LayeredCache; local-only runs do NO upfront key pass and NO local probing — byte-identical). Stable-key-only (a task whose inputs could match an upstream output is skipped → lazy read-through; gate shared with 17d via stable-keys.ts). At-most-once per key (prefetch + get share the inflight map; a settled-false miss blocks a second probe). Provenance stays remote (outcome cache-hit-remote) ONLY for genuinely remote-pulled hashes — local-first: a hash local already holds is skipped before the GET (no redundant download on a warm-local run) and stays local. Caller awaits the prefetch pool before cache.close() so no ingest hits a closed DB
17cBackground remote uploads: LayeredCache.save PUTs in the background; run() drains before cache.close()cache/layered-cache.tsA task’s outcome (and its dependents) never wait on upload latency; uploads race alongside the rest of the runEvery upload must be drained before the process exits (a short run still ships all artifacts); failures log via onRemoteError, never propagate; the uploaded bytes stay the verbatim local artifact (no repack)
17dLocal short-circuit + two-tier restore-ahead scheduler: up-front stable-key classify + ONE local probe per task; confirmed hits become a restore tier the scheduler runs ahead of their deps as LOW-priority worker backfillorchestrator/local-shortcircuit.ts, orchestrator/stable-keys.ts, graph/scheduler.ts−6.6% on a mixed slow-upstream/warm-downstream workload; parity on all-hit warm runs (probe reuse means zero double work)Local-only (never under a LayeredCache — prefetch owns those; an awaited up-front remote GET would sit on the critical path). Probe reuse — execute-task consumes preProbed, exactly 1 cache.get per stable task. Misses own the worker pool (execReady drains first). A graph with any outputs.workspaceFiles disables the restore tier. Never throws — degrades to the plain schedule
17e--dry / --graph remote prediction via HEAD existence probeorchestrator/plan.ts, cache/remote-cache.tsPlanning a remote-warm graph costs headers, not artifact downloads + local ingestPlanning stays side-effect-free on artifact storage: no download, no ingest; a predicted hit-remote = the artifact exists remotely
17fPure-SQL cache.get: stdout in the entries row; artifact untouched on probe; accessed_at bumps batch at flushcache/cache.tsHit cost no longer scales with artifact size (118 ms → 5 ms for four ~70 MB binaries); one batched UPDATE instead of per-hit writesThe entries-row stdout and the artifact’s stdout entry must stay identical (the artifact copy is what remote round-trips)
#WhatWhereWhy / effectInvariant to preserve
18Logger buffers chunks as string[], joins on flush (was += accumulation)orchestrator/logger.ts+= was O(N²) over total bytes for chatty tasksPer-task ordering within a stream must be append-only
19Memoized Bun.color ANSI lookupsorchestrator/colors.tsCalled thousands of times with one of four hex stringsCache key is the color string; gating (NO_COLOR etc.) happens before lookup
20Bun.Glob for filter matching + recursive listing (was hand-rolled regex / readdir recursion)workspace/filter.ts, cache/layered-cache.tsNative glob engineGlob semantics are now Bun’s — brace/bracket behavior changes with Bun upgrades
21Concurrent project discovery (Promise.all over package globs)workspace/workspace.tsWas serializedDedupe pass after must keep deterministic order
22AbortSignal.timeout for remote-cache fetchescache/remote-cache.tsDrops the manual controller + setTimeout ceremonyCatch both AbortError and TimeoutError
23toPosix fast path when path.sep === '/'util/paths.tsSkips split/join on the dominant platformWindows is unsupported anyway; revisit if that changes
24Hoisted dynamic imports out of per-task pathsexecute-task.ts, layered-cache.tsawait import() per task was measurable
25Bun.spawn everywhere (with resourceUsage())exec/runner.tsNative spawn + free cpu_ms / peak-RSS capture per child
25bStatus-region force-floor coalescing (forced redraws within 30 ms collapse into one trailing draw)orchestrator/status-line.ts6,540 forced redraws ≈ 6.7 MB of ANSI on a 3,270-task warm run → ~20 KBThe FINAL state must always land (the trailing draw); the first draw after idle stays immediate

June 2026 scaling pass (1090-package stress repo: 10.2 s → 0.62 s)

Section titled “June 2026 scaling pass (1090-package stress repo: 10.2 s → 0.62 s)”
#WhatWhereWhy / effectInvariant to preserve
26Bitset transitive-dependent closure for scheduler priority (was Set-DFS)graph/scheduler.ts:computeReverseDepCount8.5 s → ms on dense 100-layer graphs; O(E·N/32)Counts stay EXACT (popcount); own Kahn pass — never trust Map insertion order
27Bitset package-graph closures, sorted-name indexing (sort-free materialization)workspace/package-graph.ts68 ms → ~ms at 1090 projectsCyclic dep graphs fall back wholesale to legacy DFS
28Binary-search git-file partitioning (was O(P·F) startsWith)cache/inputs.ts:populateGitFilesCache54 ms → ~5 ms at 1090×9kPrefix ranges need the array sorted with the SAME comparator as the search
29Parallel project discovery; one readdir replaces 4 config exists-probesworkspace/workspace.ts:listProjectsserial I/O → concurrentDuplicate-name check stays a deterministic sequential pass
30Frontier ^task expansion (nearest holder per path, sparse bridging kept) — v19graph/task-graph.ts8.5× fewer edges on dense graphs; shrinks group-hash sorts, dep sorts, closures, upstream folds at onceHolder’s own dependsOn owns deeper ordering (Turbo parity, documented stop)
31Git blob OIDs as input-file hashes; dirty files get in-process blob OIDs; both git spawns concurrent — v20cache/inputs.ts, cache/cache.ts:hashFileClean-tree hashing: zero reads/stats/SQLite. 3.2× warm run-phase at 15k files (245→76 ms)A file’s key contribution must never flip across dirty↔clean (uniform blob-OID domain); symlinks/conflict stages never trusted
32Early cutoff: downstream keys fold upstream OUTPUT content identity — v21 REVERTED in v22orchestrator/upstream.ts (now folds the upstream INPUT key)Removed: pure-input transitive hashing replaced it (owner: “rely only on task input hashes”). Identical-output rebuilds re-run dependents again — rare in practice; see caching.md.n/a — no output content participates in any cache key in v22
33Scoped config loading: only in-scope projects + their transitive dep closure evaluateorchestrator/prepare.ts1090 config evals (~200 ms) → only the run’s closure; single-task wall 0.32 s → ~0.19 s on the 1090-package repoBoundary geometry still considers EVERY config-bearing project (loaded or not); a broken out-of-scope config surfacing only when it enters scope is the documented Turbo-like semantic

Known headroom (deliberately not taken yet)

Section titled “Known headroom (deliberately not taken yet)”

From benchmarks.md and the perf-pass backlog — candidates with a measured or suspected win, parked until profiled:

  • Batched cache-entry lookup for the all-hits path (Nx does one getBatch query; we probe per task). A previous getMetaBatch attempt was removed in the 2026-05 dead-code pass because it never wired into execute-task.ts — re-attempt only with the wiring.
  • Group-task reverse-index for ^build fan-out re-resolution.
  • Memoized taskConfigHash for projects sharing a preset config object (hash by object identity per run).
  • relPosix fast path for ASCII workspace-root prefixes in the enumeration loop.

For the record: restores are tar extracts with a stat-check fast path (see #12/#13) — there is no hardlink-based restore, and tar hardlink entries are rejected as a security measure.