Skip to content
GitHubRSS

Why vx is fast

vx’s speed isn’t a microbenchmark trick; it comes from a handful of structural decisions. This page explains them. The exhaustive, source-cited catalog lives in Optimizations and the raw numbers in Benchmarks.

These are reproducible on your own machine, not marketing figures:

  • The runner’s overhead on a cold build, the number to read first. On the 3,270-task workspace the tasks alone take 3m 38s under an ideal schedule; vx finishes in 3m 46s (+0:08), Turborepo in 5m 13s (+1:35), Nx in 34m 44s (+31:06) — one unit for every runner. A runner that adds seconds to a three-minute build is a different tool from one that adds half an hour, and the per-package figure (8 ms, 88 ms and 1,712 ms per package) is how each grows with the codebase.

  • vx alone — bun packages/vx-bench/run.ts [projects] measures vx across fresh / warm-no-restore / warm-restore. A 100-project workspace replays fully-cached in 74 ms whole-process (1,000 projects in 172 ms), and a restore costs about the same as an untouched tree; the current floors are in Benchmarks.

  • Head-to-head vs Turborepo and Nx — bun packages/vx-bench/compare.ts scaffolds one repo (1,090 packages, 100 dependency layers, a build, installDeps and test task each: 3,270 tasks) and runs all three runners across the same three cache states. vx leads on the warm paths; the committed results live in Benchmarks. Run it yourself — every number here is a command away.

These are correctness and speed wins — they make the cache both safer and faster:

  1. Sparse ^task bridging. ^build walks through dependency packages that don’t declare the task to the nearest one that does, so sparse task coverage doesn’t need no-op filler tasks, and the graph carries no placeholder nodes for the packages walked through.
  2. Resolved-config hashing. vx hashes the evaluated vx.config.ts object, so imports, presets, and computed values participate in the key. Static-JSON config can’t see them.
  3. Strict output ownership. Declared outputs are wiped before exec and restore, so the tree ends as the cached snapshot, with no stale files. The one exception is a task that adds files to an upstream task’s outputs: it cleans nothing before exec and only the files it recorded before a restore. Turborepo/Nx restore additively.
  4. Daemonless. No background process, no staleness window, no socket to corrupt — and the fastest warm/cached runs in the head-to-head benchmark all the same.
  • Hashes come straight from git’s index. One git ls-files -s spawn yields the tracked file list and every clean file’s blob OID; a concurrent git status prunes anything that diverges and lists untracked files. Clean-tree key derivation costs zero reads, zero stats, zero database lookups. Dirty files get the identical blob OID computed in-process, so a key never flips across a commit boundary — a class of spurious miss the others accept.
  • Bitset graph algorithms. Scheduler priority and the package graph use packed-bitset closures with popcount instead of set-union DFS. On a 3,270-task graph this turned an 8.5 s priority computation into single-digit milliseconds.
  • A scheduler tick that re-scans nothing. Ready tasks come off an exact most-blocked-first binary heap; a completion decrements its direct dependents’ counters and pushes the ones that reach zero, so the run costs one pass over the edges plus an O(log N) heap operation per task, never a re-scan of the graph per completion.
  • Stat-check restore skips. A warm-on-warm restore is N stats with zero writes and zero decompression — fingerprints in SQLite tell vx the tree is already current.
  • One artifact format end to end. Local and remote move the same tar.zst bytes — metadata rides SQLite locally and the remote’s own record on the wire, so nothing is repacked at the boundary.
  • In-process tar, atomic publish, single-transaction SQL, and collision-hardened xxh3 key derivation round it out.

Speed by subtraction is still speed:

  • No daemon / project-graph process — the cold numbers say it isn’t needed.
  • A config-eval cache only where it is provably sound — configs are programs, so the cache is gated, not heuristic: a config that reads the environment, the clock, or anything non-deterministic is refused the cache outright (denied identifiers, including escaped and aliased spellings), and one that passes is keyed by the git blob ids of its whole import closure. The load configs stage is 16–25 ms per 1,000 configs served from it, against ~200 ms of evaluations; a refused config simply evaluates live.
  • No filesystem-tracing auto-inputs. Not a gap — a position. A traced input set describes what the task read that time, on that machine, which is not the same as what it depends on; and it cannot be known before the task runs, which is exactly when the key is needed. vx asks you to declare inputs and gives you a boundary to check them against: inside the workspace, a task with sandbox reads only what allow.read grants, plus node_modules and the linked packages it depends on. Reads outside the workspace stay open, and cache.inputs grants nothing, so the sandbox checks the input list only where allow.read mirrors it. Guessing is replaced by a boundary.