Skip to content
GitHubRSS

Benchmarks

Empirical overhead numbers vs. Turborepo and Nx on synthetic workspaces. Updated as the runners evolve.

The number that matters most to a developer is the warm no-op run: every task a cache hit, nothing to restore. packages/vx-bench/generate.ts workspaces, vx run build --all, this machine (macOS arm64, Bun 1.4.0), best of 5:

ProjectsBefore (2026-09-02 morning)Wave 1Wave 2Wave 5Wave 6Wave 7What changed
100105 ms92 ms79 ms78 ms77 ms74 msgit overlaps config load; .git/HEAD read replaces a git spawn; batched probe; no output walk; git first
1000380–450 ms270 ms242 ms237 ms204 ms172 ms+ cached pure-config evals; one worktree walk; readdir discovery; batched probe; no output walk; git first

Wave 7 (2026-09-03) starts the unscoped git enumeration the moment the workspace root is known — ahead of the workspace config, discovery and the cache open, ~25 ms earlier than before — so the status walk that is the warm run’s wall floor overlaps everything. Interleaved A/B against a worktree at the previous commit, three rounds of three: 1000 projects min 195 → 172 ms.

Wave 6 (2026-09-03) removed the per-hit output walk: for dist/**-shaped globs, the mtimes of every directory under the prefix, recorded at the last save or restore, prove the output set unchanged (docs/caching.md § A current tree). Interleaved A/B against a worktree at the previous commit, three rounds of three: 1000 projects min 224 → 204 ms.

Wave 3 (batched short-circuit probe, output rows carried on the entry, memoised Bun.Globs) measured on the graph WITH dependencies — vx run test --all, 2000 tasks, interleaved arms against an immutable worktree of the previous commit: 327–329 ms → 308–314 ms.

Where Wave 2’s 242 ms at 1000 projects went (VX_TIMING=1, see below; Waves 6 and 7 took 70 ms off it, mostly the run-graph phase and the git overlap): discovery 22 ms, config load 31 ms (all cache hits) overlapped with git’s one worktree walk (status -uall, ~57 ms, the critical path), the run-graph phase 78 ms (1000 hits: probe, output glob, stat check), history recording 12 ms, cache open 9 ms, and ~50 ms of process start + module load + exit outside the table. The 1.7 s cold run is the 1000 cp commands.

Reproduce: bun packages/vx-bench/run.ts 1000 5.

The table above is one machine’s. A four-core Linux container — sharing nothing with it — reads, medians with the full spread:

ProjectsWarmRestoreCold
1000271 ms (243–315)1031 ms (1023–1051)3147 ms (2905–3388)
5000807 ms (760–809)3931 ms (3462–4172)14 181 ms (13 798–14 986)

Slower in absolute terms, as a shared container should be, and those numbers are not comparable to the ones above — different hardware. What does compare is the SCALING: 5× the projects costs 2.98× the warm run here against 2.97× on the other machine (restore 3.81× against 3.97×, cold 4.51× against 4.99×). The sub-linear warm curve is the code’s, not one box’s.

One number to take from this before optimising against the harness: the warm arm spreads ±13 % about its median on identical code, because every rep is a whole CLI invocation. An A/B here needs an effect bigger than that, and a control arm beside it.

The same container class on 2026-09-23, after the cold-path work of items 615 (a config round’s evaluations written once per table) and 622 (output-directory snapshots landed in one transaction), medians with the full spread, five reps at 1,000 and three at 5,000:

ProjectsWarmRestoreCold
1000239 ms (231–265)906 ms (831–1009)2634 ms (2418–2814)
5000711 ms (702–733)3253 ms (2993–3630)10 950 ms (10 896–11 547)

Against the 2026-09-20 rows: cold −16 % at 1,000 and −23 % at 5,000, restore −12 % and −17 %; the warm rows moved within the ±13 % spread and claim nothing — the warm path is unchanged since item 589. Scaling 5× the projects: warm 2.97×, restore 3.59×, cold 4.16×.

Two tools, and they answer different questions:

  • VX_TIMING=1 vx run … prints a stage table to stderr at the end of the run — startup, workspace config, discover projects, package graph, open cache, load configs, git enumeration, build graph, plugin stages, classify + probe, run graph, record history, output dir snapshots, close, and in a dry run plan — with each stage’s own and cumulative time, plus accumulated per-task spans (cache.get, output glob, output stat and task hash among them; modules/timing.md lists every one). This is the first thing to read: it says WHICH stage moved. The per-task spans run under the scheduler’s concurrency, so they over-count (a span’s wall includes time yielded to other tasks); compare them to each other, not to the stage total.
  • bun --cpu-prof --cpu-prof-dir=/tmp/prof packages/vx/src/bin.ts run … then bun packages/vx-bench/profile-summary.ts /tmp/prof/*.cpuprofile gives self time by function and by file. Good for finding a hot loop; unreliable about where an await waited (it attributes the wait to whatever frame was on the stack).
  • strace -f -o <file> around a run, then bun packages/vx-bench/strace-vx.ts <file> [<before-file>]: the count of vx’s OWN syscalls (the tasks’ shells are told apart by PID), as a diff against the same run on a git worktree of the base. A syscall count is deterministic where this container’s wall time is not: the restore path’s three spare round trips per artifact (item 627) were three rows of that table — readlink 1,003 → 3, mkdir 2,002 → 1,002, newfstatat −2,000 — and the save’s two mkdirs of the cache directory (630) one row, before either was a number. A write per round trip is the thread pool’s wake, so that row counts round trips. bun --preload ./packages/vx-bench/sqlite-tally.ts … is the same idea for SQLite statements; restore-bench.ts and save-bench.ts in the same package restore or save every artifact sequentially, where a per-artifact change of tens of microseconds shows above the four-worker run’s noise (and one of a few microseconds does not: 630).

Three measurement lessons from this wave, recorded so they are not re-learned. A compiled Bun 1.4.0 binary resolves on-disk packages by <pkg>/index.ts only and ignores exports, so packages/vx-bench/compare.ts measured nothing (“vx skipped”) until the packages gained root shims — if the vx row ever reads n/a again, read the skip line first. A micro-benchmark of a sync call in isolation (statSync 2 µs vs stat 13 µs) does not predict the run — the async forms run in parallel on the thread pool under the scheduler’s concurrency, and switching the warm-hit path to sync calls made the 1000-project run 40 ms SLOWER. And Bun 1.4.0’s --compile binaries carry a signature this macOS rejects (SIGKILL on launch); an ad-hoc codesign -s - --force repairs it, which the release workflow now does on a macOS runner.

Head-to-head, 2026-09-03 (46 packages, packages/vx-bench/compare.ts 10 5 1)

Section titled “Head-to-head, 2026-09-03 (46 packages, packages/vx-bench/compare.ts 10 5 1)”

Same workspace, identical commands, every runner pinned to concurrency 10, Turbo with no daemon (it uses none for turbo run since 2.9), vx as its compiled binary. Median of 1, this machine (macOS arm64, Bun 1.4.0). The harness then gave Nx npm where Turbo had bun (§ Why Nx is slower), so the Nx row is slower than a fair one. Every runner runs as in CI (CI=1), so Nx’s daemon is off:

RunnerVersionFresh (cold)Warm (no restore)Warm (restore)
vx0.0.010.45 s76 ms83 ms
vx (frozen)0.0.010.49 s83 ms88 ms
turbo2.10.1210.58 s71 ms97 ms
nx23.2.019.66 s540 ms531 ms

Read it honestly: at 46 packages Turborepo 2.10 and vx are within a few milliseconds of each other on a fully-cached run, and neither keeps a process between runs: Turbo 2.10 uses no daemon for turbo run (its docs say so from 2.9), so both work out what changed on every invocation. vx wins the restore case and ties the cold one; Nx is 7× off. The remaining fixed cost at this size is process start + git, not the pipeline.

The same 46-package run on the four-core Linux container (2026-09-24, item 735, the fixed harness: Nx runs its scripts with bun as Turbo does; CI=1, so Nx’s daemon is off; a different machine, so only the ratios compare with the table above). Turbo 2.11.3 (no daemon for turbo run) and Nx 23.2.1, vx as its compiled binary, median of 3. The CPU column is user + system of the invocation and every child it waited for; a daemon that outlives the invocation would not be counted, and none runs here:

RunnerVersionFresh (cold)Warm (no restore)Warm (restore)CPU, cold
vx0.0.010.29 s79 ms104 ms846 ms
vx (frozen)0.0.010.27 s74 ms95 ms826 ms
turbo2.11.310.43 s86 ms (1.1×)136 ms (1.3×)1.33 s
nx23.2.122.08 s844 ms (10.7×)862 ms (8.3×)43.79 s

The ideal schedule is 10.00 s, so vx and Turbo both sit on the critical path cold, and warm they are within a few milliseconds at this size (the 2026-09-23 run of the old harness read Turbo at 112 ms). The same box under the old harness read Nx at 27.21 s cold and 1m 2s of CPU.

The same harness at 476 packages / 1,428 graph nodes (packages/vx-bench/compare.ts 20 25 1, 2026-09-02, same machine; a mid-size data point — the committed packages/vx-bench/RESULTS.md is the 3,270-task run below):

RunnerFresh (cold)Warm (no restore)Warm (restore)
vx1m 40s297 ms416 ms
vx (frozen)1m 40s285 ms399 ms
turbo1m 40s342 ms (1.2×)612 ms (1.5×)
nx3m 23s1.38 s (4.7×)1.33 s (3.2×)

The same size on the four-core Linux container (2026-09-25, after items 744, 753 and 754, the fixed harness, median of 1; ideal schedule 1m 36s):

RunnerFresh (cold)Warm (no restore)Warm (restore)CPU, cold
vx1m 37s225 ms367 ms7.06 s
vx (frozen)1m 37s209 ms377 ms6.94 s
turbo1m 39s247 ms (1.1×)392 ms (1.1×)13.84 s
nx2m 27s1.84 s (8.2×)1.89 s (5.2×)7m 23s

Read it honestly: the 2026-09-24 run on this box had Turbo 2.11 winning both warm columns (303 and 446 ms against vx’s 376 and 478). Item 744 found the cost in vx’s stable-key pass (a string set copied per task per dep, now a bitset), item 753 coalesced the 952 task lines into a write per turn of the event loop, and item 754 ranked a warm run’s tasks over the work that actually waits; vx now leads both warm columns, by 10% and 7% in one rep each, which this container’s noise can close. The min-of-N interleaved read of the warm no-op on the same workspace, with compiled binaries, is vx 165–175 ms against Turbo 226 ms (item 753’s profile).

Two things, both per task, and both measured on the 46-package workspace (cold, 138 tasks, min of 3 interleaved arms, 2026-09-24, item 735):

ArmWallCPU
the old harness: nx:run-script, Nx picked npm28.82 s69.80 s
nx:run-script with bun (the fixed harness)21.37 s41.18 s
nx:run-commands: the same commands, no fork11.65 s3.05 s
  1. A Node process per task. nx:run-script runs every task in a freshly forked Node process (nx/bin/run-executor.js: 138 of them in one cold run, counted with strace -f -e execve) that loads Nx before it runs the script. That is ~270 ms of CPU per task. On four cores at concurrency 10 it saturates the CPU, so every task on the 20-task critical path waits for a core: 9.7 s of wall and 38 s of CPU, the gap between the last two rows. nx:run-commands runs its command from the Nx process itself (NX_RUN_COMMANDS_DIRECTLY), and with it Nx lands at 11.65 s against an ideal of 10.00. A package-based Nx repo gets nx:run-script for every inferred package.json script.
  2. npm, which the harness chose by accident. nx:run-script runs the script through the package manager Nx detects from a lockfile. The generated workspace had none, so Nx fell back to npm, while Turbo read packageManager and ran bun run. npm run costs 202 ms of CPU per task against bun run’s 4 ms (min of 10, one task alone): 7.5 s of wall and 29 s of CPU, the gap between the first two rows. This one was the harness’s fault, not Nx’s. nx.json now says cli.packageManager: 'bun'.

Warm is spread thin, with no single cause. With no daemon (CI), each invocation starts three plugin workers to build the project graph (~280 ms each, in parallel; nx show projects alone is 0.56 s), then hashes and replays the cached tasks. Turning the daemon on moves the graph into it and saves 40–160 ms (0.81 s against 0.85, min of 5, in a probe; 682 ms against 844, median of 3, in the harness), a fifth of the warm gap at most.

The npm share grows with the graph. At 1,090 packages (3,270 tasks, compare.ts 100 11, same box, one cold run each) Nx took 20m 39s and 76 min of CPU with npm, and 7m 22s and 23 min of CPU with bun: npm was two thirds of the old harness’s Nx number there, which is the run the site quoted. The site’s Nx column (34m 44s cold, § A real monorepo, the macOS machine) paid npm per task and is not a fair one; its vx and Turbo columns are unaffected. Next 18 re-runs it. The whole 3,270-task shape on this box with the fixed harness (2026-09-25, after items 744, 753 and 754, median of 1; ideal schedule 3m 38s):

RunnerFresh (cold)Warm (no restore)Warm (restore)CPU, cold
vx3m 40s359 ms653 ms16.05 s
vx (frozen)3m 41s292 ms651 ms16.46 s
turbo5m 4s (1.4×)431 ms (1.2×)722 ms (1.1×)33.27 s (2.1×)
nx6m 59s (1.9×)4.50 s (12.5×)4.60 s (7.0×)20m 55s (78.2×)

The day before, Turbo 2.11 won both warm columns here (496 and 856 ms against vx’s 678 and 971). Item 744 found why: vx’s stable-key pass copied a string set per task per dep, now a bitset; items 753 and 754 then took the per-line stdout writes and the whole-graph priority closure off the warm path. vx now leads every column at this size, in one rep each.

Refuted, each within ±0.3 s of the fixed harness’s 21.2 s cold: the per-task pseudo-terminal (NX_NATIVE_COMMAND_RUNNER=false) and the output style (NX_TUI=false, --outputStyle=stream). So is the daemon: the fixed harness read Nx cold at 22.25 s with NX_DAEMON=true and 22.08 s without. Parallelism is honoured: with cheap tasks, the nx:run-commands arm sits 1.65 s over the ideal schedule. The ~4 git probes each forked task runs cost ~8 ms of it. These tables used to say Nx’s daemon was on; CI=1 had always turned it off, and the harness keeps it off on purpose: it simulates CI.

A real monorepo: 3,270 tasks, 100 layers (2026-09-03)

Section titled “A real monorepo: 3,270 tasks, 100 layers (2026-09-03)”

The shape that actually stresses a task runner: 100 dependency layers, ~11 packages per layer, ~30 deps per package, three tasks each (build + installDeps + test, sleep 1 for build and test) — 3,270 task nodes, 1,090 packages. Same repo, same hardware, same task commands; every runner pinned to concurrency 10. bun packages/vx-bench/compare.ts 100 11 1, this machine (macOS arm64, 10 cores), Turbo 2.10.12, Nx 23.2.0. The committed packages/vx-bench/RESULTS.md / packages/vx-bench/results.json are this run.

vxTurborepoNx
Cold (nothing cached)3m 46s5m 13s (1.4×)34m 44s (9.2×)
Warm, nothing to rebuild510ms760ms (1.5×)3.59s (7.0×)
Warm, restore outputs777ms1.17s (1.5×)4.15s (5.3×)
CPU burned, cold (user+sys)34.61s1m 13s (2.1×)114m 06s (197.8×)
CPU burned, warm (user+sys)1.34s4.40s (3.3×)5.54s (4.1×)
Baseline (theoretical best)3m 38s cold; 0 warm, restore, CPU——
Measured floors (context)git walk 67ms · walk + raw copy 352ms · task shells 33.15s——

Baseline is the theoretical best case, so each row shows its overhead: cold is the tasks’ own durations list-scheduled on 10 workers along the exact dependency graph (critical path 1m 40s, total work ÷ workers 3m 38s); a cached run, a restore and the CPU a runner burns are 0 in theory, so every measured number in those rows is the runner. vx’s cold overhead over the ideal schedule is 8.45s on 3,270 tasks (8 ms per package); Turborepo’s is 1m 35s (88 ms per package) and Nx’s 31m 06s (1,712 ms per package) — the number to read first, in one unit for every runner: a runner that adds seconds to a three-minute build is a different tool from one that adds half an hour. For context, the measured floors row gives what the cheapest possible implementation of each step costs on this machine: one git status -uall walk (the cost of asking what changed), that walk plus a raw copy of every output file, and the task shells themselves under xargs -P 10 (which vary by about two seconds between runs; vx’s cold CPU sits within that noise).

CPU is user + system time of the invocation and every child it waited for. The tasks are sleep, so this is the runner’s own work; a daemon that outlives the invocation (Nx’s) is not counted, so Nx’s CPU is a floor.

Methodology note: a synthetic graph with sleep-based tasks isolates runner overhead from real compilation. All three runners are configured identically — same commands, the same src/** inputs and dist/** outputs, the same concurrency. (Hashing **/* instead would include each task’s own output in its inputs and break caching for everyone.) An earlier run of this shape (June 2026, a 4-core Linux box) read cold 3m 48s / 8m 18s / 8m 27s and CPU 22.7 s / 1,250 s / 2,038 s; cold wall time depends on how many cores the runners’ overhead competes with the tasks for, which is why the CPU row is the one that travels.

Reproducible head-to-head (vx vs Turborepo vs Nx)

Section titled “Reproducible head-to-head (vx vs Turborepo vs Nx)”

packages/vx-bench/compare.ts scaffolds one shared monorepo matching the shape above — layers × perLayer packages, ~30 deps each, three tasks (build + installDeps + test) with the identical shell command, src/** inputs, and dist/** outputs for every runner — then runs vx, Turbo, and Nx across three cache states. Fairness is deliberate: vx runs as the compiled binary real users install (not TS source); the workspace is git-committed with node_modules/.turbo/.nx ignored; every runner is pinned to the same concurrency; and runners are measured strictly one at a time, daemons stopped between them, so they never fight for CPU. build/test sleep 1 s so a warm hit visibly skips the work.

Terminal window
bun packages/vx-bench/compare.ts # 100 layers × 11 (3,270 nodes) — the full shape (slow)
bun packages/vx-bench/compare.ts 10 5 1 # 46 packages, 10 layers — quick
BASELINE_ONLY=1 bun packages/vx-bench/compare.ts # recompute only the baseline floors against the committed rows (~9 min)
bun packages/vx-bench/update-site.ts # rewrite the landing page, the README's bench sentence and chart, and this doc's stress section from results.json (--check to verify)
BUILD_SLEEP=0 bun packages/vx-bench/compare.ts 20 11 2 # deep graph, pure framework overhead

It writes packages/vx-bench/RESULTS.md (committed, so the numbers can be referenced from a commit).

Since 2026-09-03 the table also carries CPU columns and a baseline row — the theoretical best case (an ideal schedule of the tasks, one git walk, a raw copy of the outputs, the commands under xargs), so each runner’s row reads as overhead above it. The definitions live in the generated packages/vx-bench/RESULTS.md. The 46-package quick run is the head-to-head table above (2026-09-03); an earlier run of that shape used to sit here with a different verdict on the warm row, and one page carrying both was a contradiction, so it is gone.

vx lock + --frozen is measured as its own row: it executes the frozen vx-lock.json graph with zero per-run config evaluation. Read the row as a tie, not a win: since the config-evaluation cache (2026-09-02) the plain warm run evaluates nothing either for a config the purity gate can prove pure, and it serves the same validated object from cache.db without re-validating it, while --frozen parses the whole lock and re-validates every entry (the lock is hand-editable, so it is a boundary). What frozen still skips is the per-config identity stat; what it still pays is the lock’s own parse. Measured 2026-09-12 on the 1,000-project bench, compiled binary, 12 interleaved reps: plain min 154 / median 177 ms, frozen 148 / 165 — ~5%, the identity stats; the load configs stage reads 20–25 ms plain against 6–10 ms of lock read plus 12–14 ms of load frozen. The 83 vs 76 ms above (2026-09-03, median of 1) is the same tie under a laptop’s noise. --frozen is for what it guarantees — what runs is what was locked, whatever a config would read from the environment — and for configs the gate cannot prove pure, which evaluate live on every plain run and come from the lock under --frozen. In your repo: vx lock, then commit vx-lock.json.

How the overhead scales with the workspace (2026-09-10)

Section titled “How the overhead scales with the workspace (2026-09-10)”

The per-package figure above is one size. This is vx alone at three sizes of the same synthetic shape (packages/vx-bench/run.ts, the generator’s workspace, one build per package whose command is a mkdir and a cp, so the clock is the runner and almost nothing else), the compiled Linux binary at commit 1a35ec3, median of 3 on a 4-core Intel Xeon container:

PackagesCold (nothing cached)per packageWarm, nothing to rebuildper packageWarm, restore outputsper package
100310 ms3.1 ms56 ms0.56 ms126 ms1.26 ms
300762 ms2.5 ms100 ms0.33 ms269 ms0.90 ms
1,0002,091 ms2.1 ms178 ms0.18 ms808 ms0.81 ms

Ten times the packages costs 6.7× the cold time and 3.2× the warm time: the per-package cost falls as the fixed cost (process start, the workspace read, the cache open) is spread over more of them, and nothing in the run grows faster than the graph. A cold run at 1,000 packages is two seconds; a warm one is under two hundred milliseconds. These are the runner’s own costs on a trivial task; the 3,270-task table at the top, where each task sleeps a second and every runner is scheduled the same way, is where the same shape is compared against Turborepo and Nx.

A real Turbo repo: solidjs/solid (2026-09-10)

Section titled “A real Turbo repo: solidjs/solid (2026-09-10)”

Not a synthetic workspace: solidjs/solid at b25c557 (5 packages, pnpm 9, Turbo 2.10.10 as the repo’s own devDependency, Node 22), with vx put on top of the repo’s own turbo.json through turbo() (then @vzn/vx-turbo, now @vzn/vx-migrate) — a two-line vx.workspace.mjs, no config rewritten. Both tools see the same graph: build is four executed tasks (solid-js#types, #link, #build, solid-element#build; Turbo lists three more build nodes for packages with no such script and runs nothing for them), and test test-types is seven. Both restore the identical 64 output files. vx as its compiled binary, Turbo with no daemon (2.10 uses none for turbo run, so every run pays its own discovery, the same footing vx is on; the --no-daemon the script passes is ignored), four cores, Linux, arms interleaved. Medians (3 reps for build, 2 for test); the script is packages/vx-bench/real/turbo-repo.sh.

build (4 tasks)vxTurbo 2.10.10
cold (caches + outputs wiped)40.6 s45.5 s (1.12×)
warm, outputs wiped (restore)66 ms127 ms (1.9×)
warm, nothing wiped (no-op)51 ms95 ms (1.9×)
test test-types (7 tasks)vxTurbo 2.10.10
cold53.6 s58.2 s (1.09×)
warm, restore80 ms166 ms (2.1×)
warm, no-op59 ms93 ms (1.6×)

Read it honestly: the cold rows are rollup, tsc and vitest — the runner is a few percent of them, and the 4–5 s gap is Turbo’s per-task work around the same commands (its ** default inputs hashed per package, its log capture, its cache write), not measured to the frame here. The warm rows are the product: with everything cached, vx answers in 50–80 ms where Turbo takes 95–170 ms, and the restore case — what a CI job or a fresh checkout does — is where the ratio is widest. Neither tool has a daemon to turn on for this: Turbo’s docs deprecate it for turbo run from 2.9, and vx has none.

The same footing as the solid run, on the largest Turbo repos on GitHub: the repo’s own turbo.json, vx on top through @vzn/vx-migrate’s turbo() with a two-line vx.workspace.mjs, both tools scoped by the repo’s own filters, four cores, Linux, arms interleaved, medians of three reps. Both tools at Turbo’s default of 10 workers (vx’s default is the core count; medusa’s own script says --concurrency=100% and both get it). Turbo’s dry-run and vx’s --dry plan the same pkg#task set on every repo. One binary for all five (the restore and warm-hit fixes of STATUS 141 are in it). The script is packages/vx-bench/real/turbo-repo.sh; noop2 is a second consecutive no-op. Every repo’s revision, toolchain, scope and bench-side adjustment is in packages/vx-bench/real/REPOS.md.

What the harness does that the first attempt did not, all of it for Turbo’s benefit as much as vx’s: every git-ignored artifact outside the installs and the two caches is cleaned before a cold and a restore arm (Turbo does not clean outputs; the first run leaked .turbo logs, prebuilt files and 42 stale tsconfig.tsbuildinfo files between arms); the repo’s root node_modules/.bin is on PATH as the repo’s own yarn build would have it (bare, Turbo lost medusa’s rollup); npm trusts this container’s proxy CA (cal.com’s embed build runs npx); and astro’s build inputs exclude its own outputs (below).

Two tasks needed a bench-side output list, declared in the repo’s vx.workspace.mjs as a ten-line project-stage plugin and named here so nobody reads them as the plugin’s own mapping: medusa’s build outputs are */** minus !src/** in turbo.json, which vx (no output negation) runs uncached, so the bench names dist/** and .medusa/**; and cal.com’s @calcom/web#build writes 110 symlinks to node_modules directories under .next/node_modules, which vx’s artifact format does not store, so the bench names the rest of .next — everything next start reads — and Turbo’s artifact carries the 110 links too. An artifact still stores no directory symlink; the save refuses one by name.

withastro/astro (32 build tasks, pnpm 10, Turbo 2.10.2)

Section titled “withastro/astro (32 build tasks, pnpm 10, Turbo 2.10.2)”

astro’s build declares inputs: ["**/*", …], and an explicit Turbo input glob matches the filesystem, gitignored or not — so as shipped, each package’s own dist/** is in its hash. Turbo’s first run after a restore then rebuilds the whole graph (59.6 s on the first harness), and because 19 of the 32 builds are not byte-reproducible the run after that rebuilds those 19 again (48.8 s, 13 cached), forever. Probed with --dry=json: appending one byte to a gitignored dist/index.js changes the task hash. vx excludes a task’s declared outputs from its inputs on the same config. The table is on the fixed config — !dist/**/* and !src/**/*.prebuilt* added to the build, build:ci and prebuild inputs — so Turbo is measured at its best.

buildvxTurbo 2.10.2
cold (caches + outputs wiped)52.0 s62.9 s (1.21×)
warm, outputs wiped (restore)887 ms1.58 s (1.78×)
warm, nothing wiped (no-op)627 ms1.17 s (1.86×)
second no-op632 ms1.18 s (1.86×)

payloadcms/payload (45 build tasks, pnpm 10, Turbo 2.10.4)

Section titled “payloadcms/payload (45 build tasks, pnpm 10, Turbo 2.10.4)”

Scoped as the repo’s own build:all (the four templates excluded). Turbo’s default inputs (the git-tracked files), no explicit glob.

buildvxTurbo 2.10.4
cold (caches + outputs wiped)126.9 s127.5 s (1.00×)
warm, outputs wiped (restore)3.44 s3.46 s (1.01×)
warm, nothing wiped (no-op)256 ms237 ms (0.93×)
second no-op275 ms267 ms (0.97×)

Parity, and the doc says so: the no-op rows are within noise of each other and Turbo takes both. The difference in what the two runs DO is not noise: on every hit vx loads the 14,430 recorded output rows and stats every file (~36 ms) to prove the outputs are intact, Turbo checks nothing on disk — delete a file under dist and turbo run build still prints a hit. Before STATUS 141 this repo read 129 s / 6.3 s / 372 ms / 325 ms for vx.

medusajs/medusa (83 build + build:plugin tasks, yarn 3, Turbo 1.13.4)

Section titled “medusajs/medusa (83 build + build:plugin tasks, yarn 3, Turbo 1.13.4)”

The repo’s own --concurrency=100% for both. 24k tracked files.

build build:pluginvxTurbo 1.13.4
cold (caches + outputs wiped)308 s315 s (1.02×)
warm, outputs wiped (restore)3.86 s7.16 s (1.85×)
warm, nothing wiped (no-op)947 ms3.29 s (3.5×)
second no-op953 ms3.12 s (3.3×)

Before STATUS 141 vx’s no-op here was 2.5 s: every task carried the same globalDependencies literal and resolved it against the whole enumeration, 76 of 83 were hashed twice, and the absent .medusa/** prefix refused every directory snapshot.

n8n-io/n8n (70 build tasks, pnpm 12, Turbo 2.9.18, Node 24)

Section titled “n8n-io/n8n (70 build tasks, pnpm 12, Turbo 2.9.18, Node 24)”
buildvxTurbo 2.9.18
cold (caches + outputs wiped)137.2 s133.7 s (0.97×)
warm, outputs wiped (restore)8.27 s12.5 s (1.52×)
warm, nothing wiped (no-op)842 ms1.36 s (1.62×)
second no-op839 ms1.34 s (1.60×)

The cold row is n8n-nodes-base#build (55 s alone) plus what fits around it; a first vx rep read 193 s in the disk’s slow phase and the median absorbed it.

calcom/cal.com (13 build tasks, yarn 3, Turbo 2.7.1, scope @calcom/web...)

Section titled “calcom/cal.com (13 build tasks, yarn 3, Turbo 2.7.1, scope @calcom/web...)”

Three of the thirteen are cache: false in turbo.json (prisma’s generate among them, ~12 s together) and run on every arm under both tools, so the warm rows have a 12 s floor. .env from the example with the two empty secrets filled, SKIP_DB_MIGRATIONS=1, no database.

buildvxTurbo 2.7.1
cold (caches + outputs wiped)245.7 s250.7 s (1.02×)
warm, outputs wiped (restore)17.4 s19.9 s (1.14×)
warm, nothing wiped (no-op)14.7 s18.5 s (1.26×)
second no-op14.5 s17.8 s (1.22×)

Wide graphs (2026-09-11, one rep, 3 workers)

Section titled “Wide graphs (2026-09-11, one rep, 3 workers)”

Each repo’s whole task set, not only build: the graph a team actually runs. One rep per arm and both tools at 3 workers, because the session’s memory cgroup allows 13.3 GiB and the wide sets’ typecheck and lint processes run at 3.6–5.7 GB resident each. Same harness, same cleanup, same scope as the build tables. Two sets ran; three could not, and the reasons are the repos’ own (below), the same under both tools.

payloadcms/payload — build lint, 89 tasks. payload’s lint is cache: false in its own turbo.json, so the 44 lint tasks run on every arm under both tools (~225 s at 3 workers); the three warm rows are that floor, and the cold row is the floor plus the 45 builds.

build lintvxTurbo 2.10.4
cold (caches + outputs wiped)318 s334 s (1.05×)
warm, outputs wiped (restore)229 s223 s (0.97×)
warm, nothing wiped (no-op)227 s228 s (1.01×)
second no-op226 s226 s (1.00×)

calcom/cal.com — build lint, 24 tasks, scope @calcom/web.... type-check is cache: false in the repo’s turbo.json and stays out; the three uncached build tasks (~12 s) run on every arm as in the build table, so the warm rows are the runner plus that floor.

build lintvxTurbo 2.7.1
cold (caches + outputs wiped)246.0 s236.8 s (0.96×)
warm, outputs wiped (restore)16.2 s19.7 s (1.22×)
warm, nothing wiped (no-op)14.3 s17.7 s (1.23×)
second no-op14.3 s18.0 s (1.26×)

medusajs/medusa — build build:plugin test, 157 tasks: dropped. medusa’s test declares no dependsOn, so on a cold tree both tools start tests before the packages they import are built: vx lost @medusajs/auth#test (Cannot find module '@medusajs/framework/awilix', and it passed on the restore arm once the build existed), Turbo lost @medusajs/dashboard#test (Failed to resolve entry for package "@medusajs/admin-vite-plugin") on all four arms. A repo configuration gap the two runners expose identically; not a number for either.

withastro/astro — build test, 55 tasks: dropped, the same gap: test depends on ^test only, and its tests import their own package’s dist. Turbo lost @astrojs/internal-helpers#test, upgrade#test and telemetry#test on every arm; vx’s scheduler happened to run the builds first and lost @astrojs/language-server#test, which imports packages/astro/dist across packages, plus @astrojs/ts-plugin#test, which downloads VS Code and cannot behind this proxy.

n8n-io/n8n — build typecheck lint, 220 tasks: dropped (owner). The editor-ui’s vue-tsc and eslint at 3.6–5.7 GB resident each do not fit the cgroup beside anything else.

The Turbo footing for Nx: the repo’s own Nx (NX_DAEMON=false, NX_NO_CLOUD=true) against vx on the vx.config files bunx @vzn/vx-migrate --from nx wrote from the repo’s exported graph, both tools at the worker count the repo’s nx.json sets, medians of three interleaved reps, the same cleanup and arm logs as turbo-repo.sh. The script is packages/vx-bench/real/nx-repo.sh — the Turbo harness runs the Turbo repos and this one runs these, and each names the other only for the footing they share. Parity is the task graph: nx run-many … --graph against vx --dry=json plan the same project#target set. Every revision, toolchain and bench-side rule (the configs rewritten to .mjs, the pinned package manager on PATH) is in packages/vx-bench/real/REPOS.md § Nx repos.

TanStack/query (25 build tasks, pnpm 11, Nx 22.1.3, parallel: 5)

Section titled “TanStack/query (25 build tasks, pnpm 11, Nx 22.1.3, parallel: 5)”

Scoped as the repo’s own build script (examples/** and integrations/** excluded). Every target is nx:run-script.

buildvxNx 22.1.3
cold (caches + outputs wiped)47.4 s55.1 s (1.16×)
warm, outputs wiped (restore)656 ms1.98 s (3.02×)
warm, nothing wiped (no-op)174 ms1.88 s (10.8×)
second no-op156 ms1.89 s (12.1×)

strapi/strapi (39 build tasks, yarn 4, Nx 20.8.4, parallel: 8, --nx-ignore-cycles)

Section titled “strapi/strapi (39 build tasks, yarn 4, Nx 20.8.4, parallel: 8, --nx-ignore-cycles)”

Every package script is yarn’s run -T <root bin>, so each task is yarn run build under both tools (STATUS 142); the per-package build is npm-run-all clean --parallel build:code build:types, a rollup and a tsc into one dist. The cold row is 39 of those at eight workers on four cores; the warm rows are the runner.

buildvxNx 20.8.4
cold (caches + outputs wiped)238.2 s245.7 s (1.03×)
warm, outputs wiped (restore)1.61 s3.94 s (2.44×)
warm, nothing wiped (no-op)341 ms4.02 s (11.8×)
second no-op349 ms4.07 s (11.7×)

novuhq/novu (37 build tasks, pnpm 11, Nx 21.3.11, parallel: 4)

Section titled “novuhq/novu (37 build tasks, pnpm 11, Nx 21.3.11, parallel: 4)”

Scoped as the repo’s own build script (nextjs and nestjs excluded). The builds carry their prebuild / postbuild hooks under both tools (STATUS 144); the first vx cold rep read 335 s against 290–291 s for the other two, the disk’s slow phase, absorbed by the median.

buildvxNx 21.3.11
cold (caches + outputs wiped)291.2 s299.4 s (1.03×)
warm, outputs wiped (restore)3.10 s9.05 s (2.92×)
warm, nothing wiped (no-op)655 ms8.55 s (13.1×)
second no-op642 ms8.69 s (13.5×)

TanStack/router (85 build + test:build tasks, pnpm 11, Nx 23.2.0, parallel: 5)

Section titled “TanStack/router (85 build + test:build tasks, pnpm 11, Nx 23.2.0, parallel: 5)”

Scoped as the repo’s own build script (examples/** and e2e/** excluded): 43 build (vite build) and 42 test:build (publint + attw --pack) over 42 packages and a benchmark. Seven packages reach their core only through peerDependencies; the first vx cold run built router-devtools-core before router-core’s dist existed and failed, which is STATUS 149 (a workspace peer orders the build unless it closes a cycle) — the rows below are on that fix. Nx 23 keeps its cache per user outside the repo; the harness pins it inside so the cold arm is cold (REPOS.md).

build test:buildvxNx 23.2.0
cold (caches + outputs wiped)199.9 s226.7 s (1.13×)
warm, outputs wiped (restore)1.00 s3.33 s (3.32×)
warm, nothing wiped (no-op)505 ms3.21 s (6.4×)
second no-op513 ms3.30 s (6.4×)

refinedev/refine (35 build tasks, pnpm 9, Nx 18.2.2, parallel: 3)

Section titled “refinedev/refine (35 build tasks, pnpm 9, Nx 18.2.2, parallel: 3)”

The library builds only (tsup && node ../shared/generate-declarations.js): examples/** out as the repo’s own scope, and the two Next apps out of both tools — refine-ui fetches Google Fonts through next/font (no egress here) and live-previews is one 280 s next build that would be the whole cold arm (REPOS.md). No parallel in nx.json, so both tools run at Nx’s default of 3.

buildvxNx 18.2.2
cold (caches + outputs wiped)104.2 s115.3 s (1.11×)
warm, outputs wiped (restore)636 ms2.23 s (3.5×)
warm, nothing wiped (no-op)184 ms2.11 s (11.5×)
second no-op183 ms2.18 s (11.9×)

Where a real no-op goes (refine, VX_TIMING=1, 180 ms wall): ~15 ms of runtime boot, 27 ms startup, 34 ms of git enumeration (two spawns scoped to the 35 packages), 34 ms deriving the 35 keys and probing them in one query, 42 ms proving the outputs intact — 6,790 files under 35 dist/**, stat’ed at ~6 µs each, the one cost that scales with the repo’s output size rather than its task count — and 5 ms of history (STATUS 153).

Each @vzn/vx-migrate plugin run live on a public repo against the repo’s own tool, same box (4 cores, Bun 1.4.2), both under unshare -n, medians of three interleaved reps. vx runs the tasks the plugin maps with no vx.config written.

kindspells/astro-shield (moon 1.41.7, moon(), 6cb8dff)

Section titled “kindspells/astro-shield (moon 1.41.7, moon(), 6cb8dff)”

astro-shield:build and :lint (lint.biome, lint.tsc, lint.publint after build): the same four commands under both tools. moon runs with MOON_TOOLCHAIN_FORCE_GLOBALS=1 (no toolchain download); with the network up, its warm run here was a flat 4.2 s, its version check timing out behind this box’s proxy, so the table leaves the network out for both.

build + lintvxmoon 1.41.7
cold (caches + outputs wiped)2.64 s3.78 s (1.43×)
warm, outputs wiped (restore)125 ms2.46 s (19.7×)
warm, nothing wiped (no-op)111 ms2.37 s (21.4×)

moon’s warm floor is its toolchain and dependency-install check (~2.1 s between its last cache write and the next task runner, in its own --log debug). The repo’s test.unit fetches over HTTPS and fails under vx on this box: the proxy’s CA reaches moon’s tasks with the whole environment and not vx’s isolated one, the gap moon() names in its note.

Roadmap 2.5: astro and refine again, same revisions, harnesses and scopes as above, now through the adoption plugins with nothing written (turbo() / nx() in vx.workspace.mjs). Compiled vx at main 889a95c, Bun 1.4.2, the 4-core container, medians of three interleaved reps. This box is slower than 2026-09-11’s (Nx cold 178 s here, 115 s then), so compare ratios, not absolute numbers.

astro build (32)vxTurbo 2.10.2
cold76.0 s86.4 s (1.14×)
restore1.74 s2.38 s (1.37×)
no-op786 ms1.86 s (2.36×)
second no-op792 ms2.02 s (2.55×)
refine build (35, + 14 nx-input twins)vxNx 18.2.2
cold158.8 s178.2 s (1.12×)
restore1.35 s2.83 s (2.10×)
no-op396 ms2.87 s (7.2×)
second no-op386 ms2.73 s (7.1×)

vx leads every row. refine’s warm ratios fell from ~11.5× because nx() maps the whole graph on every run: against the same repo with vx-migrate --from nx configs written, a warm no-op reads 396 ms median (min 364) under nx() and 291 (min 272) written, 15 interleaved rounds, A/A (a second nx() copy) 393 (min 350). Of a 377 ms no-op, load configs is 90 ms; the plugin maps all 206 projects when the run asks for 35.

Day-end A/B, main 889a95c against fa419f76 (item 960), 1,000 synthetic projects warm, source runs: a tie. n=15 medians 404 / 422, A/A 424; n=21 418 / 420, A/A 401. Stage mins over 15 runs move within the A/A spread.

Later the same day (main 3e927f9c, after the turbo() / nx() mapping cache and --filter discovering once), no-op, 15 interleaved rounds, A/A beside: astro 763 → 715 ms median (A/A 719). refine did not move (375 → 396, A/A 410, git defaults): nx()’s graph key now runs a whole-repo git status -uall each run (80–99 ms alone on refine), which ate what the mapping cache saved. Restore is not bound by worker count: refine at --concurrency 6 1,057 ms median against 1,148 at 3, A/A 1,183. Under a git config that weakens stat (core.checkStat=minimal, core.trustctime=false) every input hashes instead of trusting git’s OIDs, by design: refine’s no-op read 448 ms there against 396.

astro again on main 4b7c396a, after A-11 (nested-project boundaries by ancestor lookup; the same filtered no-op read 449 → 263 ms median against its parent, A/A 260), three interleaved reps:

astro build (32)vxTurbo 2.10.2
cold48.0 s57.9 s (1.21×)
restore478 ms1.43 s (2.98×)
no-op249 ms1.38 s (5.55×)
second no-op264 ms1.35 s (5.13×)

refine on the same commit, three interleaved reps. Nx itself ran faster on this box than at 889a95c (cold 178 → 106 s), so read the ratios; vx’s no-op fell 396 → 303 ms, while item 1075’s whole-repo git status still costs ~90 ms of it (a G lead):

refine build (35)vxNx 18.2.2
cold96.2 s105.9 s (1.10×)
restore1.06 s1.19 s (1.13×)
no-op303 ms1.16 s (3.84×)
second no-op271 ms1.19 s (4.40×)

Where vx’s own headroom went, on the same 1090-package / 3,270-node graph, fully cached (vx run build test --all):

MilestoneNo-restoreRestore
Set-closure scheduler priority (before)10.2 s—
+ bitset scheduler closure1.27 s1.59 s
+ discovery / package-graph fixes1.03 s1.34 s
+ frontier ^task expansion (v19, 8.5× fewer edges)0.62 s0.87 s

Input hashing then moved to git blob OIDs (v20, git ls-files -s): clean files cost zero reads/stats, dropping the warm run-phase from ~245 ms to ~76 ms (3.2×) at 500 projects × 30 files, and cold runs never read committed file contents at all. The decision history lives in git (the log was retired 2026-09-02); the shipped-optimization catalog with invariants is optimizations.md, and the engineering tour is comparison.md § Where vx is ahead.

Config evaluation was the largest fixed cost of a warm run (loadProjectConfig ~199 ms of a ~517 ms warm wall at 1,000 projects, 2026-09-02, before the day’s work). The resolved-config evaluation cache that was first rejected here shipped the same day behind the static purity gate it needed (caching.md § Config evaluation cache): a config whose import closure is provably pure is served validated from cache.db, keyed on the blob id of every file in the closure; anything the gate cannot prove evaluates live. load configs is 16–25 ms per 1,000 configs since. vx run --frozen is not a faster version of that path (see the head-to-head above, 2026-09-12): it is the env-independent one, and the eval-free one for impure configs.

Source vs binary. The runner invokes bun packages/vx/src/bin.ts by default, which pays ~40 ms of transpile per run that the --bytecode release binary does not (2026-09-09: 114 vs 71 ms on a two-package workspace; 20 projects warm 109 vs 64 ms). Set VX_BIN=<path> to time the shipped binary instead.