Benchmarks
Empirical overhead numbers vs. Turborepo and Nx on synthetic workspaces. Updated as the runners evolve.
Warm-run overhead (2026-09-02)
Section titled “Warm-run overhead (2026-09-02)”The number that matters most to a developer is the warm no-op run: every
task a cache hit, nothing to restore. packages/vx-bench/generate.ts workspaces,
vx run build --all, this machine (macOS arm64, Bun 1.4.0), best of 5:
| Projects | Before (2026-09-02 morning) | Wave 1 | Wave 2 | Wave 5 | Wave 6 | Wave 7 | What changed |
|---|---|---|---|---|---|---|---|
| 100 | 105 ms | 92 ms | 79 ms | 78 ms | 77 ms | 74 ms | git overlaps config load; .git/HEAD read replaces a git spawn; batched probe; no output walk; git first |
| 1000 | 380–450 ms | 270 ms | 242 ms | 237 ms | 204 ms | 172 ms | + cached pure-config evals; one worktree walk; readdir discovery; batched probe; no output walk; git first |
Wave 7 (2026-09-03) starts the unscoped git enumeration the moment the
workspace root is known — ahead of the workspace config, discovery and
the cache open, ~25 ms earlier than before — so the status walk that is
the warm run’s wall floor overlaps everything. Interleaved A/B against a
worktree at the previous commit, three rounds of three: 1000 projects
min 195 → 172 ms.
Wave 6 (2026-09-03) removed the per-hit output walk: for dist/**-shaped
globs, the mtimes of every directory under the prefix, recorded at the
last save or restore, prove the output set unchanged (docs/caching.md
§ A current tree). Interleaved A/B against a worktree at the previous
commit, three rounds of three: 1000 projects min 224 → 204 ms.
Wave 3 (batched short-circuit probe, output rows carried on the entry,
memoised Bun.Globs) measured on the graph WITH dependencies —
vx run test --all, 2000 tasks, interleaved arms against an immutable
worktree of the previous commit: 327–329 ms → 308–314 ms.
Where Wave 2’s 242 ms at 1000 projects went (VX_TIMING=1, see
below; Waves 6 and 7 took 70 ms off it, mostly the run-graph phase and
the git overlap): discovery 22 ms, config load 31 ms (all cache hits) overlapped
with git’s one worktree walk (status -uall, ~57 ms, the critical
path), the run-graph phase 78 ms (1000 hits: probe, output glob, stat
check), history recording 12 ms, cache open 9 ms, and ~50 ms of process
start + module load + exit outside the table. The 1.7 s cold run is the
1000 cp commands.
Reproduce: bun packages/vx-bench/run.ts 1000 5.
A second machine, same shape (2026-09-20)
Section titled “A second machine, same shape (2026-09-20)”The table above is one machine’s. A four-core Linux container — sharing nothing with it — reads, medians with the full spread:
| Projects | Warm | Restore | Cold |
|---|---|---|---|
| 1000 | 271 ms (243–315) | 1031 ms (1023–1051) | 3147 ms (2905–3388) |
| 5000 | 807 ms (760–809) | 3931 ms (3462–4172) | 14 181 ms (13 798–14 986) |
Slower in absolute terms, as a shared container should be, and those numbers are not comparable to the ones above — different hardware. What does compare is the SCALING: 5× the projects costs 2.98× the warm run here against 2.97× on the other machine (restore 3.81× against 3.97×, cold 4.51× against 4.99×). The sub-linear warm curve is the code’s, not one box’s.
One number to take from this before optimising against the harness: the warm arm spreads ±13 % about its median on identical code, because every rep is a whole CLI invocation. An A/B here needs an effect bigger than that, and a control arm beside it.
The same container class on 2026-09-23, after the cold-path work of items 615 (a config round’s evaluations written once per table) and 622 (output-directory snapshots landed in one transaction), medians with the full spread, five reps at 1,000 and three at 5,000:
| Projects | Warm | Restore | Cold |
|---|---|---|---|
| 1000 | 239 ms (231–265) | 906 ms (831–1009) | 2634 ms (2418–2814) |
| 5000 | 711 ms (702–733) | 3253 ms (2993–3630) | 10 950 ms (10 896–11 547) |
Against the 2026-09-20 rows: cold −16 % at 1,000 and −23 % at 5,000, restore −12 % and −17 %; the warm rows moved within the ±13 % spread and claim nothing — the warm path is unchanged since item 589. Scaling 5× the projects: warm 2.97×, restore 3.59×, cold 4.16×.
Profiling a run
Section titled “Profiling a run”Two tools, and they answer different questions:
VX_TIMING=1 vx run …prints a stage table to stderr at the end of the run —startup,workspace config,discover projects,package graph,open cache,load configs,git enumeration,build graph,plugin stages,classify + probe,run graph,record history,output dir snapshots,close, and in a dry runplan— with each stage’s own and cumulative time, plus accumulated per-task spans (cache.get,output glob,output statandtask hashamong them;modules/timing.mdlists every one). This is the first thing to read: it says WHICH stage moved. The per-task spans run under the scheduler’s concurrency, so they over-count (a span’s wall includes time yielded to other tasks); compare them to each other, not to the stage total.bun --cpu-prof --cpu-prof-dir=/tmp/prof packages/vx/src/bin.ts run …thenbun packages/vx-bench/profile-summary.ts /tmp/prof/*.cpuprofilegives self time by function and by file. Good for finding a hot loop; unreliable about where anawaitwaited (it attributes the wait to whatever frame was on the stack).strace -f -o <file>around a run, thenbun packages/vx-bench/strace-vx.ts <file> [<before-file>]: the count of vx’s OWN syscalls (the tasks’ shells are told apart by PID), as a diff against the same run on agit worktreeof the base. A syscall count is deterministic where this container’s wall time is not: the restore path’s three spare round trips per artifact (item 627) were three rows of that table —readlink1,003 → 3,mkdir2,002 → 1,002,newfstatat−2,000 — and the save’s twomkdirs of the cache directory (630) one row, before either was a number. Awriteper round trip is the thread pool’s wake, so that row counts round trips.bun --preload ./packages/vx-bench/sqlite-tally.ts …is the same idea for SQLite statements;restore-bench.tsandsave-bench.tsin the same package restore or save every artifact sequentially, where a per-artifact change of tens of microseconds shows above the four-worker run’s noise (and one of a few microseconds does not: 630).
Three measurement lessons from this wave, recorded so they are not
re-learned. A compiled Bun 1.4.0 binary resolves on-disk packages by
<pkg>/index.ts only and ignores exports, so packages/vx-bench/compare.ts
measured nothing (“vx skipped”) until the packages gained root shims —
if the vx row ever reads n/a again, read the skip line first. A micro-benchmark of a sync call in isolation (statSync
2 µs vs stat 13 µs) does not predict the run — the async forms run in
parallel on the thread pool under the scheduler’s concurrency, and
switching the warm-hit path to sync calls made the 1000-project run
40 ms SLOWER. And Bun 1.4.0’s --compile binaries carry a signature this
macOS rejects (SIGKILL on launch); an ad-hoc codesign -s - --force
repairs it, which the release workflow now does on a macOS runner.
Head-to-head, 2026-09-03 (46 packages, packages/vx-bench/compare.ts 10 5 1)
Section titled “Head-to-head, 2026-09-03 (46 packages, packages/vx-bench/compare.ts 10 5 1)”Same workspace, identical commands, every runner pinned to concurrency
10, Turbo with no daemon (it uses none for turbo run since 2.9), vx as
its compiled binary. Median of 1,
this machine (macOS arm64, Bun 1.4.0). The harness then gave Nx npm
where Turbo had bun (§ Why Nx is slower), so the Nx row is slower than a
fair one. Every runner runs as in CI (CI=1), so Nx’s daemon is off:
| Runner | Version | Fresh (cold) | Warm (no restore) | Warm (restore) |
|---|---|---|---|---|
| vx | 0.0.0 | 10.45 s | 76 ms | 83 ms |
| vx (frozen) | 0.0.0 | 10.49 s | 83 ms | 88 ms |
| turbo | 2.10.12 | 10.58 s | 71 ms | 97 ms |
| nx | 23.2.0 | 19.66 s | 540 ms | 531 ms |
Read it honestly: at 46 packages Turborepo 2.10 and vx are within a few
milliseconds of each other on a fully-cached run, and neither keeps a
process between runs: Turbo 2.10 uses no daemon for turbo run (its docs
say so from 2.9), so both work out what changed on every invocation. vx wins
the restore case and ties the cold one; Nx is 7× off. The remaining
fixed cost at this size is process start + git, not the pipeline.
The same 46-package run on the four-core Linux container (2026-09-24,
item 735, the fixed harness: Nx runs its scripts with bun as Turbo does;
CI=1, so Nx’s daemon is off; a different machine, so only the ratios compare
with the table above). Turbo 2.11.3 (no daemon for turbo run) and Nx
23.2.1, vx as its compiled binary, median of 3. The CPU column is user +
system of the invocation and every child it waited for; a daemon that
outlives the invocation would not be counted, and none runs here:
| Runner | Version | Fresh (cold) | Warm (no restore) | Warm (restore) | CPU, cold |
|---|---|---|---|---|---|
| vx | 0.0.0 | 10.29 s | 79 ms | 104 ms | 846 ms |
| vx (frozen) | 0.0.0 | 10.27 s | 74 ms | 95 ms | 826 ms |
| turbo | 2.11.3 | 10.43 s | 86 ms (1.1×) | 136 ms (1.3×) | 1.33 s |
| nx | 23.2.1 | 22.08 s | 844 ms (10.7×) | 862 ms (8.3×) | 43.79 s |
The ideal schedule is 10.00 s, so vx and Turbo both sit on the critical path cold, and warm they are within a few milliseconds at this size (the 2026-09-23 run of the old harness read Turbo at 112 ms). The same box under the old harness read Nx at 27.21 s cold and 1m 2s of CPU.
The same harness at 476 packages / 1,428 graph nodes
(packages/vx-bench/compare.ts 20 25 1, 2026-09-02, same machine; a mid-size data
point — the committed packages/vx-bench/RESULTS.md is the 3,270-task run below):
| Runner | Fresh (cold) | Warm (no restore) | Warm (restore) |
|---|---|---|---|
| vx | 1m 40s | 297 ms | 416 ms |
| vx (frozen) | 1m 40s | 285 ms | 399 ms |
| turbo | 1m 40s | 342 ms (1.2×) | 612 ms (1.5×) |
| nx | 3m 23s | 1.38 s (4.7×) | 1.33 s (3.2×) |
The same size on the four-core Linux container (2026-09-25, after items 744, 753 and 754, the fixed harness, median of 1; ideal schedule 1m 36s):
| Runner | Fresh (cold) | Warm (no restore) | Warm (restore) | CPU, cold |
|---|---|---|---|---|
| vx | 1m 37s | 225 ms | 367 ms | 7.06 s |
| vx (frozen) | 1m 37s | 209 ms | 377 ms | 6.94 s |
| turbo | 1m 39s | 247 ms (1.1×) | 392 ms (1.1×) | 13.84 s |
| nx | 2m 27s | 1.84 s (8.2×) | 1.89 s (5.2×) | 7m 23s |
Read it honestly: the 2026-09-24 run on this box had Turbo 2.11 winning both warm columns (303 and 446 ms against vx’s 376 and 478). Item 744 found the cost in vx’s stable-key pass (a string set copied per task per dep, now a bitset), item 753 coalesced the 952 task lines into a write per turn of the event loop, and item 754 ranked a warm run’s tasks over the work that actually waits; vx now leads both warm columns, by 10% and 7% in one rep each, which this container’s noise can close. The min-of-N interleaved read of the warm no-op on the same workspace, with compiled binaries, is vx 165–175 ms against Turbo 226 ms (item 753’s profile).
Why Nx is slower
Section titled “Why Nx is slower”Two things, both per task, and both measured on the 46-package workspace (cold, 138 tasks, min of 3 interleaved arms, 2026-09-24, item 735):
| Arm | Wall | CPU |
|---|---|---|
the old harness: nx:run-script, Nx picked npm | 28.82 s | 69.80 s |
nx:run-script with bun (the fixed harness) | 21.37 s | 41.18 s |
nx:run-commands: the same commands, no fork | 11.65 s | 3.05 s |
- A Node process per task.
nx:run-scriptruns every task in a freshly forked Node process (nx/bin/run-executor.js: 138 of them in one cold run, counted withstrace -f -e execve) that loads Nx before it runs the script. That is ~270 ms of CPU per task. On four cores at concurrency 10 it saturates the CPU, so every task on the 20-task critical path waits for a core: 9.7 s of wall and 38 s of CPU, the gap between the last two rows.nx:run-commandsruns its command from the Nx process itself (NX_RUN_COMMANDS_DIRECTLY), and with it Nx lands at 11.65 s against an ideal of 10.00. A package-based Nx repo getsnx:run-scriptfor every inferredpackage.jsonscript. - npm, which the harness chose by accident.
nx:run-scriptruns the script through the package manager Nx detects from a lockfile. The generated workspace had none, so Nx fell back to npm, while Turbo readpackageManagerand ranbun run.npm runcosts 202 ms of CPU per task againstbun run’s 4 ms (min of 10, one task alone): 7.5 s of wall and 29 s of CPU, the gap between the first two rows. This one was the harness’s fault, not Nx’s.nx.jsonnow sayscli.packageManager: 'bun'.
Warm is spread thin, with no single cause. With no daemon (CI), each
invocation starts three plugin workers to build the project graph
(~280 ms each, in parallel; nx show projects alone is 0.56 s), then
hashes and replays the cached tasks. Turning the daemon on moves the
graph into it and saves 40–160 ms (0.81 s against 0.85, min of 5, in a
probe; 682 ms against 844, median of 3, in the harness), a fifth of the
warm gap at most.
The npm share grows with the graph. At 1,090 packages (3,270 tasks,
compare.ts 100 11, same box, one cold run each) Nx took 20m 39s and
76 min of CPU with npm, and 7m 22s and 23 min of CPU with bun: npm was
two thirds of the old harness’s Nx number there, which is the run the
site quoted. The site’s Nx column (34m 44s cold, § A real monorepo, the
macOS machine) paid npm per task and is not a fair one; its vx and Turbo
columns are unaffected. Next 18 re-runs it. The whole 3,270-task shape
on this box with the fixed harness (2026-09-25, after items 744, 753 and
754, median of 1; ideal schedule 3m 38s):
| Runner | Fresh (cold) | Warm (no restore) | Warm (restore) | CPU, cold |
|---|---|---|---|---|
| vx | 3m 40s | 359 ms | 653 ms | 16.05 s |
| vx (frozen) | 3m 41s | 292 ms | 651 ms | 16.46 s |
| turbo | 5m 4s (1.4×) | 431 ms (1.2×) | 722 ms (1.1×) | 33.27 s (2.1×) |
| nx | 6m 59s (1.9×) | 4.50 s (12.5×) | 4.60 s (7.0×) | 20m 55s (78.2×) |
The day before, Turbo 2.11 won both warm columns here (496 and 856 ms against vx’s 678 and 971). Item 744 found why: vx’s stable-key pass copied a string set per task per dep, now a bitset; items 753 and 754 then took the per-line stdout writes and the whole-graph priority closure off the warm path. vx now leads every column at this size, in one rep each.
Refuted, each within ±0.3 s of the fixed harness’s 21.2 s cold: the
per-task pseudo-terminal (NX_NATIVE_COMMAND_RUNNER=false) and the
output style (NX_TUI=false, --outputStyle=stream). So is the daemon:
the fixed harness read Nx cold at 22.25 s with NX_DAEMON=true and
22.08 s without. Parallelism is honoured: with cheap tasks, the
nx:run-commands arm sits 1.65 s over the ideal schedule. The ~4 git
probes each forked task runs cost ~8 ms of it. These tables used to say
Nx’s daemon was on; CI=1 had always turned it off, and the harness
keeps it off on purpose: it simulates CI.
A real monorepo: 3,270 tasks, 100 layers (2026-09-03)
Section titled “A real monorepo: 3,270 tasks, 100 layers (2026-09-03)”The shape that actually stresses a task runner: 100 dependency layers,
~11 packages per layer, ~30 deps per package, three tasks each
(build + installDeps + test, sleep 1 for build and test) — 3,270
task nodes, 1,090 packages. Same repo, same hardware, same task commands;
every runner pinned to concurrency 10. bun packages/vx-bench/compare.ts 100 11 1,
this machine (macOS arm64, 10 cores), Turbo 2.10.12, Nx 23.2.0.
The committed packages/vx-bench/RESULTS.md / packages/vx-bench/results.json are this run.
| vx | Turborepo | Nx | |
|---|---|---|---|
| Cold (nothing cached) | 3m 46s | 5m 13s (1.4×) | 34m 44s (9.2×) |
| Warm, nothing to rebuild | 510ms | 760ms (1.5×) | 3.59s (7.0×) |
| Warm, restore outputs | 777ms | 1.17s (1.5×) | 4.15s (5.3×) |
| CPU burned, cold (user+sys) | 34.61s | 1m 13s (2.1×) | 114m 06s (197.8×) |
| CPU burned, warm (user+sys) | 1.34s | 4.40s (3.3×) | 5.54s (4.1×) |
| Baseline (theoretical best) | 3m 38s cold; 0 warm, restore, CPU | — | — |
| Measured floors (context) | git walk 67ms · walk + raw copy 352ms · task shells 33.15s | — | — |
Baseline is the theoretical best case, so each row shows its overhead:
cold is the tasks’ own durations list-scheduled on 10 workers along the
exact dependency graph (critical path 1m 40s, total work ÷
workers 3m 38s); a cached run, a restore and the CPU a
runner burns are 0 in theory, so every measured number in those rows is
the runner. vx’s cold overhead over the ideal schedule is
8.45s on 3,270 tasks (8 ms per package); Turborepo’s is
1m 35s (88 ms per package) and Nx’s 31m 06s
(1,712 ms per package) — the number to read first, in one unit for every
runner: a runner that adds seconds to a three-minute build is a
different tool from one that adds half an hour. For context, the
measured floors row gives what the cheapest possible implementation
of each step costs on this machine: one git status -uall walk (the
cost of asking what changed), that walk plus a raw copy of every output
file, and the task shells themselves under xargs -P 10 (which vary by
about two seconds between runs; vx’s cold CPU sits within that noise).
CPU is user + system time of the invocation and every child it
waited for. The tasks are sleep, so this is the runner’s own work; a
daemon that outlives the invocation (Nx’s) is not counted, so Nx’s CPU
is a floor.
Methodology note: a synthetic graph with
sleep-based tasks isolates runner overhead from real compilation. All three runners are configured identically — same commands, the samesrc/**inputs anddist/**outputs, the same concurrency. (Hashing**/*instead would include each task’s own output in its inputs and break caching for everyone.) An earlier run of this shape (June 2026, a 4-core Linux box) read cold 3m 48s / 8m 18s / 8m 27s and CPU 22.7 s / 1,250 s / 2,038 s; cold wall time depends on how many cores the runners’ overhead competes with the tasks for, which is why the CPU row is the one that travels.
Reproducible head-to-head (vx vs Turborepo vs Nx)
Section titled “Reproducible head-to-head (vx vs Turborepo vs Nx)”packages/vx-bench/compare.ts scaffolds one shared monorepo matching the shape
above — layers × perLayer packages, ~30 deps each, three tasks
(build + installDeps + test) with the identical shell command,
src/** inputs, and dist/** outputs for every runner — then runs vx,
Turbo, and Nx across three cache states. Fairness is deliberate: vx runs
as the compiled binary real users install (not TS source); the
workspace is git-committed with node_modules/.turbo/.nx ignored;
every runner is pinned to the same concurrency; and runners are
measured strictly one at a time, daemons stopped between them, so they
never fight for CPU. build/test sleep 1 s so a warm hit visibly
skips the work.
bun packages/vx-bench/compare.ts # 100 layers × 11 (3,270 nodes) — the full shape (slow)bun packages/vx-bench/compare.ts 10 5 1 # 46 packages, 10 layers — quickBASELINE_ONLY=1 bun packages/vx-bench/compare.ts # recompute only the baseline floors against the committed rows (~9 min)bun packages/vx-bench/update-site.ts # rewrite the landing page, the README's bench sentence and chart, and this doc's stress section from results.json (--check to verify)BUILD_SLEEP=0 bun packages/vx-bench/compare.ts 20 11 2 # deep graph, pure framework overheadIt writes packages/vx-bench/RESULTS.md
(committed, so the numbers can be referenced from a commit).
Since 2026-09-03 the table also carries CPU columns and a baseline
row — the theoretical best case (an ideal schedule of the tasks, one git
walk, a raw copy of the outputs, the commands under xargs), so each
runner’s row reads as overhead above it. The definitions live in the
generated packages/vx-bench/RESULTS.md. The 46-package quick run is
the head-to-head table above (2026-09-03); an earlier run of that shape
used to sit here with a different verdict on the warm row, and one page
carrying both was a contradiction, so it is gone.
vx lock + --frozen is measured as its own row: it executes the
frozen vx-lock.json graph with zero per-run config evaluation. Read
the row as a tie, not a win: since the config-evaluation cache
(2026-09-02) the plain warm run evaluates nothing either for a config
the purity gate can prove pure, and it serves the same validated object
from cache.db without re-validating it, while --frozen parses the
whole lock and re-validates every entry (the lock is hand-editable, so
it is a boundary). What frozen still skips is the per-config identity
stat; what it still pays is the lock’s own parse. Measured 2026-09-12
on the 1,000-project bench, compiled binary, 12 interleaved reps:
plain min 154 / median 177 ms, frozen 148 / 165 — ~5%, the identity
stats; the load configs stage reads 20–25 ms plain against 6–10 ms
of lock read plus 12–14 ms of load frozen. The 83 vs 76 ms above
(2026-09-03, median of 1) is the same tie under a laptop’s noise.
--frozen is for what it guarantees — what runs is what was locked,
whatever a config would read from the environment — and for configs
the gate cannot prove pure, which evaluate live on every plain run and
come from the lock under --frozen. In your repo: vx lock, then
commit vx-lock.json.
How the overhead scales with the workspace (2026-09-10)
Section titled “How the overhead scales with the workspace (2026-09-10)”The per-package figure above is one size. This is vx alone at three
sizes of the same synthetic shape (packages/vx-bench/run.ts, the
generator’s workspace, one build per package whose command is a
mkdir and a cp, so the clock is the runner and almost nothing else),
the compiled Linux binary at commit 1a35ec3, median of 3 on a 4-core
Intel Xeon container:
| Packages | Cold (nothing cached) | per package | Warm, nothing to rebuild | per package | Warm, restore outputs | per package |
|---|---|---|---|---|---|---|
| 100 | 310 ms | 3.1 ms | 56 ms | 0.56 ms | 126 ms | 1.26 ms |
| 300 | 762 ms | 2.5 ms | 100 ms | 0.33 ms | 269 ms | 0.90 ms |
| 1,000 | 2,091 ms | 2.1 ms | 178 ms | 0.18 ms | 808 ms | 0.81 ms |
Ten times the packages costs 6.7× the cold time and 3.2× the warm time: the per-package cost falls as the fixed cost (process start, the workspace read, the cache open) is spread over more of them, and nothing in the run grows faster than the graph. A cold run at 1,000 packages is two seconds; a warm one is under two hundred milliseconds. These are the runner’s own costs on a trivial task; the 3,270-task table at the top, where each task sleeps a second and every runner is scheduled the same way, is where the same shape is compared against Turborepo and Nx.
A real Turbo repo: solidjs/solid (2026-09-10)
Section titled “A real Turbo repo: solidjs/solid (2026-09-10)”Not a synthetic workspace: solidjs/solid at b25c557 (5 packages,
pnpm 9, Turbo 2.10.10 as the repo’s own devDependency, Node 22), with
vx put on top of the repo’s own turbo.json through turbo() (then @vzn/vx-turbo, now @vzn/vx-migrate) —
a two-line vx.workspace.mjs, no config rewritten. Both tools see the
same graph: build is four executed tasks (solid-js#types, #link,
#build, solid-element#build; Turbo lists three more build nodes
for packages with no such script and runs nothing for them), and
test test-types is seven. Both restore the identical 64 output files.
vx as its compiled binary, Turbo with no daemon (2.10 uses none for
turbo run, so every run pays its own discovery, the same footing vx is
on; the --no-daemon the script passes is ignored), four
cores, Linux, arms interleaved. Medians (3 reps for build, 2 for
test); the script is packages/vx-bench/real/turbo-repo.sh.
build (4 tasks) | vx | Turbo 2.10.10 |
|---|---|---|
| cold (caches + outputs wiped) | 40.6 s | 45.5 s (1.12×) |
| warm, outputs wiped (restore) | 66 ms | 127 ms (1.9×) |
| warm, nothing wiped (no-op) | 51 ms | 95 ms (1.9×) |
test test-types (7 tasks) | vx | Turbo 2.10.10 |
|---|---|---|
| cold | 53.6 s | 58.2 s (1.09×) |
| warm, restore | 80 ms | 166 ms (2.1×) |
| warm, no-op | 59 ms | 93 ms (1.6×) |
Read it honestly: the cold rows are rollup, tsc and vitest — the
runner is a few percent of them, and the 4–5 s gap is Turbo’s per-task
work around the same commands (its ** default inputs hashed per
package, its log capture, its cache write), not measured to the frame
here. The warm rows are the product: with everything cached, vx
answers in 50–80 ms where Turbo takes 95–170 ms, and the restore case
— what a CI job or a fresh checkout does — is where the ratio is
widest. Neither tool has a daemon to turn on for this: Turbo’s docs
deprecate it for turbo run from 2.9, and vx has none.
Five real Turbo repos (2026-09-11)
Section titled “Five real Turbo repos (2026-09-11)”The same footing as the solid run, on the largest Turbo repos on GitHub:
the repo’s own turbo.json, vx on top through @vzn/vx-migrate’s turbo() with a
two-line vx.workspace.mjs, both tools scoped by the repo’s own
filters, four cores, Linux, arms interleaved, medians of three reps.
Both tools at Turbo’s default of 10 workers (vx’s default is the core
count; medusa’s own script says --concurrency=100% and both get it).
Turbo’s dry-run and vx’s --dry plan the same pkg#task set on every
repo. One binary for all five (the restore and warm-hit fixes of
STATUS 141 are in it). The script is packages/vx-bench/real/turbo-repo.sh;
noop2 is a second consecutive no-op. Every repo’s revision, toolchain,
scope and bench-side adjustment is in packages/vx-bench/real/REPOS.md.
What the harness does that the first attempt did not, all of it for
Turbo’s benefit as much as vx’s: every git-ignored artifact outside the
installs and the two caches is cleaned before a cold and a restore arm
(Turbo does not clean outputs; the first run leaked .turbo logs,
prebuilt files and 42 stale tsconfig.tsbuildinfo files between
arms); the repo’s root node_modules/.bin is on PATH as the repo’s own
yarn build would have it (bare, Turbo lost medusa’s rollup); npm
trusts this container’s proxy CA (cal.com’s embed build runs npx);
and astro’s build inputs exclude its own outputs (below).
Two tasks needed a bench-side output list, declared in the repo’s
vx.workspace.mjs as a ten-line project-stage plugin and named here so
nobody reads them as the plugin’s own mapping: medusa’s build outputs
are */** minus !src/** in turbo.json, which vx (no output negation)
runs uncached, so the bench names dist/** and .medusa/**; and
cal.com’s @calcom/web#build writes 110 symlinks to node_modules
directories under .next/node_modules, which vx’s artifact format does
not store, so the bench names the rest of .next — everything
next start reads — and Turbo’s artifact carries the 110 links too.
An artifact still stores no directory symlink; the save refuses one by name.
withastro/astro (32 build tasks, pnpm 10, Turbo 2.10.2)
Section titled “withastro/astro (32 build tasks, pnpm 10, Turbo 2.10.2)”astro’s build declares inputs: ["**/*", …], and an explicit Turbo
input glob matches the filesystem, gitignored or not — so as shipped,
each package’s own dist/** is in its hash. Turbo’s first run after a
restore then rebuilds the whole graph (59.6 s on the first harness),
and because 19 of the 32 builds are not byte-reproducible the run after
that rebuilds those 19 again (48.8 s, 13 cached), forever. Probed with
--dry=json: appending one byte to a gitignored dist/index.js changes
the task hash. vx excludes a task’s declared outputs from its inputs on
the same config. The table is on the fixed config — !dist/**/* and
!src/**/*.prebuilt* added to the build, build:ci and prebuild
inputs — so Turbo is measured at its best.
build | vx | Turbo 2.10.2 |
|---|---|---|
| cold (caches + outputs wiped) | 52.0 s | 62.9 s (1.21×) |
| warm, outputs wiped (restore) | 887 ms | 1.58 s (1.78×) |
| warm, nothing wiped (no-op) | 627 ms | 1.17 s (1.86×) |
| second no-op | 632 ms | 1.18 s (1.86×) |
payloadcms/payload (45 build tasks, pnpm 10, Turbo 2.10.4)
Section titled “payloadcms/payload (45 build tasks, pnpm 10, Turbo 2.10.4)”Scoped as the repo’s own build:all (the four templates excluded).
Turbo’s default inputs (the git-tracked files), no explicit glob.
build | vx | Turbo 2.10.4 |
|---|---|---|
| cold (caches + outputs wiped) | 126.9 s | 127.5 s (1.00×) |
| warm, outputs wiped (restore) | 3.44 s | 3.46 s (1.01×) |
| warm, nothing wiped (no-op) | 256 ms | 237 ms (0.93×) |
| second no-op | 275 ms | 267 ms (0.97×) |
Parity, and the doc says so: the no-op rows are within noise of each
other and Turbo takes both. The difference in what the two runs DO is
not noise: on every hit vx loads the 14,430 recorded output rows and
stats every file (~36 ms) to prove the outputs are intact, Turbo checks
nothing on disk — delete a file under dist and turbo run build still
prints a hit. Before STATUS 141 this repo read 129 s / 6.3 s / 372 ms /
325 ms for vx.
medusajs/medusa (83 build + build:plugin tasks, yarn 3, Turbo 1.13.4)
Section titled “medusajs/medusa (83 build + build:plugin tasks, yarn 3, Turbo 1.13.4)”The repo’s own --concurrency=100% for both. 24k tracked files.
build build:plugin | vx | Turbo 1.13.4 |
|---|---|---|
| cold (caches + outputs wiped) | 308 s | 315 s (1.02×) |
| warm, outputs wiped (restore) | 3.86 s | 7.16 s (1.85×) |
| warm, nothing wiped (no-op) | 947 ms | 3.29 s (3.5×) |
| second no-op | 953 ms | 3.12 s (3.3×) |
Before STATUS 141 vx’s no-op here was 2.5 s: every task carried the
same globalDependencies literal and resolved it against the whole
enumeration, 76 of 83 were hashed twice, and the absent .medusa/**
prefix refused every directory snapshot.
n8n-io/n8n (70 build tasks, pnpm 12, Turbo 2.9.18, Node 24)
Section titled “n8n-io/n8n (70 build tasks, pnpm 12, Turbo 2.9.18, Node 24)”build | vx | Turbo 2.9.18 |
|---|---|---|
| cold (caches + outputs wiped) | 137.2 s | 133.7 s (0.97×) |
| warm, outputs wiped (restore) | 8.27 s | 12.5 s (1.52×) |
| warm, nothing wiped (no-op) | 842 ms | 1.36 s (1.62×) |
| second no-op | 839 ms | 1.34 s (1.60×) |
The cold row is n8n-nodes-base#build (55 s alone) plus what fits
around it; a first vx rep read 193 s in the disk’s slow phase and the
median absorbed it.
calcom/cal.com (13 build tasks, yarn 3, Turbo 2.7.1, scope @calcom/web...)
Section titled “calcom/cal.com (13 build tasks, yarn 3, Turbo 2.7.1, scope @calcom/web...)”Three of the thirteen are cache: false in turbo.json (prisma’s
generate among them, ~12 s together) and run on every arm under both
tools, so the warm rows have a 12 s floor. .env from the example with
the two empty secrets filled, SKIP_DB_MIGRATIONS=1, no database.
build | vx | Turbo 2.7.1 |
|---|---|---|
| cold (caches + outputs wiped) | 245.7 s | 250.7 s (1.02×) |
| warm, outputs wiped (restore) | 17.4 s | 19.9 s (1.14×) |
| warm, nothing wiped (no-op) | 14.7 s | 18.5 s (1.26×) |
| second no-op | 14.5 s | 17.8 s (1.22×) |
Wide graphs (2026-09-11, one rep, 3 workers)
Section titled “Wide graphs (2026-09-11, one rep, 3 workers)”Each repo’s whole task set, not only build: the graph a team actually
runs. One rep per arm and both tools at 3 workers, because the session’s
memory cgroup allows 13.3 GiB and the wide sets’ typecheck and lint
processes run at 3.6–5.7 GB resident each. Same harness, same cleanup,
same scope as the build tables. Two sets ran; three could not, and the
reasons are the repos’ own (below), the same under both tools.
payloadcms/payload — build lint, 89 tasks. payload’s lint is
cache: false in its own turbo.json, so the 44 lint tasks run on every
arm under both tools (~225 s at 3 workers); the three warm rows are
that floor, and the cold row is the floor plus the 45 builds.
build lint | vx | Turbo 2.10.4 |
|---|---|---|
| cold (caches + outputs wiped) | 318 s | 334 s (1.05×) |
| warm, outputs wiped (restore) | 229 s | 223 s (0.97×) |
| warm, nothing wiped (no-op) | 227 s | 228 s (1.01×) |
| second no-op | 226 s | 226 s (1.00×) |
calcom/cal.com — build lint, 24 tasks, scope @calcom/web....
type-check is cache: false in the repo’s turbo.json and stays out;
the three uncached build tasks (~12 s) run on every arm as in the build
table, so the warm rows are the runner plus that floor.
build lint | vx | Turbo 2.7.1 |
|---|---|---|
| cold (caches + outputs wiped) | 246.0 s | 236.8 s (0.96×) |
| warm, outputs wiped (restore) | 16.2 s | 19.7 s (1.22×) |
| warm, nothing wiped (no-op) | 14.3 s | 17.7 s (1.23×) |
| second no-op | 14.3 s | 18.0 s (1.26×) |
medusajs/medusa — build build:plugin test, 157 tasks: dropped.
medusa’s test declares no dependsOn, so on a cold tree both tools
start tests before the packages they import are built: vx lost
@medusajs/auth#test (Cannot find module '@medusajs/framework/awilix',
and it passed on the restore arm once the build existed), Turbo lost
@medusajs/dashboard#test (Failed to resolve entry for package "@medusajs/admin-vite-plugin") on all four arms. A repo configuration
gap the two runners expose identically; not a number for either.
withastro/astro — build test, 55 tasks: dropped, the same gap:
test depends on ^test only, and its tests import their own package’s
dist. Turbo lost @astrojs/internal-helpers#test, upgrade#test and
telemetry#test on every arm; vx’s scheduler happened to run the builds
first and lost @astrojs/language-server#test, which imports
packages/astro/dist across packages, plus @astrojs/ts-plugin#test,
which downloads VS Code and cannot behind this proxy.
n8n-io/n8n — build typecheck lint, 220 tasks: dropped (owner). The
editor-ui’s vue-tsc and eslint at 3.6–5.7 GB resident each do not fit
the cgroup beside anything else.
Real Nx repos (2026-09-11)
Section titled “Real Nx repos (2026-09-11)”The Turbo footing for Nx: the repo’s own Nx (NX_DAEMON=false,
NX_NO_CLOUD=true) against vx on the vx.config files
bunx @vzn/vx-migrate --from nx wrote from the repo’s exported graph,
both tools at the worker count the repo’s nx.json sets, medians of
three interleaved reps, the same cleanup and arm logs as
turbo-repo.sh. The script is packages/vx-bench/real/nx-repo.sh —
the Turbo harness runs the Turbo repos and this one runs these, and
each names the other only for the footing they share. Parity is the
task graph: nx run-many … --graph
against vx --dry=json plan the same project#target set. Every
revision, toolchain and bench-side rule (the configs rewritten to
.mjs, the pinned package manager on PATH) is in
packages/vx-bench/real/REPOS.md § Nx repos.
TanStack/query (25 build tasks, pnpm 11, Nx 22.1.3, parallel: 5)
Section titled “TanStack/query (25 build tasks, pnpm 11, Nx 22.1.3, parallel: 5)”Scoped as the repo’s own build script (examples/** and
integrations/** excluded). Every target is nx:run-script.
build | vx | Nx 22.1.3 |
|---|---|---|
| cold (caches + outputs wiped) | 47.4 s | 55.1 s (1.16×) |
| warm, outputs wiped (restore) | 656 ms | 1.98 s (3.02×) |
| warm, nothing wiped (no-op) | 174 ms | 1.88 s (10.8×) |
| second no-op | 156 ms | 1.89 s (12.1×) |
strapi/strapi (39 build tasks, yarn 4, Nx 20.8.4, parallel: 8, --nx-ignore-cycles)
Section titled “strapi/strapi (39 build tasks, yarn 4, Nx 20.8.4, parallel: 8, --nx-ignore-cycles)”Every package script is yarn’s run -T <root bin>, so each task is
yarn run build under both tools (STATUS 142); the per-package
build is npm-run-all clean --parallel build:code build:types,
a rollup and a tsc into one dist. The cold row is 39 of those at
eight workers on four cores; the warm rows are the runner.
build | vx | Nx 20.8.4 |
|---|---|---|
| cold (caches + outputs wiped) | 238.2 s | 245.7 s (1.03×) |
| warm, outputs wiped (restore) | 1.61 s | 3.94 s (2.44×) |
| warm, nothing wiped (no-op) | 341 ms | 4.02 s (11.8×) |
| second no-op | 349 ms | 4.07 s (11.7×) |
novuhq/novu (37 build tasks, pnpm 11, Nx 21.3.11, parallel: 4)
Section titled “novuhq/novu (37 build tasks, pnpm 11, Nx 21.3.11, parallel: 4)”Scoped as the repo’s own build script (nextjs and nestjs
excluded). The builds carry their prebuild / postbuild hooks under
both tools (STATUS 144); the first vx cold rep read 335 s against
290–291 s for the other two, the disk’s slow phase, absorbed by the
median.
build | vx | Nx 21.3.11 |
|---|---|---|
| cold (caches + outputs wiped) | 291.2 s | 299.4 s (1.03×) |
| warm, outputs wiped (restore) | 3.10 s | 9.05 s (2.92×) |
| warm, nothing wiped (no-op) | 655 ms | 8.55 s (13.1×) |
| second no-op | 642 ms | 8.69 s (13.5×) |
TanStack/router (85 build + test:build tasks, pnpm 11, Nx 23.2.0, parallel: 5)
Section titled “TanStack/router (85 build + test:build tasks, pnpm 11, Nx 23.2.0, parallel: 5)”Scoped as the repo’s own build script (examples/** and e2e/**
excluded): 43 build (vite build) and 42 test:build (publint +
attw --pack) over 42 packages and a benchmark. Seven packages reach
their core only through peerDependencies; the first vx cold run
built router-devtools-core before router-core’s dist existed and
failed, which is STATUS 149 (a workspace peer orders the build unless
it closes a cycle) — the rows below are on that fix. Nx 23 keeps its
cache per user outside the repo; the harness pins it inside so the
cold arm is cold (REPOS.md).
build test:build | vx | Nx 23.2.0 |
|---|---|---|
| cold (caches + outputs wiped) | 199.9 s | 226.7 s (1.13×) |
| warm, outputs wiped (restore) | 1.00 s | 3.33 s (3.32×) |
| warm, nothing wiped (no-op) | 505 ms | 3.21 s (6.4×) |
| second no-op | 513 ms | 3.30 s (6.4×) |
refinedev/refine (35 build tasks, pnpm 9, Nx 18.2.2, parallel: 3)
Section titled “refinedev/refine (35 build tasks, pnpm 9, Nx 18.2.2, parallel: 3)”The library builds only (tsup && node ../shared/generate-declarations.js):
examples/** out as the repo’s own scope, and the two Next apps out of
both tools — refine-ui fetches Google Fonts through next/font
(no egress here) and live-previews is one 280 s next build that
would be the whole cold arm (REPOS.md). No parallel in nx.json,
so both tools run at Nx’s default of 3.
build | vx | Nx 18.2.2 |
|---|---|---|
| cold (caches + outputs wiped) | 104.2 s | 115.3 s (1.11×) |
| warm, outputs wiped (restore) | 636 ms | 2.23 s (3.5×) |
| warm, nothing wiped (no-op) | 184 ms | 2.11 s (11.5×) |
| second no-op | 183 ms | 2.18 s (11.9×) |
Where a real no-op goes (refine, VX_TIMING=1, 180 ms wall): ~15 ms
of runtime boot, 27 ms startup, 34 ms of git enumeration (two spawns
scoped to the 35 packages), 34 ms deriving the 35 keys and probing
them in one query, 42 ms proving the outputs intact — 6,790 files
under 35 dist/**, stat’ed at ~6 µs each, the one cost that scales
with the repo’s output size rather than its task count — and 5 ms of
history (STATUS 153).
Adoption paths on real repos (2026-09-28)
Section titled “Adoption paths on real repos (2026-09-28)”Each @vzn/vx-migrate plugin run live on a public repo against the
repo’s own tool, same box (4 cores, Bun 1.4.2), both under
unshare -n, medians of three interleaved reps. vx runs the tasks the
plugin maps with no vx.config written.
kindspells/astro-shield (moon 1.41.7, moon(), 6cb8dff)
Section titled “kindspells/astro-shield (moon 1.41.7, moon(), 6cb8dff)”astro-shield:build and :lint (lint.biome, lint.tsc,
lint.publint after build): the same four commands under both tools.
moon runs with MOON_TOOLCHAIN_FORCE_GLOBALS=1 (no toolchain download);
with the network up, its warm run here was a flat 4.2 s, its version
check timing out behind this box’s proxy, so the table leaves the
network out for both.
build + lint | vx | moon 1.41.7 |
|---|---|---|
| cold (caches + outputs wiped) | 2.64 s | 3.78 s (1.43×) |
| warm, outputs wiped (restore) | 125 ms | 2.46 s (19.7×) |
| warm, nothing wiped (no-op) | 111 ms | 2.37 s (21.4×) |
moon’s warm floor is its toolchain and dependency-install check
(~2.1 s between its last cache write and the next task runner, in its
own --log debug). The repo’s test.unit fetches over HTTPS and fails
under vx on this box: the proxy’s CA reaches moon’s tasks with the whole
environment and not vx’s isolated one, the gap moon() names in its
note.
Real repos re-measured (2026-09-27)
Section titled “Real repos re-measured (2026-09-27)”Roadmap 2.5: astro and refine again, same revisions, harnesses and
scopes as above, now through the adoption plugins with nothing written
(turbo() / nx() in vx.workspace.mjs). Compiled vx at main 889a95c,
Bun 1.4.2, the 4-core container, medians of three interleaved reps.
This box is slower than 2026-09-11’s (Nx cold 178 s here, 115 s then),
so compare ratios, not absolute numbers.
astro build (32) | vx | Turbo 2.10.2 |
|---|---|---|
| cold | 76.0 s | 86.4 s (1.14×) |
| restore | 1.74 s | 2.38 s (1.37×) |
| no-op | 786 ms | 1.86 s (2.36×) |
| second no-op | 792 ms | 2.02 s (2.55×) |
refine build (35, + 14 nx-input twins) | vx | Nx 18.2.2 |
|---|---|---|
| cold | 158.8 s | 178.2 s (1.12×) |
| restore | 1.35 s | 2.83 s (2.10×) |
| no-op | 396 ms | 2.87 s (7.2×) |
| second no-op | 386 ms | 2.73 s (7.1×) |
vx leads every row. refine’s warm ratios fell from ~11.5× because
nx() maps the whole graph on every run: against the same repo with
vx-migrate --from nx configs written, a warm no-op reads 396 ms
median (min 364) under nx() and 291 (min 272) written, 15
interleaved rounds, A/A (a second nx() copy) 393 (min 350). Of a
377 ms no-op, load configs is 90 ms; the plugin maps all 206
projects when the run asks for 35.
Day-end A/B, main 889a95c against fa419f76 (item 960), 1,000 synthetic projects warm, source runs: a tie. n=15 medians 404 / 422, A/A 424; n=21 418 / 420, A/A 401. Stage mins over 15 runs move within the A/A spread.
Later the same day (main 3e927f9c, after the turbo() / nx() mapping
cache and --filter discovering once), no-op, 15 interleaved rounds,
A/A beside: astro 763 → 715 ms median (A/A 719). refine did not move
(375 → 396, A/A 410, git defaults): nx()’s graph key now runs a
whole-repo git status -uall each run (80–99 ms alone on refine), which
ate what the mapping cache saved. Restore is not bound by worker count:
refine at --concurrency 6 1,057 ms median against 1,148 at 3, A/A
1,183. Under a git config that weakens stat (core.checkStat=minimal,
core.trustctime=false) every input hashes instead of trusting git’s
OIDs, by design: refine’s no-op read 448 ms there against 396.
astro again on main 4b7c396a, after A-11 (nested-project boundaries by ancestor lookup; the same filtered no-op read 449 → 263 ms median against its parent, A/A 260), three interleaved reps:
astro build (32) | vx | Turbo 2.10.2 |
|---|---|---|
| cold | 48.0 s | 57.9 s (1.21×) |
| restore | 478 ms | 1.43 s (2.98×) |
| no-op | 249 ms | 1.38 s (5.55×) |
| second no-op | 264 ms | 1.35 s (5.13×) |
refine on the same commit, three interleaved reps. Nx itself ran
faster on this box than at 889a95c (cold 178 → 106 s), so read the
ratios; vx’s no-op fell 396 → 303 ms, while item 1075’s whole-repo
git status still costs ~90 ms of it (a G lead):
refine build (35) | vx | Nx 18.2.2 |
|---|---|---|
| cold | 96.2 s | 105.9 s (1.10×) |
| restore | 1.06 s | 1.19 s (1.13×) |
| no-op | 303 ms | 1.16 s (3.84×) |
| second no-op | 271 ms | 1.19 s (4.40×) |
Performance history
Section titled “Performance history”Where vx’s own headroom went, on the same 1090-package / 3,270-node graph,
fully cached (vx run build test --all):
| Milestone | No-restore | Restore |
|---|---|---|
| Set-closure scheduler priority (before) | 10.2 s | — |
| + bitset scheduler closure | 1.27 s | 1.59 s |
| + discovery / package-graph fixes | 1.03 s | 1.34 s |
+ frontier ^task expansion (v19, 8.5× fewer edges) | 0.62 s | 0.87 s |
Input hashing then moved to git blob OIDs (v20, git ls-files -s): clean
files cost zero reads/stats, dropping the warm run-phase from ~245 ms to
~76 ms (3.2×) at 500 projects × 30 files, and cold runs never read
committed file contents at all. The decision history lives in git (the log was retired 2026-09-02);
the shipped-optimization catalog with invariants is
optimizations.md, and the engineering tour is
comparison.md § Where vx is ahead.
Known headroom
Section titled “Known headroom”Config evaluation was the largest fixed cost of a warm run
(loadProjectConfig ~199 ms of a ~517 ms warm wall at 1,000 projects,
2026-09-02, before the day’s work). The resolved-config evaluation cache
that was first rejected here shipped the same day behind the static
purity gate it needed (caching.md § Config evaluation cache): a config
whose import closure is provably pure is served validated from
cache.db, keyed on the blob id of every file in the closure; anything
the gate cannot prove evaluates live. load configs is 16–25 ms per
1,000 configs since. vx run --frozen is not a faster version of that
path (see the head-to-head above, 2026-09-12): it is the env-independent
one, and the eval-free one for impure configs.
Source vs binary. The runner invokes bun packages/vx/src/bin.ts by default, which pays ~40 ms of transpile per run that the --bytecode release binary does not (2026-09-09: 114 vs 71 ms on a two-package workspace; 20 projects warm 109 vs 64 ms). Set VX_BIN=<path> to time the shipped binary instead.