Skip to content

The whole machine, not one core at a time.

Pipelined compilation on every core under one jobserver, memory-aware scheduling, and a timeline of the chain that sets the pace.

faster than Cargo overall, across 32 measurements
2.8×
to find out nothing changed on a 249-package workspace (Cargo: 225 ms)
53 ms
Cargo's speed on a fresh checkout of a 25-crate workspace
108×

Measured, not promised.

Startup, builds from scratch, finding out nothing changed, rebuilds after an edit, cache reuse, tests and inspection commands — Rune against Cargo, measured with rune-bench.

Rune's own workspace · 249 packages rune cargo
Nothing changedbuild
53 ms
225 ms
Nothing changedcheck
64 ms
234 ms
Nothing changedtest --no-run
56 ms
272 ms
Dependency treetree
9 ms
184 ms
Full build from scratch
112.5 s
122.8 s
Fresh machine, shared cache
0.65 s
full rebuild
fresh checkout
108×
switching back to a branch
34×
running the tests
9×
faster overall, 32 measurements
2.8×

On a generated 25-crate workspace. Cold builds and rebuilds after an edit are on par: both run the same compiler — what changes is everything around it.

Every core, all the way to the end.

What stable Rust cannot parallelize is a chain of crates compiled one after another, and a single crate's front end. Everything else, Rune overlaps.

Pipelined

A crate starts the moment the .rmeta of each dependency exists, and the chain that sets the build's length goes first, weighted by measured compile times.

Build scripts as their own step

A build script starts once its build-dependencies are built, overlapping the rest of its package's dependencies instead of waiting behind them.

One jobserver for everything

About 1.3 compilers per core share one GNU jobserver: a compiler alone at the end takes the idle slots for its codegen threads, and the total never exceeds the budget.

Never behind codegen

Starting a job never waits on the jobserver pipe: codegen threads borrow idle tokens, and a job that starts while they hold them all owes one. The critical path never waits.

Memory-aware

Each compiler's peak memory is measured and remembered, so a crate known to need 2 GiB waits until 2 GiB are free. Many cores and little memory: no swapping, no OOM kills.

Polite when asked

build.low-priority = true runs every compiler at a lower priority so the rest of the machine stays responsive; rune watch always does.

See what sets the pace.

A build longer than 10 s that kept fewer than 60 % of the cores busy ends with one line naming the chain that held it up. Splitting the largest crate of that chain is usually what helps.

finding the slow part
rune build --timings  # every phase, and the chain in full  # .rune/timings/rune-timing.html: a timeline of every step  # .rune/fingerprints/trace.json: a Chrome tracerune explain  # which units the next build compiles, and whyrune doctor  # what slows your build down, with the fix

On a nightly compiler, build.frontend-threads = "auto" gives crates that start while cores are idle a parallel front end.

Measure it on your project.

Each tool builds its own copy, runs alternate between them so a warming machine favours neither, and reports give medians with their spread. Variables that could tilt the result — sccache included — are removed for both.

rune-bench
cargo build --release -p rune-cli -p rune-bench  # in Rune's repository./target/release/rune-bench run --project ~/src/my-app./target/release/rune-bench compare before/results.json after/results.json

How rune-bench measures

Try it on your project.

Nothing to migrate and nothing to undo: your project keeps working with plain cargo.

$ curl -fsSL https://www.runepm.com/install.sh | sh