How the runner benchmarks work.
Every runner runs the same suites from the same harness. Values are medians, SaaS runners are priced at list price, RunsOn at EC2 spot with on-demand alongside, and every failure stays on the record.
Builds
Three real projects, built cold on every runner. The build time is the job's workload phases, wall clock.
- Rustspotify-player
-
- Cold build of aome510/spotify-player, the project Namespace uses for its own Rust caching benchmark: checkout, apt dependencies, toolchain install, cargo fmt, cargo test, and two cargo clippy passes.
- No cache.
- TypeScriptopencode
-
- Cold clone, bun install and typecheck of anomalyco/opencode, at the commit and Bun version ComputeSDK's DAX sandbox benchmark uses.
- Unlike DAX, typecheck runs
turbo typecheck --concurrency=4on every runner, whatever its size: turbo's default parallelism needs more than 16 GB and gets OOM-killed on runners without swap. - Totals are therefore not directly comparable with DAX.
- DockerPostHog image
-
- docker buildx build of PostHog's production Dockerfile from a pinned commit, with no build cache: frontend and plugin-server builds, Python dependencies, and layer assembly.
- The source is fetched first, shallow (depth 1, no tags), in its own timed 'fetch source' phase, and the build uses that local checkout as its context.
Hardware and disk
Synthetic scores explain the build times; they never set the order.
- Suite
- CPU (sysbench, 7-Zip, PassMark), memory bandwidth, disk (fio in the job workspace), network downloads, a docker pull, and pgbench against a PostgreSQL service container.
- Disk
-
- fio runs inside the job workspace, so it measures the disk a build writes to.
- 4k random reads at queue depth 32 and 1, sequential 1M reads, 4k writes each followed by fdatasync.
- Swap
- Each job records its swap size and kind, shown in GiB as measured; a runner without that record is not described. A sizing line cites swap only for runners at the RAM size it talks about.
- Location
-
- Download from GitHub and docker pull depend on where a runner sits. Each job's public address gives its country and region; RunsOn also reports its EC2 zone.
- Avrea: Uusimaa, FI; Hauts-de-France, FR
- Blacksmith: Arizona, US
- AWS CodeBuild: Virginia, US
- GitHub: Virginia, US; Iowa, US +10 more
- Namespace: Virginia, US
- RunsOn: Virginia, US (us-east-1a/1b)
- StarSling: California, US; Virginia, US
- Ubicloud: Saxony, DE; Hesse, DE
- Warpbuild: Virginia, US; Illinois, US +3 more (us-east-1a/1c/1d)
Queue
How fast a provider absorbs a spike of jobs.
- Measured
- Time from job creation to job start, from the GitHub Jobs API.
- Burst
- A burst of 15 identical jobs per runner, all queued at the same instant, each holding its runner for 60 seconds.
- Since 4 Oct 2026 each RunsOn job requests its own label (runs-on=<run id>-<runner id>-<iteration>), so a RunsOn runner launched for one job can no longer be taken by a sibling; before, a runner missing from a burst left one job waiting for RunsOn to retry, minutes later.
- Queue time shows how fast a provider absorbs a spike: how quickly it scales, and the concurrency limits of the plan this account is on.
- Runs on its own, never alongside another suite, in shards of one runner per provider, one shard after another (a provider's x64 and arm64 runners never share a shard, since some providers cap concurrency per account).
- Shown
- The p50 wait with the p10–p90 spread, across every job of every burst (15 per burst); a runner that sat out a burst, or lost jobs to a spot interruption, says so on its row. Profiles add how long until every job had started.
- Left out
- RunsOn bursts started before 4 Oct 2026, 08:54 UTC: every job of a RunsOn burst asked for the same label, so GitHub treated the burst's runners as interchangeable: a job could start on a runner launched for another, and a runner missing from the burst left one job waiting about three minutes for a retry. Since harness commit 52fd7b6, each job asks for its own label. Those runs stay visible, faded, under “Over time” on the RunsOn profile; they count nowhere. RunsOn runs this benchmark, so a result it leaves out is listed here with the reason.
- Plan limit
- Fewer jobs ran at once than were queued, and later ones started in waves as earlier ones finished (Avrea 8 vCPU, 12 at once; Namespace 2, 4 and 8 vCPU, 2 to 8 at once). That wait is the benchmark account's plan limit: shown, marked, never ranked.
Cache
One file through the cache, saved and restored on the same runner.
- Suite
- A 4 GB random file saved with a cache action, deleted, restored right away on the same runner, and checked to be complete:
- actions/cache (GitHub's cache service, or whatever the provider routes it to) and the provider's own cache action where the provider still recommends one (Blacksmith archived useblacksmith/cache: actions/cache is routed to its cache).
- No pause between save and restore, so a cache that isn't read-after-write consistent shows up as a failure.
- File generation is not timed.
- Actions
- actions/cache v6.1.0 and WarpBuilds/cache v2.0.0
- tmpfs
- tmpfs variants are shown but not ranked in cache and workspace disk comparisons: the workspace sits in RAM, which speeds up save and restore and is not a disk. Builds rank them as usual.
Price and cost
Pay-as-you-go rates: SaaS list prices matched to each runner's shape; RunsOn at EC2 prices. Variable cost only.
- Price
- Every provider is priced at its marginal pay-as-you-go rate, what the next job costs: no included minutes, plan credits, free allowances, prepaid units, reserved instances or savings plans. SaaS runners: list price per minute, matched on provider, architecture, tier, vCPU and memory, before any scaling; Namespace at its overage rate, $0.0015 per unit-minute. RunsOn: EC2 spot or on-demand prices, never reserved or savings-plan prices (see EC2 spot below).
- Cost per build
- Job duration × price per minute, the seconds as they ran, never rounded up to a billing increment.
- Fixed costs
- Cost per build is variable cost only. Licenses and plan fees, RunsOn's license included, are flat: listed once on each provider's page. What they add per job depends on how many jobs you run, so no single per-job share would be right; the calculator adds them at your volume.
- EC2 spot
-
- RunsOn runners are priced at EC2 spot, RunsOn's default: the 7-day average of the cheapest us-east-1 zone's spot price, fetched 5 Oct 2026. The RunsOn finder API records every zone's spot price every 90 minutes; the cheapest zone's at each fetch, averaged over the last 7 days, is what an hour costs without one day's swing or one historical low. It is not a median over the zones and days of the runs, and a spot price changes here only once it moves by 1% or more. The EBS volume is added.
- The on-demand figure is always shown alongside. With “RunsOn price: On-demand”, costs and cost ranks use it instead; every other provider has one list price.
- 292 RunsOn results are costed from RunsOn's own per-job EC2 and EBS cost, as each job ran; their price basis says “measured per job”. That cost covers the instance time RunsOn billed: from the runner's start, before the job, to a shutdown allowance after it, at least 60 seconds. A job that fell back to on-demand is costed at on-demand. The on-demand figure prices the same billed time at the on-demand rate, plus the job's EBS, so it is never below the measured one.
- 2371 jobs ran on spot and 578 on on-demand, RunsOn's fallback when spot instances are interrupted (its circuit breaker) or this benchmark account's spot quota is reached; 20 were interrupted, shown as failures (“spot interruption”).
- Scope
- A cost ratio compares a managed SaaS runner with a self-hosted runner on your own AWS account. Neither Warpbuild BYOC nor the effort of running a self-hosted stack was measured.
- Elsewhere
- The CPU page uses the same prices for CPU per dollar, each matched to the runner's exact shape, RunsOn at spot or on-demand.
EC2 CPU tab
EC2 instance types at 2 vCPU, scored and priced per region.
- Source
-
- One row per EC2 instance type at its 2 vCPU size, from one suite: the one that measured it in the most runs; a tie goes to the EC2 CPU sweep, then the hardware suite, then EC2 storage. ec2-cpu.json uses the same row.
- Today: 42 from the EC2 CPU sweep (1 run each) and 12 from the hardware suite (4–5 runs each). RunsOn's own runner types run in the hardware suite every few days, so the tab, the CPU page and the runner panels show the same figure.
- The EC2 CPU sweep runs once a month, so its window is 75 days (three sweeps); the other suites keep 30. Values are medians over the runs of the current workload definition.
- Capacity of the jobs behind the tab: EC2 CPU sweep: 0 of 42 jobs on spot, 42 on-demand; hardware suite: 35 of 60 jobs on spot, 25 on-demand.
- Ranks
- Every measured type is listed and sortable from its first run. A place (the # column, a leader, a ranked card or a stated finding) needs 3 runs and exists on the scores and the per-dollar figures only, never on a price, a band or a name; below 3 runs the row shows its run count. A type within 5% of the type heading its group shares that place (“=02”, “joint #2”), so near-identical scores are never ranked apart; a leader's near ties are named with it, and a lead needs 5% and no faster type still under 3 runs, or it reads “fastest measured”.
- Prices
- Per region (us-east-1, us-west-2, eu-west-1 and eu-central-1), from the RunsOn finder API (Linux), refreshed daily: spot is the 7-day average of the cheapest zone's spot price, the figure the runner pages price RunsOn with (see EC2 spot), and every saving, cheapest spot and CPU per dollar uses it. On-demand is the list price. The interruption band is AWS Spot Advisor's. A price the refresh could not fetch keeps the previous refresh's, dated; with none, it shows “—”, and a missing band “n/a”.
- CPU per dollar
- CPU Mark, or PassMark single-thread, per dollar of an hour of the instance at that spot price, the instance alone (the CPU page adds the EBS volume). Burstable t types are listed and ranked on CPU but left out of both: their score is a burst on CPU credits that an hour at that price does not sustain.
Ranks
Every size races; each provider counts once.
- Runs
- A run is one workflow run; the iterations inside it count once. A variant needs 3 runs to be its provider's bar or to take a place, a shade or a × ratio. Below that it shows its run count, unranked.
- Provider rank
- A runner is ranked against every other provider's best variant of the same arch, at any size, so a provider with five machines takes one place. The vCPU count beside every name shows when a bigger machine wins.
- Size toggle
- The benchmarks page opens at the reference size, 4 vCPU, so each provider fields one machine of the same size; the CPU and cache pages open at “Any size”, where every variant races. A size keeps those runners only and ranks within it; 8 to 12 vCPU count as one size (“8–12 vCPU”), since some machine families have no 8 vCPU shape. A size no other provider ran has no rank there.
- Suggested default
- On the RunsOn page: among RunsOn's 4 vCPU variants on the default EBS volume with a confirmed rate that finished all three builds, the one with the lowest cost across them at spot (or at the price the price pill picks), within 10% of the fastest one's build time.
- Shading
- The top fifth of providers in a column is tinted, the bottom fifth faint (at least one each, from three providers up). Bold is the best machine.
- Near ties
- A second provider within 5% of the leader, or inside its run-to-run spread, is named with it. Head-to-heads call values within 5% even.
- Clear lead
- A leader sentence or a highlighted bar also needs a 5% margin over the next provider, and no faster result still under 3 runs. A first place without that margin reads “fastest measured”.
- Columns
- 13 in the all-runners table: two summary columns, “Cost for all builds” and “Time for all builds”, then Rust, TypeScript, Docker, 1 thread, rand write, QD1 read, GitHub DL, cache restore, pgbench, queue and price.
- Summary columns
- “Cost for all builds” adds up a runner's three build costs, “Time for all builds” its three build times, only for runners that finished all three builds. The time is the table's default order, fastest first, ranked and shaded like a build column on the fewest runs of its three builds; runners missing a build follow, ordered by their Rust build. The cost is shaded cheap to expensive among the runners shown with 3+ runs.
Reading the charts
Bars rank a statistic; box plots show how far it moves.
- Multipliers
- Each ranked bar ends with its value over the top ranked bar's: ×1.08 is a time 8% longer, ×0.52 a disk with about half the IOPS. A bar under 3 runs shows its run count there instead. Build bars are split into phases in run order, lightest first.
- Timing
- Build bars time the workload steps. “Whole job” adds runner overhead; “Wait + job” adds the queue: each run's shortest wait, the median across runs, burst runs excluded. A plan-capped wait is shown, never ranked. A lead there needs 5% over every other bar.
- Best value
- The time-against-cost chart joins the runners no other runner beats on both, with a 2% noise margin. Only runners with 3 runs of every build in the view take part. Axes are linear from zero by default, log on request.
- Box plots
- One row per runner, every job a dot. The whisker spans min to max, the box p10 to p90, the tick is the p50; rows with the narrowest box come first, each with its n. A box needs 5 samples: a row with fewer is its dots and whisker, unranked.
Failures
A failed job is a result: it stays on the record and never takes a rank.
- Why
- Each failed cell says why: the runner's catalog note for that suite, else the job's own error message, else the phase it stopped in.
- Cell states
-
- failed: the job failed; the cell links to it.
- pending: still queued or running when the data was fetched.
- no runner: no runner picked up the job.
- –: the runner is not part of that suite.
- n/a: the step did not report the value.
- n/a with a ?, Docker build: not run on purpose: the PostHog Docker build needs more than 8 GB of memory; it ran out of memory on 8 GB runners. Never counted as a failure.
- n/a with a ?, Docker build: not run on purpose: needs more than 32 GB of RAM with tmpfs, where Docker's data root lives in RAM; PostHog's layers plus the build's own memory exceeded it and the runner was shut down mid-build. The 64 GB tmpfs runner builds it. Never counted as a failure.
Freshness and publication
- Medians
- Values are medians across the runs of the current workload definition; runs of an older definition are not aggregated.
- Schedule
-
- Builds every three days, on days 1, 4, 7… of the month: Rust at 02:00, TypeScript at 04:00, Docker at 06:00 UTC.
- Hardware (02:00) and cache (04:00 UTC) the day after; the burst alone the day after that, at 15:00 UTC.
- PassMark single-thread on every 2 vCPU runner daily at 08:00 UTC; EC2 storage on the 1st of each month, the EC2 CPU sweep on the 10th.
- This site picks up each run when the harness publishes it, and checks daily at 06:30 UTC.
- Stale data
- Once the newest run is more than 10 days old, the runner pages say so under their title.
- Publication
- Run by RunsOn, which sells one of the runners measured; the harness and raw results are public. The harness is open at runs-on-demo/benchmark-public, with every run and job log. A profile is indexed once its builds have 3 runs there.
- Raw data
- runners.json, plus one JSON file per provider, linked from each profile.
Methodology changelog
1 change to a workload definition, newest first: once a run with the change lands, earlier runs stop counting for that suite.