Fastest GitHub Actions runners: disk I/O.
fio inside the job workspace of 89 runners from 9 providers, 2 to 12 vCPU. Latency first: a build waits on small reads and synced writes far more often than it streams one big file.
How this shows up in real builds
Latency per provider
Each provider's best disk. QD1 is one 4k read at a time, the way a build reads thousands of small files. tmpfs is RAM, not a disk: add it to see the gap on the same scale.
tmpfs for scale: 1,156,332 on RunsOn m8a.2xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
tmpfs for scale: 680,179 on RunsOn m9g.xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
What the numbers say
Extremes on x64, and RAM against disk.
- On x64, one small read at a time (4k, QD1) comes back 85,326 times a second on Blacksmith 4 vCPU and 1,569 on RunsOn m8i.2xlarge · EBS gp3, 54× apart; p99 latency is 14 µs against 1,012 µs.
- Synced 4k writes on x64 run from 495 per second on RunsOn m8i.2xlarge · EBS gp3 (workspace mounted nobarrier) to 20,425 on Blacksmith 4 vCPU (workspace mounted nobarrier).
- r8a.xlarge with tmpfs ($0.0021/min at spot) answers 1,071,527 QD1 reads and 431,794 synced writes per second: 88× and 36× the best local NVMe at 4 vCPU (i7i.xlarge, $0.0019) and 586× and 800× default EBS gp3 (m8a.xlarge, $0.0015); its workspace is 31 GiB, from RAM. (EC2 storage suite)
More findings20: EBS, arm64, NVMe vs EBS, provisioned gp3, tmpfs, storage classes
- Every EBS-backed workspace (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations) answers between 1,569 and 1,844 QD1 reads and 495 to 562 synced writes per second, whoever runs it and whatever throughput it is provisioned for: each request is a round trip to network storage.
- On arm64, one small read at a time (4k, QD1) comes back 26,866 times a second on Warpbuild xfast (Apple M4 Pro) 6 vCPU and 1,660 on RunsOn m9g.large · EBS gp3, 16× apart; p99 latency is 40 µs against 963 µs.
- Synced 4k writes on arm64 run from 537 per second on RunsOn c8g.large · EBS gp3 (workspace mounted nobarrier) to 17,054 on RunsOn m9gd.2xlarge · local NVMe.
- Same Xeon 6975P-C CPU, different disk: m8id.2xlarge on local NVMe answers 11,063 QD1 reads and 7,665 synced writes per second, m8i.2xlarge on EBS gp3 1,569 and 495.
- Same Xeon 6975P-C CPU, different disk: m8id.large on local NVMe answers 10,991 QD1 reads and 5,549 synced writes per second, m8i-flex.large on EBS gp3 1,745 and 528.
- Same Xeon 6975P-C CPU, different disk: m8id.xlarge on local NVMe answers 11,033 QD1 reads and 7,853 synced writes per second, m8i-flex.xlarge on EBS gp3 1,765 and 527.
- Same Neoverse-V3 CPU, different disk: m9gd.2xlarge on local NVMe answers 13,522 QD1 reads and 17,054 synced writes per second, m9g.2xlarge on EBS gp3 1,788 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.xlarge on local NVMe answers 11,717 QD1 reads and 11,264 synced writes per second, m9g.xlarge on EBS gp3 1,713 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.large on local NVMe answers 13,477 QD1 reads and 7,176 synced writes per second, m9g.large on EBS gp3 1,660 and 542.
- Provisioning gp3 at 1000 MiB/s on m8a.2xlarge takes sequential reads from 402 MiB/s to 1,002 MiB/s (2.5×); QD1 reads barely move (1,791 → 1,792 IOPS): it buys bandwidth, not latency.
- Provisioning gp3 at 1000 MiB/s on m8azn.xlarge takes sequential reads from 401 MiB/s to 1,001 MiB/s (2.5×); QD1 reads barely move (1,775 → 1,775 IOPS): it buys bandwidth, not latency.
- tmpfs on m8a.2xlarge is not a disk at all: 1,156,332 QD1 reads and 432,842 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.2xlarge is not a disk at all: 653,143 QD1 reads and 301,463 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.xlarge is not a disk at all: 680,179 QD1 reads and 303,962 synced writes per second, from RAM. The workspace is 15 GiB, sized by memory.
- tmpfs on r8a.2xlarge is not a disk at all: 1,151,153 QD1 reads and 433,820 synced writes per second, from RAM. The workspace is 62 GiB, sized by memory.
- tmpfs: 653k–1.2M QD1 reads and 301k–434k synced writes per second across 5 runners (5 RunsOn configurations).
- VM disk: 5.9k–85k QD1 reads and 578–20k synced writes per second across 42 runners from 7 providers (6 Avrea runners, 6 Blacksmith runners, 3 GitHub runners, 11 Namespace runners, 2 StarSling x64 runners, 9 Ubicloud runners and 5 Warpbuild runners).
- local NVMe: 11k–14k QD1 reads and 5.5k–17k synced writes per second across 9 runners (9 RunsOn configurations).
- network disk: 6.8k–8.6k QD1 reads and 3.3k–4.7k synced writes per second across 2 runners (2 GitHub runners).
- EBS: 1.6k–1.8k QD1 reads and 495–562 synced writes per second across 31 runners from 3 providers (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations).
Reading these numbers
Most of a build is small, dependent I/O, not one big stream.
Why latency before throughput
-
git checkout,npm install,cargo buildand test runners read thousands of small files, each waiting for the last: the QD1 column. QD32 and sequential show what a parallel compiler, a bigdocker buildor a cache archive can pull; read them after. - Package managers, databases and SQLite-backed tests flush their writes. Each flush waits for the storage: network disks pay a round trip every time. That is sync write.
Every runner
Each provider's best disk first; its other machines sit behind +N more. Bars are log scale.
| One request at a time | 4k, 32 in flight | Sequential, 1M | |||||
|---|---|---|---|---|---|---|---|
| Runner | QD1 read ↑ | p99 ↓ | Sync write ↑ | Read ↑ | Write ↑ | Read ↑ | Write ↑ |
| x64 | |||||||
| Blacksmith | 14 µs | 873,112 | 254,574 | 6.9 GiB/s | 1.9 GiB/s | ||
| 8 vCPU | 19 µs | 422,170 | 207,013 | 6.0 GiB/s | 1.9 GiB/s | ||
| 2 vCPU | 18 µs | 563,990 | 149,081 | 5.1 GiB/s | 1.9 GiB/s | ||
| Namespace 8x32 | 13 µs | 1,112,926 | 412,276 | 4.2 GiB/s | 3.0 GiB/s | ||
| 8x16 | 12 µs | 1,124,263 | 476,969 | 4.1 GiB/s | 2.9 GiB/s | ||
| 4x16 | 12 µs | 1,011,926 | 436,056 | 4.1 GiB/s | 2.9 GiB/s | ||
| 4x8 | 13 µs | 1,015,510 | 403,484 | 4.1 GiB/s | 3.0 GiB/s | ||
| 2x8 | 13 µs | 350,533 | 222,621 | 4.1 GiB/s | 2.9 GiB/s | ||
| StarSling | 17 µs | 265,725 | 15,331 | 2.5 GiB/s | 2.0 GiB/s | ||
| 4 vCPU | 17 µs | 227,686 | 11,527 | 3.3 GiB/s | 2.8 GiB/s | ||
| Avrea | 18 µs | 439,265 | 213,725 | 5.9 GiB/s | 3.4 GiB/s | ||
| 8 vCPU | 23 µs | 245,805 | 139,439 | 7.4 GiB/s | 3.3 GiB/s | ||
| 2 vCPU | 17 µs | 454,466 | 184,016 | 5.4 GiB/s | 3.4 GiB/s | ||
| Warpbuild | 58 µs | 144,184 | 66,887 | 7.6 GiB/s | 4.6 GiB/s | ||
| 8 vCPU | 60 µs | 116,390 | 45,671 | 7.8 GiB/s | 4.5 GiB/s | ||
| 4 vCPU | 64 µs | 124,074 | 38,883 | 7.8 GiB/s | 5.1 GiB/s | ||
| RunsOn i7i.2xlarge | 228 µs | 300,466 | 165,110 | 2.0 GiB/s | 1.5 GiB/s | ||
| m8azn.3xlarge | 782 µs | 2,997 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| c8a.2xlarge | 840 µs | 2,996 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8a.2xlarge | 856 µs | 2,995 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8a.2xlarge | 897 µs | 15,996 | 15,992 | 1,001 MiB/s | 1,001 MiB/s | ||
| r8a.2xlarge | <1 µs | 4,231,165 | 3,828,376 | 18.6 GiB/s | 10.1 GiB/s | ||
| m8a.2xlarge | <1 µs | 4,258,470 | 3,787,945 | 18.4 GiB/s | 9.9 GiB/s | ||
| m8azn.xlarge | 963 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8azn.xlarge | 856 µs | 15,994 | 15,991 | 1,001 MiB/s | 1,001 MiB/s | ||
| m8id.2xlarge | 199 µs | 134,258 | 67,075 | 905 MiB/s | 432 MiB/s | ||
| m8i.2xlarge | 1,012 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m8a.xlarge | 840 µs | 2,994 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8i-flex.2xlarge | 848 µs | 2,991 | 2,995 | 401 MiB/s | 401 MiB/s | ||
| c8a.xlarge | 856 µs | 2,995 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| r8a.xlarge | <1 µs | 4,113,746 | 3,786,762 | 18.2 GiB/s | 9.9 GiB/s | ||
| m8azn.large | 856 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8i-flex.xlarge | 774 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8id.xlarge | 226 µs | 67,089 | 33,526 | 451 MiB/s | 216 MiB/s | ||
| m8i.xlarge | 930 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.xlarge | 301 µs | 150,053 | 82,487 | 1,010 MiB/s | 788 MiB/s | ||
| r8a.large | 881 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8a.large | 774 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| c8a.large | 856 µs | 2,994 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8id.large | 228 µs | 33,542 | 16,753 | 226 MiB/s | 108 MiB/s | ||
| m8i.large | 872 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.large | 212 µs | 75,023 | 41,229 | 503 MiB/s | 395 MiB/s | ||
| m8i-flex.large | 881 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| t8i.medium | 791 µs | 3,000 | 2,998 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud premium | 358 µs | 374,058 | 28,158 | 3.5 GiB/s | 1.4 GiB/s | ||
| standard 8 vCPU | 469 µs | 264,258 | 11,860 | 2.6 GiB/s | 944 MiB/s | ||
| premium 4 vCPU | 569 µs | 232,993 | 16,062 | 2.7 GiB/s | 882 MiB/s | ||
| standard 4 vCPU | 1,122 µs | 168,254 | 8,677 | 2.2 GiB/s | 483 MiB/s | ||
| standard 2 vCPU | 1,106 µs | 112,843 | 54,236 | 2.6 GiB/s | 809 MiB/s | ||
| premium 2 vCPU | 1,106 µs | 138,040 | 64,867 | 2.7 GiB/s | 867 MiB/s | ||
| GitHub | 202 µs | 38,954 | 18,761 | 784 MiB/s | 532 MiB/s | ||
| 4-core | 251 µs | 19,580 | 11,712 | 394 MiB/s | 394 MiB/s | ||
| 2-core | 289 µs | 9,378 | 8,712 | 199 MiB/s | 199 MiB/s | ||
| AWS CodeBuild large | 750 µs | 2,988 | 2,970 | 252 MiB/s | 252 MiB/s | ||
| medium | 758 µs | 2,993 | 2,968 | 252 MiB/s | 252 MiB/s | ||
| small | 823 µs | 2,989 | 2,968 | 252 MiB/s | 252 MiB/s | ||
| arm64 | |||||||
| Warpbuild xfast (Apple M4 Pro) | 40 µs | 339,443 | 5,906 | 22.4 GiB/s | 27.2 GiB/s | ||
| xfast (Apple M4 Pro) 12 vCPU | 49 µs | 334,760 | 5,979 | 19.8 GiB/s | 26.8 GiB/s | ||
| 8 vCPU | 782 µs | 4,993 | 4,995 | 401 MiB/s | 401 MiB/s | ||
| 4 vCPU | 807 µs | 4,195 | 4,193 | 301 MiB/s | 302 MiB/s | ||
| 2 vCPU | 832 µs | 4,195 | 4,193 | 302 MiB/s | 301 MiB/s | ||
| Blacksmith | 36 µs | 236,954 | 28,207 | 1.6 GiB/s | 944 MiB/s | ||
| 4 vCPU | 37 µs | 167,257 | 25,998 | 1.6 GiB/s | 822 MiB/s | ||
| 2 vCPU | 43 µs | 77,706 | 27,167 | 1.6 GiB/s | 898 MiB/s | ||
| Namespace Apple silicon | 85 µs | 148,271 | 17,731 | 31.2 GiB/s | 15.4 GiB/s | ||
| 8x16 | 93 µs | 188,957 | 61,821 | 965 MiB/s | 773 MiB/s | ||
| 8x32 | 98 µs | 179,145 | 40,849 | 942 MiB/s | 756 MiB/s | ||
| 4x16 | 90 µs | 144,335 | 44,153 | 978 MiB/s | 738 MiB/s | ||
| 4x8 | 89 µs | 145,399 | 29,784 | 994 MiB/s | 808 MiB/s | ||
| 2x8 | 106 µs | 62,389 | 24,244 | 916 MiB/s | 696 MiB/s | ||
| Avrea | 53 µs | 137,885 | 15,642 | 41.6 GiB/s | 15.4 GiB/s | ||
| 4 vCPU | 55 µs | 146,383 | 15,032 | 28.2 GiB/s | 13.7 GiB/s | ||
| 2 vCPU | 59 µs | 165,024 | 13,185 | 31.4 GiB/s | 15.7 GiB/s | ||
| RunsOn m9gd.2xlarge | 81 µs | 174,573 | 87,210 | 1.1 GiB/s | 560 MiB/s | ||
| m9g.2xlarge | 807 µs | 2,995 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m9g.2xlarge | <1 µs | 2,651,587 | 2,495,659 | 14.3 GiB/s | 9.1 GiB/s | ||
| c8g.2xlarge | 848 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | 938 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | <1 µs | 2,710,177 | 2,526,279 | 15.5 GiB/s | 9.3 GiB/s | ||
| m9gd.xlarge | 216 µs | 87,221 | 43,583 | 585 MiB/s | 280 MiB/s | ||
| c8g.xlarge | 897 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m9gd.large | 81 µs | 43,606 | 21,782 | 293 MiB/s | 141 MiB/s | ||
| m9g.large | 963 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| c8g.large | 815 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud standard | 334 µs | 93,715 | 55,422 | 937 MiB/s | 637 MiB/s | ||
| standard 8 vCPU | 358 µs | 222,352 | 101,505 | 875 MiB/s | 438 MiB/s | ||
| standard 4 vCPU | 449 µs | 164,716 | 26,833 | 798 MiB/s | 428 MiB/s | ||
| GitHub | 202 µs | 19,595 | 11,786 | 394 MiB/s | 337 MiB/s | ||
| 2-core | 230 µs | 9,377 | 6,115 | 199 MiB/s | 163 MiB/s | ||
- IOPS are 4k blocks, medians of finished jobs
- bold best disk measured on the arch (tmpfs is RAM, left out)
- , : cheaper, less durable flushes
Pick the disk per job on RunsOn
-
The instance type sets the storage under the workspace: EBS gp3 by default, provisioned gp3 for bandwidth, local NVMe
on the
dandifamilies, or tmpfs when the job fits in RAM. Pick it with thefamilylabel; each one, measured: EC2 storage benchmark.
How it's measured
fio runs inside the job, in GITHUB_WORKSPACE: the disk a build writes to, on the filesystem the provider
mounted there. Every runner is compared, whatever its shape.
fio settingsblock sizes, queue depths, mounts
- QD1 random read: 4k blocks, one request in flight; IOPS and p99 latency.
- Sync write: 4k writes, each followed by
fdatasync. - Random read and write: 4k blocks, 4 jobs with 32 requests in flight each.
- Sequential read and write: 1M blocks, 32 requests in flight.
- The workspace filesystem, its size and its mount options come from the job's own mount table.
- Earlier versions of this page used a different harness and other fio settings on 2 vCPU runners; those numbers are gone rather than mixed in.