Fastest GitHub Actions runners: disk I/O.
fio inside the job workspace of 89 runners from 9 providers, 2 to 12 vCPU. Latency first: a build waits on small reads and synced writes far more often than it streams one big file.
How this shows up in real builds
Latency per provider
Each provider's best disk. QD1 is one 4k read at a time, the way a build reads thousands of small files. tmpfs is RAM, not a disk: add it to see the gap on the same scale.
tmpfs for scale: 1,127,245 on RunsOn m8a.2xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
tmpfs for scale: 674,926 on RunsOn m9g.xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
What the numbers say
Extremes on x64, and RAM against disk.
- On x64, one small read at a time (4k, QD1) comes back 84,854 times a second on Blacksmith 4 vCPU and 1,559 on RunsOn m8i.2xlarge · EBS gp3, 54× apart; p99 latency is 14 µs against 1,036 µs.
- Synced 4k writes on x64 run from 470 per second on RunsOn m8i.2xlarge · EBS gp3 (workspace mounted nobarrier) to 18,160 on Namespace 8x32.
- r8a.xlarge with tmpfs ($0.0021/min at spot) answers 1,054,041 QD1 reads and 912,188 synced writes per second: 87× and 77× the best local NVMe at 4 vCPU (i7i.xlarge, $0.0019) and 577× and 1,692× default EBS gp3 (m8a.xlarge, $0.0015); its workspace is 31 GiB, from RAM. (EC2 storage suite)
More findings20: EBS, arm64, NVMe vs EBS, provisioned gp3, tmpfs, storage classes
- Every EBS-backed workspace (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations) answers between 1,559 and 1,842 QD1 reads and 470 to 562 synced writes per second, whoever runs it and whatever throughput it is provisioned for: each request is a round trip to network storage.
- On arm64, one small read at a time (4k, QD1) comes back 27,054 times a second on Warpbuild xfast (Apple M4 Pro) 6 vCPU and 1,660 on RunsOn m9g.large · EBS gp3, 16× apart; p99 latency is 39 µs against 963 µs.
- Synced 4k writes on arm64 run from 538 per second on Warpbuild 2 vCPU (workspace mounted nobarrier) to 16,257 on RunsOn m9gd.2xlarge · local NVMe.
- Same Xeon 6975P-C CPU, different disk: m8id.2xlarge on local NVMe answers 11,168 QD1 reads and 7,657 synced writes per second, m8i.2xlarge on EBS gp3 1,559 and 470.
- Same Xeon 6975P-C CPU, different disk: m8id.large on local NVMe answers 10,850 QD1 reads and 5,551 synced writes per second, m8i-flex.large on EBS gp3 1,745 and 527.
- Same Xeon 6975P-C CPU, different disk: m8id.xlarge on local NVMe answers 11,056 QD1 reads and 7,841 synced writes per second, m8i-flex.xlarge on EBS gp3 1,576 and 527.
- Same Neoverse-V3 CPU, different disk: m9gd.2xlarge on local NVMe answers 14,947 QD1 reads and 16,257 synced writes per second, m9g.2xlarge on EBS gp3 1,720 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.xlarge on local NVMe answers 12,610 QD1 reads and 11,264 synced writes per second, m9g.xlarge on EBS gp3 1,818 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.large on local NVMe answers 13,526 QD1 reads and 7,159 synced writes per second, m9g.large on EBS gp3 1,660 and 543.
- Provisioning gp3 at 1000 MiB/s on m8a.2xlarge takes sequential reads from 402 MiB/s to 1,002 MiB/s (2.5×); QD1 reads barely move (1,791 → 1,792 IOPS): it buys bandwidth, not latency.
- Provisioning gp3 at 1000 MiB/s on m8azn.xlarge takes sequential reads from 401 MiB/s to 1,002 MiB/s (2.5×); QD1 reads barely move (1,775 → 1,775 IOPS): it buys bandwidth, not latency.
- tmpfs on m8a.2xlarge is not a disk at all: 1,127,245 QD1 reads and 664,801 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.2xlarge is not a disk at all: 625,966 QD1 reads and 470,633 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.xlarge is not a disk at all: 674,926 QD1 reads and 634,741 synced writes per second, from RAM. The workspace is 15 GiB, sized by memory.
- tmpfs on r8a.2xlarge is not a disk at all: 1,098,003 QD1 reads and 877,674 synced writes per second, from RAM. The workspace is 62 GiB, sized by memory.
- tmpfs: 626k–1.1M QD1 reads and 471k–912k synced writes per second across 5 runners (5 RunsOn configurations).
- VM disk: 3.8k–85k QD1 reads and 578–18k synced writes per second across 42 runners from 7 providers (6 Avrea runners, 6 Blacksmith runners, 3 GitHub runners, 11 Namespace runners, 2 StarSling x64 runners, 9 Ubicloud runners and 5 Warpbuild runners).
- local NVMe: 11k–15k QD1 reads and 5.6k–16k synced writes per second across 9 runners (9 RunsOn configurations).
- network disk: 6.8k–8.6k QD1 reads and 3.3k–4.7k synced writes per second across 2 runners (2 GitHub runners).
- EBS: 1.6k–1.8k QD1 reads and 470–562 synced writes per second across 31 runners from 3 providers (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations).
Reading these numbers
Most of a build is small, dependent I/O, not one big stream.
Why latency before throughput
-
git checkout,npm install,cargo buildand test runners read thousands of small files, each waiting for the last: the QD1 column. QD32 and sequential show what a parallel compiler, a bigdocker buildor a cache archive can pull; read them after. - Package managers, databases and SQLite-backed tests flush their writes. Each flush waits for the storage: network disks pay a round trip every time. That is sync write.
Every runner
Each provider's best disk first; its other machines sit behind +N more. Bars are log scale.
| One request at a time | 4k, 32 in flight | Sequential, 1M | |||||
|---|---|---|---|---|---|---|---|
| Runner | QD1 read ↑ | p99 ↓ | Sync write ↑ | Read ↑ | Write ↑ | Read ↑ | Write ↑ |
| x64 | |||||||
| Blacksmith | 14 µs | 873,112 | 251,762 | 6.9 GiB/s | 1.6 GiB/s | ||
| 8 vCPU | 19 µs | 567,996 | 199,982 | 4.0 GiB/s | 1.7 GiB/s | ||
| 2 vCPU | 19 µs | 515,353 | 136,106 | 4.7 GiB/s | 1.9 GiB/s | ||
| Namespace 4x16 | 12 µs | 1,011,926 | 436,056 | 4.1 GiB/s | 2.9 GiB/s | ||
| 8x32 | 13 µs | 1,126,282 | 417,764 | 4.2 GiB/s | 2.9 GiB/s | ||
| 8x16 | 12 µs | 1,137,038 | 423,854 | 4.1 GiB/s | 2.9 GiB/s | ||
| 4x8 | 12 µs | 1,009,155 | 403,484 | 4.1 GiB/s | 3.0 GiB/s | ||
| 2x8 | 13 µs | 350,533 | 222,621 | 4.1 GiB/s | 2.9 GiB/s | ||
| StarSling | 17 µs | 262,205 | 11,527 | 2.5 GiB/s | 3.4 GiB/s | ||
| 8 vCPU | 18 µs | 318,790 | 10,959 | 2.5 GiB/s | 2.0 GiB/s | ||
| Avrea | 18 µs | 402,481 | 172,765 | 4.6 GiB/s | 3.1 GiB/s | ||
| 8 vCPU | 23 µs | 245,805 | 139,439 | 7.4 GiB/s | 5.4 GiB/s | ||
| 2 vCPU | 20 µs | 454,066 | 132,755 | 5.3 GiB/s | 3.4 GiB/s | ||
| Warpbuild | 56 µs | 151,759 | 72,474 | 7.6 GiB/s | 4.6 GiB/s | ||
| 8 vCPU | 60 µs | 98,888 | 44,952 | 7.9 GiB/s | 4.5 GiB/s | ||
| 4 vCPU | 58 µs | 143,694 | 40,569 | 7.8 GiB/s | 5.1 GiB/s | ||
| RunsOn i7i.2xlarge | 228 µs | 300,466 | 165,109 | 2.0 GiB/s | 1.5 GiB/s | ||
| m8azn.3xlarge | 766 µs | 2,997 | 2,996 | 401 MiB/s | 401 MiB/s | ||
| c8a.2xlarge | 832 µs | 2,995 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8a.2xlarge | 856 µs | 2,995 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8a.2xlarge | 897 µs | 15,995 | 15,993 | 1,001 MiB/s | 1,001 MiB/s | ||
| r8a.2xlarge | <1 µs | 4,207,889 | 3,669,702 | 18.5 GiB/s | 10.1 GiB/s | ||
| m8a.2xlarge | <1 µs | 4,214,475 | 3,756,731 | 18.4 GiB/s | 10.0 GiB/s | ||
| m8azn.xlarge | 963 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8azn.xlarge | 856 µs | 15,994 | 15,991 | 1,002 MiB/s | 1,001 MiB/s | ||
| m8id.2xlarge | 169 µs | 134,258 | 67,074 | 905 MiB/s | 432 MiB/s | ||
| m8i.2xlarge | 1,036 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m8a.xlarge | 840 µs | 2,993 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8i-flex.2xlarge | 848 µs | 2,992 | 2,995 | 401 MiB/s | 401 MiB/s | ||
| c8a.xlarge | 930 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| r8a.xlarge | <1 µs | 4,014,510 | 3,665,859 | 18.2 GiB/s | 9.9 GiB/s | ||
| m8azn.large | 856 µs | 2,992 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8i-flex.xlarge | 963 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8id.xlarge | 210 µs | 67,089 | 33,526 | 451 MiB/s | 216 MiB/s | ||
| m8i.xlarge | 930 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.xlarge | 307 µs | 150,053 | 82,488 | 1,010 MiB/s | 788 MiB/s | ||
| r8a.large | 881 µs | 2,991 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m8a.large | 770 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| c8a.large | 807 µs | 2,991 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8id.large | 229 µs | 33,542 | 16,753 | 225 MiB/s | 108 MiB/s | ||
| m8i.large | 872 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.large | 212 µs | 75,023 | 41,229 | 503 MiB/s | 395 MiB/s | ||
| m8i-flex.large | 881 µs | 2,994 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| t8i.medium | 799 µs | 3,000 | 2,998 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud premium | 408 µs | 362,984 | 19,859 | 3.1 GiB/s | 1.2 GiB/s | ||
| standard 8 vCPU | 569 µs | 264,258 | 11,860 | 2.6 GiB/s | 689 MiB/s | ||
| premium 4 vCPU | 1,679 µs | 210,832 | 16,062 | 2.7 GiB/s | 807 MiB/s | ||
| standard 4 vCPU | 2,073 µs | 157,238 | 8,344 | 2.2 GiB/s | 250 MiB/s | ||
| standard 2 vCPU | 864 µs | 106,887 | 46,311 | 2.3 GiB/s | 398 MiB/s | ||
| premium 2 vCPU | 1,106 µs | 138,040 | 64,617 | 2.8 GiB/s | 833 MiB/s | ||
| GitHub | 192 µs | 38,954 | 18,761 | 784 MiB/s | 532 MiB/s | ||
| 4-core | 251 µs | 19,587 | 11,712 | 394 MiB/s | 394 MiB/s | ||
| 2-core | 289 µs | 9,378 | 8,712 | 199 MiB/s | 199 MiB/s | ||
| AWS CodeBuild large | 799 µs | 2,988 | 2,970 | 252 MiB/s | 252 MiB/s | ||
| medium | 778 µs | 2,990 | 2,967 | 252 MiB/s | 252 MiB/s | ||
| small | 791 µs | 2,988 | 2,969 | 252 MiB/s | 252 MiB/s | ||
| arm64 | |||||||
| Warpbuild xfast (Apple M4 Pro) | 39 µs | 338,802 | 5,322 | 22.4 GiB/s | 27.2 GiB/s | ||
| xfast (Apple M4 Pro) 12 vCPU | 52 µs | 334,760 | 5,979 | 19.8 GiB/s | 26.8 GiB/s | ||
| 8 vCPU | 807 µs | 4,993 | 4,995 | 401 MiB/s | 401 MiB/s | ||
| 4 vCPU | 807 µs | 4,195 | 4,193 | 301 MiB/s | 302 MiB/s | ||
| 2 vCPU | 832 µs | 4,194 | 4,193 | 302 MiB/s | 301 MiB/s | ||
| Blacksmith | 38 µs | 223,621 | 27,644 | 1.6 GiB/s | 940 MiB/s | ||
| 4 vCPU | 39 µs | 155,975 | 24,132 | 1.7 GiB/s | 742 MiB/s | ||
| 2 vCPU | 46 µs | 77,706 | 23,744 | 1.5 GiB/s | 775 MiB/s | ||
| Namespace Apple silicon | 80 µs | 143,397 | 15,978 | 32.9 GiB/s | 15.4 GiB/s | ||
| 8x16 | 93 µs | 188,888 | 61,821 | 990 MiB/s | 748 MiB/s | ||
| 8x32 | 98 µs | 186,558 | 32,721 | 953 MiB/s | 763 MiB/s | ||
| 4x16 | 90 µs | 144,335 | 39,744 | 969 MiB/s | 738 MiB/s | ||
| 4x8 | 90 µs | 145,399 | 29,784 | 988 MiB/s | 756 MiB/s | ||
| 2x8 | 106 µs | 77,649 | 24,244 | 921 MiB/s | 736 MiB/s | ||
| Avrea | 56 µs | 137,885 | 15,642 | 36.5 GiB/s | 15.6 GiB/s | ||
| 4 vCPU | 55 µs | 146,579 | 15,032 | 28.9 GiB/s | 13.7 GiB/s | ||
| 2 vCPU | 59 µs | 165,024 | 14,173 | 31.7 GiB/s | 14.0 GiB/s | ||
| RunsOn m9gd.2xlarge | 72 µs | 174,573 | 87,210 | 1.1 GiB/s | 560 MiB/s | ||
| m9g.2xlarge | 823 µs | 2,995 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m9g.2xlarge | <1 µs | 2,537,608 | 2,410,790 | 11.6 GiB/s | 4.0 GiB/s | ||
| c8g.2xlarge | 815 µs | 2,994 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | 938 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | <1 µs | 2,690,916 | 2,487,599 | 15.1 GiB/s | 4.6 GiB/s | ||
| m9gd.xlarge | 138 µs | 87,221 | 43,583 | 585 MiB/s | 280 MiB/s | ||
| c8g.xlarge | 889 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m9gd.large | 81 µs | 43,607 | 21,782 | 293 MiB/s | 141 MiB/s | ||
| m9g.large | 963 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| c8g.large | 815 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud standard | 330 µs | 211,987 | 101,505 | 832 MiB/s | 438 MiB/s | ||
| standard 4 vCPU | 610 µs | 158,274 | 57,360 | 836 MiB/s | 351 MiB/s | ||
| standard 2 vCPU | 317 µs | 107,365 | 78,893 | 929 MiB/s | 637 MiB/s | ||
| GitHub | 202 µs | 19,588 | 11,775 | 394 MiB/s | 337 MiB/s | ||
| 2-core | 264 µs | 9,377 | 6,407 | 200 MiB/s | 163 MiB/s | ||
- IOPS are 4k blocks, medians of finished jobs
- bold best disk measured on the arch (tmpfs is RAM, left out)
- , : cheaper, less durable flushes
Pick the disk per job on RunsOn
-
The instance type sets the storage under the workspace: EBS gp3 by default, provisioned gp3 for bandwidth, local NVMe
on the
dandifamilies, or tmpfs when the job fits in RAM. Pick it with thefamilylabel; each one, measured: EC2 storage benchmark.
How it's measured
fio runs inside the job, in GITHUB_WORKSPACE: the disk a build writes to, on the filesystem the provider
mounted there. Every runner is compared, whatever its shape.
fio settingsblock sizes, queue depths, mounts
- QD1 random read: 4k blocks, one request in flight; IOPS and p99 latency.
- Sync write: 4k writes, each followed by
fdatasync. - Random read and write: 4k blocks, 4 jobs with 32 requests in flight each.
- Sequential read and write: 1M blocks, 32 requests in flight.
- The workspace filesystem, its size and its mount options come from the job's own mount table.
- Earlier versions of this page used a different harness and other fio settings on 2 vCPU runners; those numbers are gone rather than mixed in.