Fastest GitHub Actions runners: disk I/O.
fio inside the job workspace of 89 runners from 9 providers, 2 to 12 vCPU. Latency first: a build waits on small reads and synced writes far more often than it streams one big file.
How this shows up in real builds
Latency per provider
Each provider's best disk. QD1 is one 4k read at a time, the way a build reads thousands of small files. tmpfs is RAM, not a disk: add it to see the gap on the same scale.
tmpfs for scale: 1,156,332 on RunsOn m8a.2xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
tmpfs for scale: 682,673 on RunsOn m9g.xlarge · tmpfs.
volatile and nobarrier mounts make each flush cheaper by making it less durable.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.
What the numbers say
Extremes on x64, and RAM against disk.
- On x64, one small read at a time (4k, QD1) comes back 84,854 times a second on Blacksmith 4 vCPU and 1,569 on RunsOn m8i.2xlarge · EBS gp3, 54× apart; p99 latency is 14 µs against 1,012 µs.
- Synced 4k writes on x64 run from 498 per second on RunsOn m8i.2xlarge · EBS gp3 (workspace mounted nobarrier) to 20,425 on Blacksmith 4 vCPU (workspace mounted nobarrier).
- r8a.xlarge with tmpfs ($0.0021/min at spot) answers 1,071,527 QD1 reads and 431,794 synced writes per second: 88× and 36× the best local NVMe at 4 vCPU (i7i.xlarge, $0.0019) and 584× and 800× default EBS gp3 (m8a.xlarge, $0.0015); its workspace is 31 GiB, from RAM. (EC2 storage suite)
More findings20: EBS, arm64, NVMe vs EBS, provisioned gp3, tmpfs, storage classes
- Every EBS-backed workspace (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations) answers between 1,568 and 1,844 QD1 reads and 498 to 562 synced writes per second, whoever runs it and whatever throughput it is provisioned for: each request is a round trip to network storage.
- On arm64, one small read at a time (4k, QD1) comes back 26,844 times a second on Warpbuild xfast (Apple M4 Pro) 6 vCPU and 1,568 on RunsOn m9g.xlarge · EBS gp3, 17× apart; p99 latency is 41 µs against 946 µs.
- Synced 4k writes on arm64 run from 537 per second on RunsOn c8g.large · EBS gp3 (workspace mounted nobarrier) to 17,318 on RunsOn m9gd.2xlarge · local NVMe.
- Same Xeon 6975P-C CPU, different disk: m8id.2xlarge on local NVMe answers 10,957 QD1 reads and 7,657 synced writes per second, m8i.2xlarge on EBS gp3 1,569 and 498.
- Same Xeon 6975P-C CPU, different disk: m8id.large on local NVMe answers 10,991 QD1 reads and 5,549 synced writes per second, m8i-flex.large on EBS gp3 1,745 and 530.
- Same Xeon 6975P-C CPU, different disk: m8id.xlarge on local NVMe answers 11,045 QD1 reads and 7,898 synced writes per second, m8i-flex.xlarge on EBS gp3 1,765 and 527.
- Same Neoverse-V3 CPU, different disk: m9gd.2xlarge on local NVMe answers 13,464 QD1 reads and 17,318 synced writes per second, m9g.2xlarge on EBS gp3 1,788 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.xlarge on local NVMe answers 12,704 QD1 reads and 12,664 synced writes per second, m9g.xlarge on EBS gp3 1,568 and 543.
- Same Neoverse-V3 CPU, different disk: m9gd.large on local NVMe answers 13,492 QD1 reads and 7,267 synced writes per second, m9g.large on EBS gp3 1,621 and 542.
- Provisioning gp3 at 1000 MiB/s on m8a.2xlarge takes sequential reads from 402 MiB/s to 1,002 MiB/s (2.5×); QD1 reads barely move (1,791 → 1,792 IOPS): it buys bandwidth, not latency.
- Provisioning gp3 at 1000 MiB/s on m8azn.xlarge takes sequential reads from 401 MiB/s to 1,001 MiB/s (2.5×); QD1 reads barely move (1,773 → 1,764 IOPS): it buys bandwidth, not latency.
- tmpfs on m8a.2xlarge is not a disk at all: 1,156,332 QD1 reads and 432,491 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.2xlarge is not a disk at all: 653,143 QD1 reads and 301,438 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
- tmpfs on m9g.xlarge is not a disk at all: 682,673 QD1 reads and 303,857 synced writes per second, from RAM. The workspace is 15 GiB, sized by memory.
- tmpfs on r8a.2xlarge is not a disk at all: 1,154,342 QD1 reads and 433,632 synced writes per second, from RAM. The workspace is 62 GiB, sized by memory.
- tmpfs: 653k–1.2M QD1 reads and 301k–434k synced writes per second across 5 runners (5 RunsOn configurations).
- VM disk: 6.0k–85k QD1 reads and 578–20k synced writes per second across 42 runners from 7 providers (6 Avrea runners, 6 Blacksmith runners, 3 GitHub runners, 11 Namespace runners, 2 StarSling x64 runners, 9 Ubicloud runners and 5 Warpbuild runners).
- local NVMe: 11k–13k QD1 reads and 5.5k–17k synced writes per second across 9 runners (9 RunsOn configurations).
- network disk: 6.7k–8.6k QD1 reads and 3.1k–4.4k synced writes per second across 2 runners (2 GitHub runners).
- EBS: 1.6k–1.8k QD1 reads and 498–562 synced writes per second across 31 runners from 3 providers (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations).
Reading these numbers
Most of a build is small, dependent I/O, not one big stream.
Why latency before throughput
-
git checkout,npm install,cargo buildand test runners read thousands of small files, each waiting for the last: the QD1 column. QD32 and sequential show what a parallel compiler, a bigdocker buildor a cache archive can pull; read them after. - Package managers, databases and SQLite-backed tests flush their writes. Each flush waits for the storage: network disks pay a round trip every time. That is sync write.
Every runner
Each provider's best disk first; its other machines sit behind +N more. Bars are log scale.
| One request at a time | 4k, 32 in flight | Sequential, 1M | |||||
|---|---|---|---|---|---|---|---|
| Runner | QD1 read ↑ | p99 ↓ | Sync write ↑ | Read ↑ | Write ↑ | Read ↑ | Write ↑ |
| x64 | |||||||
| Blacksmith | 14 µs | 856,257 | 254,574 | 6.5 GiB/s | 1.9 GiB/s | ||
| 8 vCPU | 13 µs | 683,187 | 326,353 | 6.5 GiB/s | 1.9 GiB/s | ||
| 2 vCPU | 18 µs | 637,720 | 173,541 | 6.3 GiB/s | 2.1 GiB/s | ||
| Namespace 8x32 | 13 µs | 1,108,356 | 412,276 | 4.2 GiB/s | 3.0 GiB/s | ||
| 8x16 | 12 µs | 1,120,671 | 476,969 | 4.1 GiB/s | 2.9 GiB/s | ||
| 4x16 | 12 µs | 1,007,328 | 406,763 | 4.1 GiB/s | 2.9 GiB/s | ||
| 4x8 | 13 µs | 1,015,510 | 403,484 | 4.1 GiB/s | 2.9 GiB/s | ||
| 2x8 | 13 µs | 351,146 | 222,814 | 4.1 GiB/s | 2.9 GiB/s | ||
| StarSling | 17 µs | 318,790 | 16,225 | 2.3 GiB/s | 2.0 GiB/s | ||
| 4 vCPU | 17 µs | 262,205 | 16,650 | 3.9 GiB/s | 2.8 GiB/s | ||
| Avrea | 18 µs | 468,080 | 246,705 | 6.6 GiB/s | 3.5 GiB/s | ||
| 8 vCPU | 19 µs | 235,683 | 139,439 | 7.4 GiB/s | 3.3 GiB/s | ||
| 2 vCPU | 17 µs | 454,466 | 120,725 | 8.4 GiB/s | 6.2 GiB/s | ||
| Warpbuild | 59 µs | 176,232 | 68,626 | 7.5 GiB/s | 4.5 GiB/s | ||
| 8 vCPU | 60 µs | 120,520 | 45,671 | 7.8 GiB/s | 4.9 GiB/s | ||
| 4 vCPU | 64 µs | 124,074 | 38,883 | 7.2 GiB/s | 5.1 GiB/s | ||
| RunsOn i7i.2xlarge | 228 µs | 300,466 | 165,110 | 2.0 GiB/s | 1.5 GiB/s | ||
| m8azn.3xlarge | 782 µs | 2,997 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| c8a.2xlarge | 840 µs | 2,995 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| r8a.2xlarge | <1 µs | 4,231,165 | 3,828,376 | 18.6 GiB/s | 10.1 GiB/s | ||
| m8a.2xlarge | 897 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8a.2xlarge | 840 µs | 15,996 | 15,990 | 1,001 MiB/s | 1,001 MiB/s | ||
| m8a.2xlarge | <1 µs | 4,258,470 | 3,812,490 | 18.1 GiB/s | 10.0 GiB/s | ||
| m8azn.xlarge | 897 µs | 2,994 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8azn.xlarge | 872 µs | 15,994 | 15,991 | 1,001 MiB/s | 1,001 MiB/s | ||
| m8i.2xlarge | 1,012 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m8id.2xlarge | 228 µs | 134,258 | 67,074 | 905 MiB/s | 432 MiB/s | ||
| m8i-flex.2xlarge | 872 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m8a.xlarge | 799 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| c8a.xlarge | 832 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| r8a.xlarge | <1 µs | 4,113,746 | 3,786,762 | 18.2 GiB/s | 9.8 GiB/s | ||
| m8azn.large | 881 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m8i-flex.xlarge | 774 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| m8id.xlarge | 218 µs | 67,089 | 33,526 | 451 MiB/s | 216 MiB/s | ||
| m8i.xlarge | 930 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.xlarge | 301 µs | 150,053 | 82,487 | 1,010 MiB/s | 788 MiB/s | ||
| m8a.large | 774 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| r8a.large | 807 µs | 2,994 | 2,993 | 401 MiB/s | 401 MiB/s | ||
| c8a.large | 807 µs | 2,994 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m8id.large | 228 µs | 33,542 | 16,753 | 226 MiB/s | 108 MiB/s | ||
| m8i.large | 872 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| i7i.large | 220 µs | 75,023 | 41,229 | 503 MiB/s | 395 MiB/s | ||
| m8i-flex.large | 881 µs | 2,994 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| t8i.medium | 791 µs | 3,000 | 2,998 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud premium | 305 µs | 374,058 | 28,158 | 3.5 GiB/s | 1.4 GiB/s | ||
| standard 8 vCPU | 528 µs | 251,230 | 11,860 | 2.4 GiB/s | 944 MiB/s | ||
| standard 4 vCPU | 700 µs | 183,661 | 10,621 | 2.4 GiB/s | 494 MiB/s | ||
| premium 4 vCPU | 1,679 µs | 173,146 | 16,062 | 2.5 GiB/s | 537 MiB/s | ||
| standard 2 vCPU | 807 µs | 112,843 | 60,676 | 2.7 GiB/s | 806 MiB/s | ||
| premium 2 vCPU | 1,106 µs | 147,821 | 64,867 | 2.7 GiB/s | 867 MiB/s | ||
| GitHub | 232 µs | 38,940 | 18,331 | 784 MiB/s | 538 MiB/s | ||
| 4-core | 251 µs | 19,578 | 11,682 | 394 MiB/s | 394 MiB/s | ||
| 2-core | 289 µs | 9,374 | 8,866 | 199 MiB/s | 130 MiB/s | ||
| AWS CodeBuild medium | 758 µs | 2,989 | 2,968 | 252 MiB/s | 252 MiB/s | ||
| large | 815 µs | 2,992 | 2,971 | 252 MiB/s | 252 MiB/s | ||
| small | 823 µs | 2,988 | 2,963 | 252 MiB/s | 252 MiB/s | ||
| arm64 | |||||||
| Warpbuild xfast (Apple M4 Pro) | 41 µs | 340,451 | 5,906 | 22.4 GiB/s | 27.2 GiB/s | ||
| xfast (Apple M4 Pro) 12 vCPU | 52 µs | 326,216 | 4,587 | 20.6 GiB/s | 25.9 GiB/s | ||
| 8 vCPU | 774 µs | 4,993 | 4,995 | 401 MiB/s | 401 MiB/s | ||
| 4 vCPU | 823 µs | 4,194 | 4,192 | 302 MiB/s | 302 MiB/s | ||
| 2 vCPU | 832 µs | 4,195 | 4,193 | 302 MiB/s | 301 MiB/s | ||
| Blacksmith | 34 µs | 249,115 | 28,207 | 1.6 GiB/s | 944 MiB/s | ||
| 4 vCPU | 37 µs | 167,257 | 25,998 | 1.6 GiB/s | 842 MiB/s | ||
| 2 vCPU | 46 µs | 74,915 | 23,744 | 1.6 GiB/s | 813 MiB/s | ||
| Avrea | 53 µs | 138,687 | 16,621 | 41.6 GiB/s | 15.6 GiB/s | ||
| 4 vCPU | 59 µs | 146,383 | 13,857 | 28.9 GiB/s | 13.9 GiB/s | ||
| 2 vCPU | 58 µs | 166,284 | 14,173 | 31.4 GiB/s | 15.7 GiB/s | ||
| Namespace Apple silicon | 85 µs | 148,271 | 17,731 | 30.9 GiB/s | 15.4 GiB/s | ||
| 8x32 | 98 µs | 179,145 | 40,849 | 942 MiB/s | 763 MiB/s | ||
| 8x16 | 94 µs | 188,957 | 52,820 | 923 MiB/s | 748 MiB/s | ||
| 4x8 | 89 µs | 145,399 | 32,144 | 994 MiB/s | 739 MiB/s | ||
| 4x16 | 90 µs | 141,434 | 36,088 | 978 MiB/s | 738 MiB/s | ||
| 2x8 | 106 µs | 62,389 | 24,244 | 941 MiB/s | 696 MiB/s | ||
| RunsOn m9gd.large | 81 µs | 43,606 | 21,782 | 293 MiB/s | 141 MiB/s | ||
| m9gd.2xlarge | 84 µs | 174,573 | 87,210 | 1.1 GiB/s | 560 MiB/s | ||
| m9g.2xlarge | 807 µs | 2,995 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m9g.2xlarge | <1 µs | 2,651,587 | 2,495,659 | 14.3 GiB/s | 9.1 GiB/s | ||
| c8g.2xlarge | 1,020 µs | 2,994 | 2,991 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | 946 µs | 2,994 | 2,992 | 401 MiB/s | 401 MiB/s | ||
| m9g.xlarge | <1 µs | 2,723,038 | 2,543,221 | 15.5 GiB/s | 9.5 GiB/s | ||
| m9gd.xlarge | 134 µs | 87,221 | 43,583 | 585 MiB/s | 280 MiB/s | ||
| c8g.xlarge | 889 µs | 2,994 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| m9g.large | 987 µs | 2,992 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| c8g.large | 889 µs | 2,993 | 2,994 | 401 MiB/s | 401 MiB/s | ||
| Ubicloud standard | 334 µs | 93,227 | 55,422 | 937 MiB/s | 608 MiB/s | ||
| standard 8 vCPU | 358 µs | 222,190 | 101,505 | 886 MiB/s | 434 MiB/s | ||
| standard 4 vCPU | 371 µs | 170,177 | 26,833 | 798 MiB/s | 392 MiB/s | ||
| GitHub | 202 µs | 19,595 | 12,039 | 394 MiB/s | 336 MiB/s | ||
| 2-core | 230 µs | 9,374 | 6,017 | 199 MiB/s | 163 MiB/s | ||
- IOPS are 4k blocks, medians of finished jobs
- bold best disk measured on the arch (tmpfs is RAM, left out)
- , : cheaper, less durable flushes
Pick the disk per job on RunsOn
-
The instance type sets the storage under the workspace: EBS gp3 by default, provisioned gp3 for bandwidth, local NVMe
on the
dandifamilies, or tmpfs when the job fits in RAM. Pick it with thefamilylabel; each one, measured: EC2 storage benchmark.
How it's measured
fio runs inside the job, in GITHUB_WORKSPACE: the disk a build writes to, on the filesystem the provider
mounted there. Every runner is compared, whatever its shape.
fio settingsblock sizes, queue depths, mounts
- QD1 random read: 4k blocks, one request in flight; IOPS and p99 latency.
- Sync write: 4k writes, each followed by
fdatasync. - Random read and write: 4k blocks, 4 jobs with 32 requests in flight each.
- Sequential read and write: 1M blocks, 32 requests in flight.
- The workspace filesystem, its size and its mount options come from the job's own mount table.
- Earlier versions of this page used a different harness and other fio settings on 2 vCPU runners; those numbers are gone rather than mixed in.