self-host →

Fastest GitHub Actions runners: disk I/O.

fio inside the job workspace of 89 runners from 9 providers, 2 to 12 vCPU. Latency first: a build waits on small reads and synced writes far more often than it streams one big file.

Fastest QD1 read, x64 85,326 Blacksmith 4 vCPU · reads/s
Fastest QD1 read, arm64 26,866 Warpbuild xfast (Apple M4 Pro) 6 vCPU · reads/s
GitHub's own runner 7,964 x64 QD1 reads/s on GitHub 8-core, the baseline
Fastest synced writes, x64 20,425 Blacksmith 4 vCPU · writes/s · nobarrier mount
tmpfs, for scale 1.2M QD1 reads/s on RunsOn m8a.2xlarge · tmpfs: RAM, not a disk, never ranked

Latency per provider

Each provider's best disk. QD1 is one 4k read at a time, the way a build reads thousands of small files. tmpfs is RAM, not a disk: add it to see the gap on the same scale.

Arch
Storage
QD1 random readx64 · 4k, one request in flight · IOPS, higher is better
  1. Blacksmith4 vCPUVM disk85,326, p99 14 µs, ×1.00
  2. Namespace8 vCPU8x32 · VM disk78,677, p99 13 µs, ×0.92
  3. StarSling8 vCPUVM disk54,269, p99 17 µs, ×0.64
  4. Avrea4 vCPUVM disk47,519, p99 18 µs, ×0.56
  5. Warpbuild2 vCPUVM disk20,471, p99 58 µs, ×0.24
  6. RunsOn8 vCPUi7i.2xlarge · local NVMe12,709, p99 228 µs, ×0.15
  7. Ubicloud8 vCPUpremium · VM disk10,127, p99 358 µs, ×0.12
  8. GitHub8 vCPUVM disk7,964, p99 202 µs, ×0.09
  9. AWS CodeBuild8 vCPUlarge · EBS1,752, p99 750 µs, ×0.02

tmpfs for scale: 1,156,332 on RunsOn m8a.2xlarge · tmpfs.

Synced writesx64 · 4k + fdatasync · IOPS, higher is better
  1. Blacksmith4 vCPUVM disk · nobarrier20,425, p99 22 µs, ×1.00
  2. Namespace8 vCPU8x32 · VM disk18,166, p99 20 µs, ×0.89
  3. RunsOn2 vCPUi7i.large · local NVMe13,567, p99 31 µs, ×0.66
  4. Avrea4 vCPUVM disk8,386, p99 42 µs, ×0.41
  5. Ubicloud8 vCPUpremium · VM disk4,203, p99 338 µs, ×0.21
  6. GitHub2 vCPUVM disk · nobarrier4,000, p99 171 µs, ×0.20
  7. StarSling8 vCPUVM disk3,070, p99 28 µs, ×0.15
  8. Warpbuild2 vCPUVM disk2,056, p99 111 µs, ×0.10
  9. AWS CodeBuild2 vCPUsmall · EBS562, p99 995 µs, ×0.03

volatile and nobarrier mounts make each flush cheaper by making it less durable.

QD1 random readx64 · 4k, one request in flight · IOPS, higher is better
  1. RunsOn8 vCPUm8a.2xlarge · tmpfs · RAM1,156,332, p99 <1 µs, ×14
  2. RunsOn8 vCPUr8a.2xlarge · tmpfs · RAM1,151,153, p99 <1 µs, ×13
  3. RunsOn4 vCPUr8a.xlarge · tmpfs · RAM1,071,527, p99 <1 µs, ×13
  4. Blacksmith4 vCPUVM disk85,326, p99 14 µs, ×1.00
  5. Namespace8 vCPU8x32 · VM disk78,677, p99 13 µs, ×0.92
  6. StarSling8 vCPUVM disk54,269, p99 17 µs, ×0.64
  7. Avrea4 vCPUVM disk47,519, p99 18 µs, ×0.56
  8. Warpbuild2 vCPUVM disk20,471, p99 58 µs, ×0.24
  9. RunsOn8 vCPUi7i.2xlarge · local NVMe12,709, p99 228 µs, ×0.15
  10. Ubicloud8 vCPUpremium · VM disk10,127, p99 358 µs, ×0.12
  11. GitHub8 vCPUVM disk7,964, p99 202 µs, ×0.09
  12. AWS CodeBuild8 vCPUlarge · EBS1,752, p99 750 µs, ×0.02

Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.

Synced writesx64 · 4k + fdatasync · IOPS, higher is better
  1. RunsOn8 vCPUr8a.2xlarge · tmpfs · RAM433,820, p99 <1 µs, ×21
  2. RunsOn8 vCPUm8a.2xlarge · tmpfs · RAM432,842, p99 <1 µs, ×21
  3. RunsOn4 vCPUr8a.xlarge · tmpfs · RAM431,794, p99 <1 µs, ×21
  4. Blacksmith4 vCPUVM disk · nobarrier20,425, p99 22 µs, ×1.00
  5. Namespace8 vCPU8x32 · VM disk18,166, p99 20 µs, ×0.89
  6. RunsOn2 vCPUi7i.large · local NVMe13,567, p99 31 µs, ×0.66
  7. Avrea4 vCPUVM disk8,386, p99 42 µs, ×0.41
  8. Ubicloud8 vCPUpremium · VM disk4,203, p99 338 µs, ×0.21
  9. GitHub2 vCPUVM disk · nobarrier4,000, p99 171 µs, ×0.20
  10. StarSling8 vCPUVM disk3,070, p99 28 µs, ×0.15
  11. Warpbuild2 vCPUVM disk2,056, p99 111 µs, ×0.10
  12. AWS CodeBuild2 vCPUsmall · EBS562, p99 995 µs, ×0.03

Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.

QD1 random readarm64 · 4k, one request in flight · IOPS, higher is better
  1. Warpbuild6 vCPUxfast (Apple M4 Pro) · VM disk26,866, p99 40 µs, ×1.00
  2. Blacksmith8 vCPUVM disk26,024, p99 36 µs, ×0.97
  3. Namespace4 vCPUApple silicon · VM disk18,455, p99 85 µs, ×0.69
  4. Avrea8 vCPUVM disk18,297, p99 53 µs, ×0.68
  5. RunsOn8 vCPUm9gd.2xlarge · local NVMe13,522, p99 81 µs, ×0.50
  6. Ubicloud2 vCPUstandard · VM disk9,482, p99 334 µs, ×0.35
  7. GitHub4 vCPUnetwork disk8,635, p99 202 µs, ×0.32

tmpfs for scale: 680,179 on RunsOn m9g.xlarge · tmpfs.

Synced writesarm64 · 4k + fdatasync · IOPS, higher is better
  1. RunsOn8 vCPUm9gd.2xlarge · local NVMe17,054, p99 20 µs, ×1.00
  2. Warpbuild12 vCPUxfast (Apple M4 Pro) · VM disk · nobarrier8,636, p99 66 µs, ×0.51
  3. Blacksmith8 vCPUVM disk · nobarrier7,676, p99 60 µs, ×0.45
  4. Ubicloud2 vCPUstandard · VM disk5,425, p99 241 µs, ×0.32
  5. GitHub4 vCPUnetwork disk · nobarrier4,666, p99 257 µs, ×0.27
  6. Avrea2 vCPUVM disk4,166, p99 56 µs, ×0.24
  7. Namespace8 vCPU8x16 · VM disk4,028, p99 163 µs, ×0.24

volatile and nobarrier mounts make each flush cheaper by making it less durable.

QD1 random readarm64 · 4k, one request in flight · IOPS, higher is better
  1. RunsOn4 vCPUm9g.xlarge · tmpfs · RAM680,179, p99 <1 µs, ×25
  2. RunsOn8 vCPUm9g.2xlarge · tmpfs · RAM653,143, p99 <1 µs, ×24
  3. Warpbuild6 vCPUxfast (Apple M4 Pro) · VM disk26,866, p99 40 µs, ×1.00
  4. Blacksmith8 vCPUVM disk26,024, p99 36 µs, ×0.97
  5. Namespace4 vCPUApple silicon · VM disk18,455, p99 85 µs, ×0.69
  6. Avrea8 vCPUVM disk18,297, p99 53 µs, ×0.68
  7. RunsOn8 vCPUm9gd.2xlarge · local NVMe13,522, p99 81 µs, ×0.50
  8. Ubicloud2 vCPUstandard · VM disk9,482, p99 334 µs, ×0.35
  9. GitHub4 vCPUnetwork disk8,635, p99 202 µs, ×0.32

Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.

Synced writesarm64 · 4k + fdatasync · IOPS, higher is better
  1. RunsOn4 vCPUm9g.xlarge · tmpfs · RAM303,962, p99 <1 µs, ×18
  2. RunsOn8 vCPUm9g.2xlarge · tmpfs · RAM301,463, p99 <1 µs, ×18
  3. RunsOn8 vCPUm9gd.2xlarge · local NVMe17,054, p99 20 µs, ×1.00
  4. Warpbuild12 vCPUxfast (Apple M4 Pro) · VM disk · nobarrier8,636, p99 66 µs, ×0.51
  5. Blacksmith8 vCPUVM disk · nobarrier7,676, p99 60 µs, ×0.45
  6. Ubicloud2 vCPUstandard · VM disk5,425, p99 241 µs, ×0.32
  7. GitHub4 vCPUnetwork disk · nobarrier4,666, p99 257 µs, ×0.27
  8. Avrea2 vCPUVM disk4,166, p99 56 µs, ×0.24
  9. Namespace8 vCPU8x16 · VM disk4,028, p99 163 µs, ×0.24

Hatched: tmpfs, the workspace in RAM. Multipliers are against the best disk.

What the numbers say

Extremes on x64, and RAM against disk.

  • On x64, one small read at a time (4k, QD1) comes back 85,326 times a second on Blacksmith 4 vCPU and 1,569 on RunsOn m8i.2xlarge · EBS gp3, 54× apart; p99 latency is 14 µs against 1,012 µs.
  • Synced 4k writes on x64 run from 495 per second on RunsOn m8i.2xlarge · EBS gp3 (workspace mounted nobarrier) to 20,425 on Blacksmith 4 vCPU (workspace mounted nobarrier).
  • r8a.xlarge with tmpfs ($0.0021/min at spot) answers 1,071,527 QD1 reads and 431,794 synced writes per second: 88× and 36× the best local NVMe at 4 vCPU (i7i.xlarge, $0.0019) and 586× and 800× default EBS gp3 (m8a.xlarge, $0.0015); its workspace is 31 GiB, from RAM. (EC2 storage suite)
More findings20: EBS, arm64, NVMe vs EBS, provisioned gp3, tmpfs, storage classes
  • Every EBS-backed workspace (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations) answers between 1,569 and 1,844 QD1 reads and 495 to 562 synced writes per second, whoever runs it and whatever throughput it is provisioned for: each request is a round trip to network storage.
  • On arm64, one small read at a time (4k, QD1) comes back 26,866 times a second on Warpbuild xfast (Apple M4 Pro) 6 vCPU and 1,660 on RunsOn m9g.large · EBS gp3, 16× apart; p99 latency is 40 µs against 963 µs.
  • Synced 4k writes on arm64 run from 537 per second on RunsOn c8g.large · EBS gp3 (workspace mounted nobarrier) to 17,054 on RunsOn m9gd.2xlarge · local NVMe.
  • Same Xeon 6975P-C CPU, different disk: m8id.2xlarge on local NVMe answers 11,063 QD1 reads and 7,665 synced writes per second, m8i.2xlarge on EBS gp3 1,569 and 495.
  • Same Xeon 6975P-C CPU, different disk: m8id.large on local NVMe answers 10,991 QD1 reads and 5,549 synced writes per second, m8i-flex.large on EBS gp3 1,745 and 528.
  • Same Xeon 6975P-C CPU, different disk: m8id.xlarge on local NVMe answers 11,033 QD1 reads and 7,853 synced writes per second, m8i-flex.xlarge on EBS gp3 1,765 and 527.
  • Same Neoverse-V3 CPU, different disk: m9gd.2xlarge on local NVMe answers 13,522 QD1 reads and 17,054 synced writes per second, m9g.2xlarge on EBS gp3 1,788 and 543.
  • Same Neoverse-V3 CPU, different disk: m9gd.xlarge on local NVMe answers 11,717 QD1 reads and 11,264 synced writes per second, m9g.xlarge on EBS gp3 1,713 and 543.
  • Same Neoverse-V3 CPU, different disk: m9gd.large on local NVMe answers 13,477 QD1 reads and 7,176 synced writes per second, m9g.large on EBS gp3 1,660 and 542.
  • Provisioning gp3 at 1000 MiB/s on m8a.2xlarge takes sequential reads from 402 MiB/s to 1,002 MiB/s (2.5×); QD1 reads barely move (1,791 → 1,792 IOPS): it buys bandwidth, not latency.
  • Provisioning gp3 at 1000 MiB/s on m8azn.xlarge takes sequential reads from 401 MiB/s to 1,001 MiB/s (2.5×); QD1 reads barely move (1,775 → 1,775 IOPS): it buys bandwidth, not latency.
  • tmpfs on m8a.2xlarge is not a disk at all: 1,156,332 QD1 reads and 432,842 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
  • tmpfs on m9g.2xlarge is not a disk at all: 653,143 QD1 reads and 301,463 synced writes per second, from RAM. The workspace is 31 GiB, sized by memory.
  • tmpfs on m9g.xlarge is not a disk at all: 680,179 QD1 reads and 303,962 synced writes per second, from RAM. The workspace is 15 GiB, sized by memory.
  • tmpfs on r8a.2xlarge is not a disk at all: 1,151,153 QD1 reads and 433,820 synced writes per second, from RAM. The workspace is 62 GiB, sized by memory.
  • tmpfs: 653k–1.2M QD1 reads and 301k–434k synced writes per second across 5 runners (5 RunsOn configurations).
  • VM disk: 5.9k–85k QD1 reads and 578–20k synced writes per second across 42 runners from 7 providers (6 Avrea runners, 6 Blacksmith runners, 3 GitHub runners, 11 Namespace runners, 2 StarSling x64 runners, 9 Ubicloud runners and 5 Warpbuild runners).
  • local NVMe: 11k–14k QD1 reads and 5.5k–17k synced writes per second across 9 runners (9 RunsOn configurations).
  • network disk: 6.8k–8.6k QD1 reads and 3.3k–4.7k synced writes per second across 2 runners (2 GitHub runners).
  • EBS: 1.6k–1.8k QD1 reads and 495–562 synced writes per second across 31 runners from 3 providers (3 AWS CodeBuild x64 runners, 25 RunsOn configurations and 3 Warpbuild configurations).

Reading these numbers

Most of a build is small, dependent I/O, not one big stream.

Why latency before throughput

  • git checkout, npm install, cargo build and test runners read thousands of small files, each waiting for the last: the QD1 column. QD32 and sequential show what a parallel compiler, a big docker build or a cache archive can pull; read them after.
  • Package managers, databases and SQLite-backed tests flush their writes. Each flush waits for the storage: network disks pay a round trip every time. That is sync write.

Every runner

Each provider's best disk first; its other machines sit behind +N more. Bars are log scale.

One request at a time 4k, 32 in flight Sequential, 1M
Runner QD1 read ↑p99 ↓Sync write ↑Read ↑Write ↑Read ↑Write ↑
x6456 runners, 9 providers
Blacksmith 4 vCPU 16 GB VM disk EPYC · model hidden nobarrier 85,326 14 µs 20,425 873,112 254,574 6.9 GiB/s 1.9 GiB/s
8 vCPU 8 vCPU 32 GB VM disk EPYC · model hidden nobarrier 58,340 19 µs 13,973 422,170 207,013 6.0 GiB/s 1.9 GiB/s
2 vCPU 2 vCPU 8 GB VM disk EPYC · model hidden nobarrier 64,450 18 µs 16,985 563,990 149,081 5.1 GiB/s 1.9 GiB/s
Namespace 8x32 8 vCPU 32 GB VM disk EPYC · model hidden 78,677 13 µs 18,166 1,112,926 412,276 4.2 GiB/s 3.0 GiB/s
8x16 8 vCPU 16 GB VM disk EPYC · model hidden 77,689 12 µs 17,807 1,124,263 476,969 4.1 GiB/s 2.9 GiB/s
4x16 4 vCPU 16 GB VM disk EPYC · model hidden 78,615 12 µs 17,888 1,011,926 436,056 4.1 GiB/s 2.9 GiB/s
4x8 4 vCPU 8 GB VM disk EPYC · model hidden 78,033 13 µs 18,087 1,015,510 403,484 4.1 GiB/s 3.0 GiB/s
2x8 2 vCPU 8 GB VM disk EPYC · model hidden 77,512 13 µs 18,068 350,533 222,621 4.1 GiB/s 2.9 GiB/s
StarSling 8 vCPU 32 GB VM disk EPYC · model hidden 54,269 17 µs 3,070 265,725 15,331 2.5 GiB/s 2.0 GiB/s
4 vCPU 4 vCPU 16 GB VM disk EPYC · model hidden 53,300 17 µs 3,023 227,686 11,527 3.3 GiB/s 2.8 GiB/s
Avrea 4 vCPU 16 GB VM disk EPYC 4585PX 47,519 18 µs 8,386 439,265 213,725 5.9 GiB/s 3.4 GiB/s
8 vCPU 8 vCPU 32 GB VM disk EPYC 4585PX 42,313 23 µs 7,570 245,805 139,439 7.4 GiB/s 3.3 GiB/s
2 vCPU 2 vCPU 8 GB VM disk EPYC 4585PX 45,672 17 µs 8,163 454,466 184,016 5.4 GiB/s 3.4 GiB/s
Warpbuild 2 vCPU 8 GB VM disk Ryzen 9 9950X 20,471 58 µs 2,056 144,184 66,887 7.6 GiB/s 4.6 GiB/s
8 vCPU 8 vCPU 32 GB VM disk Ryzen 9 9950X 19,636 60 µs 1,813 116,390 45,671 7.8 GiB/s 4.5 GiB/s
4 vCPU 4 vCPU 16 GB VM disk Ryzen 9 9950X 18,708 64 µs 1,980 124,074 38,883 7.8 GiB/s 5.1 GiB/s
RunsOn i7i.2xlarge 8 vCPU 64 GB local NVMe Xeon 8559C 12,709 228 µs 11,496 300,466 165,110 2.0 GiB/s 1.5 GiB/s
m8azn.3xlarge 12 vCPU 48 GB EBS gp3 EPYC 9R05 nobarrier 1,780 782 µs 540 2,997 2,994 401 MiB/s 401 MiB/s
c8a.2xlarge 8 vCPU 16 GB EBS gp3 EPYC 9R45 nobarrier 1,761 840 µs 538 2,996 2,991 401 MiB/s 401 MiB/s
m8a.2xlarge 8 vCPU 32 GB EBS gp3 EPYC 9R45 nobarrier 1,770 856 µs 538 2,995 2,991 401 MiB/s 401 MiB/s
m8a.2xlarge 8 vCPU 32 GB gp3 1000 MiB/s EPYC 9R45 nobarrier 1,677 897 µs 538 15,996 15,992 1,001 MiB/s 1,001 MiB/s
r8a.2xlarge 8 vCPU 64 GB tmpfs EPYC 9R45 1,151,153 <1 µs 433,820 4,231,165 3,828,376 18.6 GiB/s 10.1 GiB/s
m8a.2xlarge 8 vCPU 32 GB tmpfs EPYC 9R45 1,156,332 <1 µs 432,842 4,258,470 3,787,945 18.4 GiB/s 9.9 GiB/s
m8azn.xlarge 4 vCPU 16 GB EBS gp3 EPYC 9R05 nobarrier 1,775 963 µs 541 2,994 2,992 401 MiB/s 401 MiB/s
m8azn.xlarge 4 vCPU 16 GB gp3 1000 MiB/s EPYC 9R05 nobarrier 1,775 856 µs 541 15,994 15,991 1,001 MiB/s 1,001 MiB/s
m8id.2xlarge 8 vCPU 32 GB local NVMe Xeon 6975P-C 11,063 199 µs 7,665 134,258 67,075 905 MiB/s 432 MiB/s
m8i.2xlarge 8 vCPU 32 GB EBS gp3 Xeon 6975P-C nobarrier 1,569 1,012 µs 495 2,992 2,994 401 MiB/s 401 MiB/s
m8a.xlarge 4 vCPU 16 GB EBS gp3 EPYC 9R45 nobarrier 1,828 840 µs 540 2,994 2,991 401 MiB/s 401 MiB/s
m8i-flex.2xlarge 8 vCPU 32 GB EBS gp3 Xeon 6975P-C nobarrier 1,762 848 µs 519 2,991 2,995 401 MiB/s 401 MiB/s
c8a.xlarge 4 vCPU 8 GB EBS gp3 EPYC 9R45 nobarrier 1,794 856 µs 540 2,995 2,992 401 MiB/s 401 MiB/s
r8a.xlarge 4 vCPU 32 GB tmpfs EPYC 9R45 1,071,527 <1 µs 431,794 4,113,746 3,786,762 18.2 GiB/s 9.9 GiB/s
m8azn.large 2 vCPU 8 GB EBS gp3 EPYC 9R05 nobarrier 1,782 856 µs 542 2,994 2,992 401 MiB/s 401 MiB/s
m8i-flex.xlarge 4 vCPU 16 GB EBS gp3 Xeon 6975P-C nobarrier 1,765 774 µs 527 2,994 2,993 401 MiB/s 401 MiB/s
m8id.xlarge 4 vCPU 16 GB local NVMe Xeon 6975P-C 11,033 226 µs 7,853 67,089 33,526 451 MiB/s 216 MiB/s
m8i.xlarge 4 vCPU 16 GB EBS gp3 Xeon 6975P-C nobarrier 1,658 930 µs 521 2,992 2,994 401 MiB/s 401 MiB/s
i7i.xlarge 4 vCPU 32 GB local NVMe Xeon 8559C 11,437 301 µs 12,299 150,053 82,487 1,010 MiB/s 788 MiB/s
r8a.large 2 vCPU 16 GB EBS gp3 EPYC 9R45 nobarrier 1,844 881 µs 541 2,994 2,993 401 MiB/s 401 MiB/s
m8a.large 2 vCPU 8 GB EBS gp3 EPYC 9R45 nobarrier 1,780 774 µs 541 2,994 2,992 401 MiB/s 401 MiB/s
c8a.large 2 vCPU 4 GB EBS gp3 EPYC 9R45 nobarrier 1,809 856 µs 541 2,994 2,991 401 MiB/s 401 MiB/s
m8id.large 2 vCPU 8 GB local NVMe Xeon 6975P-C 10,991 228 µs 5,549 33,542 16,753 226 MiB/s 108 MiB/s
m8i.large 2 vCPU 8 GB EBS gp3 Xeon 6975P-C nobarrier 1,736 872 µs 534 2,992 2,994 401 MiB/s 401 MiB/s
i7i.large 2 vCPU 16 GB local NVMe Xeon 8559C 12,482 212 µs 13,567 75,023 41,229 503 MiB/s 395 MiB/s
m8i-flex.large 2 vCPU 8 GB EBS gp3 Xeon 6975P-C nobarrier 1,745 881 µs 528 2,993 2,994 401 MiB/s 401 MiB/s
t8i.medium 2 vCPU 4 GB EBS gp3 Xeon 6975P-C nobarrier 1,750 791 µs 524 3,000 2,998 401 MiB/s 401 MiB/s
Ubicloud premium 8 vCPU 32 GB VM disk Ryzen 9 7950X3D 10,127 358 µs 4,203 374,058 28,158 3.5 GiB/s 1.4 GiB/s
standard 8 vCPU 8 vCPU 32 GB VM disk EPYC 9454P 8,691 469 µs 1,670 264,258 11,860 2.6 GiB/s 944 MiB/s
premium 4 vCPU 4 vCPU 16 GB VM disk Ryzen 9 7950X3D 8,406 569 µs 1,706 232,993 16,062 2.7 GiB/s 882 MiB/s
standard 4 vCPU 4 vCPU 16 GB VM disk EPYC 9454P 5,898 1,122 µs 578 168,254 8,677 2.2 GiB/s 483 MiB/s
standard 2 vCPU 2 vCPU 8 GB VM disk EPYC 9454P 6,847 1,106 µs 1,549 112,843 54,236 2.6 GiB/s 809 MiB/s
premium 2 vCPU 2 vCPU 8 GB VM disk Ryzen 9 7950X3D 6,564 1,106 µs 2,417 138,040 64,867 2.7 GiB/s 867 MiB/s
GitHub 8 vCPU 32 GB VM disk EPYC 9V74 nobarrier 7,964 202 µs 3,968 38,954 18,761 784 MiB/s 532 MiB/s
4-core 4 vCPU 16 GB network disk EPYC 9V74 nobarrier 6,843 251 µs 3,262 19,580 11,712 394 MiB/s 394 MiB/s
2-core 2 vCPU 8 GB VM disk EPYC 9V74 nobarrier 6,531 289 µs 4,000 9,378 8,712 199 MiB/s 199 MiB/s
AWS CodeBuild large 8 vCPU 16 GB EBS Xeon 8275CL 1,752 750 µs 559 2,988 2,970 252 MiB/s 252 MiB/s
medium 4 vCPU 8 GB EBS Xeon 8124M 1,751 758 µs 558 2,993 2,968 252 MiB/s 252 MiB/s
small 2 vCPU 3 GB EBS Xeon 8275CL 1,749 823 µs 562 2,989 2,968 252 MiB/s 252 MiB/s
arm6433 runners, 7 providers
Warpbuild xfast (Apple M4 Pro) 6 vCPU 14 GB VM disk nobarrier 26,866 40 µs 8,615 339,443 5,906 22.4 GiB/s 27.2 GiB/s
xfast (Apple M4 Pro) 12 vCPU 12 vCPU 28 GB VM disk nobarrier 26,302 49 µs 8,636 334,760 5,979 19.8 GiB/s 26.8 GiB/s
8 vCPU 8 vCPU 32 GB EBS Neoverse-V2 nobarrier 1,776 782 µs 539 4,993 4,995 401 MiB/s 401 MiB/s
4 vCPU 4 vCPU 16 GB EBS Neoverse-V2 nobarrier 1,825 807 µs 539 4,195 4,193 301 MiB/s 302 MiB/s
2 vCPU 2 vCPU 8 GB EBS Neoverse-V2 nobarrier 1,797 832 µs 538 4,195 4,193 302 MiB/s 301 MiB/s
Blacksmith 8 vCPU 32 GB VM disk Ampere-1a nobarrier 26,024 36 µs 7,676 236,954 28,207 1.6 GiB/s 944 MiB/s
4 vCPU 4 vCPU 16 GB VM disk Ampere-1a nobarrier 25,073 37 µs 7,280 167,257 25,998 1.6 GiB/s 822 MiB/s
2 vCPU 2 vCPU 6 GB VM disk Ampere-1a nobarrier 23,450 43 µs 6,384 77,706 27,167 1.6 GiB/s 898 MiB/s
Namespace Apple silicon 4 vCPU 16 GB VM disk CPU not reported 18,455 85 µs 3,489 148,271 17,731 31.2 GiB/s 15.4 GiB/s
8x16 8 vCPU 16 GB VM disk Ampere-1a 16,751 93 µs 4,028 188,957 61,821 965 MiB/s 773 MiB/s
8x32 8 vCPU 32 GB VM disk Ampere-1a 15,964 98 µs 3,974 179,145 40,849 942 MiB/s 756 MiB/s
4x16 4 vCPU 16 GB VM disk Ampere-1a 16,892 90 µs 3,944 144,335 44,153 978 MiB/s 738 MiB/s
4x8 4 vCPU 8 GB VM disk Ampere-1a 17,191 89 µs 3,813 145,399 29,784 994 MiB/s 808 MiB/s
2x8 2 vCPU 8 GB VM disk Ampere-1a 14,513 106 µs 3,185 62,389 24,244 916 MiB/s 696 MiB/s
Avrea 8 vCPU 32 GB VM disk Apple M5 Max · as stated by the provider 18,297 53 µs 3,848 137,885 15,642 41.6 GiB/s 15.4 GiB/s
4 vCPU 4 vCPU 16 GB VM disk Apple M5 Max · as stated by the provider 17,294 55 µs 3,881 146,383 15,032 28.2 GiB/s 13.7 GiB/s
2 vCPU 2 vCPU 8 GB VM disk Apple M5 Max · as stated by the provider 17,388 59 µs 4,166 165,024 13,185 31.4 GiB/s 15.7 GiB/s
RunsOn m9gd.2xlarge 8 vCPU 32 GB local NVMe Neoverse-V3 13,522 81 µs 17,054 174,573 87,210 1.1 GiB/s 560 MiB/s
m9g.2xlarge 8 vCPU 32 GB EBS gp3 Neoverse-V3 nobarrier 1,788 807 µs 543 2,995 2,992 401 MiB/s 401 MiB/s
m9g.2xlarge 8 vCPU 32 GB tmpfs Neoverse-V3 653,143 <1 µs 301,463 2,651,587 2,495,659 14.3 GiB/s 9.1 GiB/s
c8g.2xlarge 8 vCPU 16 GB EBS gp3 Neoverse-V2 nobarrier 1,817 848 µs 539 2,994 2,992 401 MiB/s 401 MiB/s
m9g.xlarge 4 vCPU 16 GB EBS gp3 Neoverse-V3 nobarrier 1,713 938 µs 543 2,994 2,993 401 MiB/s 401 MiB/s
m9g.xlarge 4 vCPU 16 GB tmpfs Neoverse-V3 680,179 <1 µs 303,962 2,710,177 2,526,279 15.5 GiB/s 9.3 GiB/s
m9gd.xlarge 4 vCPU 16 GB local NVMe Neoverse-V3 11,717 216 µs 11,264 87,221 43,583 585 MiB/s 280 MiB/s
c8g.xlarge 4 vCPU 8 GB EBS gp3 Neoverse-V2 nobarrier 1,703 897 µs 539 2,993 2,994 401 MiB/s 401 MiB/s
m9gd.large 2 vCPU 8 GB local NVMe Neoverse-V3 13,477 81 µs 7,176 43,606 21,782 293 MiB/s 141 MiB/s
m9g.large 2 vCPU 8 GB EBS gp3 Neoverse-V3 nobarrier 1,660 963 µs 542 2,992 2,994 401 MiB/s 401 MiB/s
c8g.large 2 vCPU 4 GB EBS gp3 Neoverse-V2 nobarrier 1,779 815 µs 537 2,993 2,994 401 MiB/s 401 MiB/s
Ubicloud standard 2 vCPU 8 GB VM disk Neoverse-N1 9,482 334 µs 5,425 93,715 55,422 937 MiB/s 637 MiB/s
standard 8 vCPU 8 vCPU 32 GB VM disk Neoverse-N1 9,278 358 µs 2,433 222,352 101,505 875 MiB/s 438 MiB/s
standard 4 vCPU 4 vCPU 16 GB VM disk Neoverse-N1 8,892 449 µs 2,132 164,716 26,833 798 MiB/s 428 MiB/s
GitHub 4 vCPU 16 GB network disk Neoverse-N2 nobarrier 8,635 202 µs 4,666 19,595 11,786 394 MiB/s 337 MiB/s
2-core 2 vCPU 8 GB VM disk Neoverse-N2 nobarrier 7,441 230 µs 3,066 9,377 6,115 199 MiB/s 163 MiB/s
  • IOPS are 4k blocks, medians of finished jobs
  • bold best disk measured on the arch (tmpfs is RAM, left out)
  • volatile, nobarrier: cheaper, less durable flushes

Pick the disk per job on RunsOn

  • The instance type sets the storage under the workspace: EBS gp3 by default, provisioned gp3 for bandwidth, local NVMe on the d and i families, or tmpfs when the job fits in RAM. Pick it with the family label; each one, measured: EC2 storage benchmark.

How it's measured

Methodology

fio runs inside the job, in GITHUB_WORKSPACE: the disk a build writes to, on the filesystem the provider mounted there. Every runner is compared, whatever its shape.

fio settingsblock sizes, queue depths, mounts
  • QD1 random read: 4k blocks, one request in flight; IOPS and p99 latency.
  • Sync write: 4k writes, each followed by fdatasync.
  • Random read and write: 4k blocks, 4 jobs with 32 requests in flight each.
  • Sequential read and write: 1M blocks, 32 requests in flight.
  • The workspace filesystem, its size and its mount options come from the job's own mount table.
  • Earlier versions of this page used a different harness and other fio settings on 2 vCPU runners; those numbers are gone rather than mixed in.