self-host →

v3.4.0

View on GitHub Upgrade guide

Spotlight

  • Fewer out-of-memory failures on small Linux runners. Runners with 16 GiB of RAM or less get compressed in-memory swap (zram); to opt out, run sudo swapoff -a as a job's first step (for example for kubeadm). Jobs that run out of memory now get a warning. The kernel also kills job processes before the GitHub runner, so these jobs no longer end as lost runners. Possibly a breaking change so let me know if you hit issues.
  • Default runners now use 8th-generation or newer instance families (m8/c8/r8 on x64, m8g/c8g/r8g on arm64), and arm64 runners also accept Graviton5 (m9g/c9g/r9g) where available. 1cpu-linux-x64 now picks 1–2 vCPU types with 4 GiB of memory. That includes burstable t8i.medium in unlimited credit mode; use runner=1cpu-linux-x64/family=c8 to avoid it. In ap-northeast-3 (Osaka), which has no 8th-generation x86 families, x64 jobs need a custom runner, for example with family=m7+c7+r7.
  • Faster Docker builds with Magic Cache (extras=s3-cache) and BuildKit type=gha caching: in our tests, an unchanged 620-layer export dropped from 55s to 30s.

CloudFormation

  • New EnableWarmPools parameter (default true). Set it to false to drop warm instances on one stack, such as a secondary region; pool= jobs still run there on new instances.
  • Parameter changes made with Use current template are now applied instead of ignored. Upgrade to v3.4.0 itself with Replace current template. Its change set flags 13 resources as Replacement: True, but only the task definition is replaced.

Terraform

  • Check before upgrading Fleet: instances carry a new runs-on-launch-token tag, and the task role gains dynamodb:ConditionCheckItem. If an SCP restricts aws:TagKeys on ec2:CreateFleet or ec2:CreateTags, allow the tag, or launches are denied. If you set permission_boundary_arn, allow the action.
  • New enable_warm_pools input (Flex), equivalent to EnableWarmPools.
  • New otel_resource_attributes input (Flex, Fleet). OTEL_RESOURCE_ATTRIBUTES now overrides RunsOn defaults such as deployment.environment, except service.name. Fixes runs-on#556.
  • New locks_table_point_in_time_recovery_enabled input (Flex) enables point-in-time recovery on the locks table. log_retention_days now also covers control-plane logs (minimum 14 days), so values above 14 keep them longer. Fixes terraform-aws-runs-on#65.
  • Slack alerts no longer fail applies on Terraform 1.5/1.6.
  • The Fleet module requires Terraform/OpenTofu 1.7.0+ and now parses with Pulumi HCL. Upgrades from v3.0.x plan to delete aws_secretsmanager_secret_version.config; this is expected.
  • The Flex budget is now tagged.

Other fixes

  • Flex and Fleet are more resilient to failures, restarts and bursts of jobs:
    • Fleet no longer kills a runner in the middle of a job. GitHub can give a Fleet runner a different job than the one it was launched for. When that original job was then cancelled (for example by concurrency) or timed out, Fleet terminated the runner, and the job it was actually running failed.
    • Fleet handles bursts of hundreds of jobs more reliably. Instances from slow or lost launches are found and terminated without holding up the fleet's other repairs. After a Spot shortage, a fleet launches on-demand for at least 5 minutes.
    • Flex restarts and deploys let in-flight launches finish, then flush alerts and telemetry. Before, a restart could occasionally leave an instance untracked until housekeeping terminated it.
    • Flex job state stays consistent under concurrent updates: completed webhooks are no longer dropped, exhausted jobs no longer relaunch, and a throttled tag write can no longer get a busy instance terminated.
    • Warm pools keep their capacity through transient errors. A failed GitHub config fetch no longer makes Flex terminate every pool instance, and partial launches, throttled starts and scale-downs no longer waste or kill instances.
    • StickyDisk caches no longer roll back to older snapshots or keep data from Spot-interrupted steps, and a deleted volume continues cold instead of stalling.
    • Boot-dead Linux Fleet runners are replaced after 15 minutes instead of 20.
    • Rare stuck jobs are fixed: a job could stay queued after a GitHub server error during runner registration (409 Already exists), or after a sibling matrix job took its runner. Addresses runs-on#546.
  • Merge queues and runs-on/spot-retry-gate now recognize Spot interruptions even when a post step outlasts the grace period. On Linux, the annotation is written on the interrupted step.
  • Flex jobs that wait for deployment approval no longer get a runner before the approval check sees them, start as soon as they are approved, and no longer log every poll.
  • Idle Fleet stacks on GitHub Enterprise Server no longer busy-poll.
  • The agent flushes telemetry on boot errors, crashes and terminations, so workflows no longer need a final sleep step. OTLP exporter headers now reach only extras=otel runners, in root-only files.
  • More accurate metrics: Fleet no longer records a job when GitHub cancels or re-issues it before a runner starts it. Flex in_progress and runner counts are correct, failed steps are marked as trace errors, and end-of-job charts now work on large runners.
  • The "⏱️ Timings" table in each job's "Set up job" step now includes the GitHub runner's startup (connecting to GitHub, then listening for jobs), so slow runner sessions show up in the job log.
  • Magic Cache: artifact actions work without runs-on/action, concurrent uploads of a key no longer fail, and cache isolation works on GHE.com.
  • extras=efs runners no longer hang for 90s at shutdown, and CloudWatch runner logs are no longer doubled. Custom images that use logind honour EC2 stops after a warm boot.
  • Fleet standby instances use Fleet names. Fixes terraform-aws-runs-on#44.
  • roc: log errors are reported, --watch stops spending GitHub quota, and stack doctor exits 1 on failure, with pass/fail/skip in checks.json. connect and interrupt need a v3.1.0+ stack.
  • Security updates for Go dependencies, including OpenTelemetry Go v1.45.0 (GO-2026-6508).

Release resources