roc CLI
Command line tool to manage and troubleshoot your RunsOn installation
RunsOn CLI (roc) is a command line tool to manage and troubleshoot your RunsOn ↗ installation.
It works with modern RunsOn installations, including current v3 stacks. The CLI assumes you run it with AWS credentials that can read the RunsOn stack resources, CloudWatch logs, SSM sessions, EC2 instance metadata, and any service used by the selected command.
Usage#
Flex
Most troubleshooting starts with the GitHub Actions job URL:
JOB_URL="https://github.com/acme/app/actions/runs/123456789/job/987654321"
AWS_PROFILE=runs-on-admin roc logs "$JOB_URL" --include=console --watchAWS_PROFILE=runs-on-admin roc connect "$JOB_URL"AWS_PROFILE=runs-on-admin roc interrupt "$JOB_URL" --waitUse logs first, connect when you need a live SSM shell, and interrupt to test Spot interruption handling.
Fleet
For Fleet, start with roc logs when you need diagnostics for a specific job URL. The v3.2 CLI resolves the Fleet claim and durable Spot-interruption evidence before it streams logs; roc logs --full exports that correlation with every attempted runner instance. roc connect and roc interrupt remain primarily Flex troubleshooting tools; use Fleet troubleshooting for capacity and recovery checks.
Installation#
mise#
mise ↗ installs CLI versions side by side. Pin the exact
version for each stack so roc changes automatically when you enter its project
directory:
mise use --pin 'github:runs-on/cli[bin=roc]@3.2.3'Install the latest stable CLI as your global default:
mise use --global 'github:runs-on/cli[bin=roc]@latest'Run an exact version once without changing your configuration:
mise x 'github:runs-on/cli[bin=roc]@3.2.3' -- roc versionDownload Binary#
Download the exact CLI version that matches your stack from the Releases ↗ page.
Example (macOS ARM64):
VERSION=v3.4.0curl -Lo ./roc https://github.com/runs-on/cli/releases/download/${VERSION}/roc_${VERSION}_darwin_arm64chmod a+x ./rocxattr -d com.apple.quarantine ./roc./roc --helpExample (Linux AMD64):
VERSION=v3.4.0curl -Lo ./roc https://github.com/runs-on/cli/releases/download/${VERSION}/roc_${VERSION}_linux_amd64chmod a+x ./roc./roc --helpGitHub Action#
You can use the RunsOn CLI in your GitHub Actions workflows by including it as a step:
- uses: runs-on/cli@main with: version: 'latest' # Optional: defaults to 'latest'Example workflow:
name: Lint RunsOn Config
on: pull_request: paths: - '.github/runs-on.yml'
jobs: lint: runs-on: ubuntu-latest steps: - uses: actions/checkout@v7
- uses: runs-on/cli@main
- name: Lint runs-on.yml run: roc lint .github/runs-on.ymlCore Commands#
roc connect#
Connect to the instance running a specific job via SSM by passing the GitHub Actions job URL.
This feature requires the AWS CLI and the AWS Session Manager plugin ↗ (session-manager-plugin on your PATH). roc checks both before it looks up the job. It works on macOS, Linux and Windows; on Windows, roc runs aws ssm start-session as a child process, so Ctrl-C goes to the remote shell.
Usage: roc connect JOB_URL [flags]
Flags: -h, --help help for connect --watch Wait for instance ID if not found
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")Example:
AWS_PROFILE=runs-on-admin roc connect https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899roc logs#
Fetch RunsOn server and instance logs for a specific job URL. Use the --include flag to specify additional streamed log types, or --full to export a complete diagnostic archive.
Usage: roc logs JOB_URL [flags]
Flags: -f, --format string Output format: long (default) or short (default "long") --full Export full diagnostic archive for the job -h, --help help for logs --include strings Include additional log types: 'run' (all logs from entire run), 'console' (EC2 instance console logs) --no-color Disable color output for streamed logs (also off when stdout is not a terminal or NO_COLOR is set) -w, --watch string[="5s"] Watch for new logs with optional interval (e.g. --watch 2s)
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")The lookback window for streamed roc logs is fixed at the last 2 hours. The command first invokes the stack’s job diagnostics resolver. It detects Flex or Fleet, correlates local job or claim state, and returns durable Spot-interruption evidence when it is available.
If roc cannot read a log source (for example, CloudWatch returns AccessDenied, or an instance’s console output is unavailable), it prints Warning: cannot read <source> logs: ... on stderr, once per distinct error. Without --watch, it still prints everything it collected, then exits with status 1 (N log source(s) failed; output is incomplete). With --watch, it keeps retrying CloudWatch sources on every interval.
Streamed logs are colored only when stdout is a terminal. Piping or redirecting the output, setting NO_COLOR, or passing --no-color turns color off.
--full writes a roc-logs-<job_id>-<timestamp>.zip archive instead of streaming. It contains the resolver response (diagnostics/resolver.json), the local job or claim record (diagnostics/local-record.json), control-plane logs for the job and run, CloudTrail events for every attempted instance, EC2 console output, and agent logs. Its time window starts five minutes before the job and ends ten minutes after it.
The resolver exposes a redacted diagnostic settings summary. It does not include credentials, OTLP endpoints or headers, policy bodies, resource tags, or other secret configuration. For ambiguous Fleet jobs where the resolver cannot fetch workflow details, roc logs can fall back to the local GitHub CLI; install gh and authenticate it with repository Actions read access.
--full cannot be combined with --watch. The job-specific command does not accept --since; use roc stack logs --since ... for stack-wide streaming.
Examples:
# Fetch logs for a specific job (default behavior)AWS_PROFILE=runs-on-admin roc logs https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --watch
# Fetch all application logs for a run (all jobs in the run)AWS_PROFILE=runs-on-admin roc logs https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --include=run --watch
# Fetch EC2 instance console logsAWS_PROFILE=runs-on-admin roc logs https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --include=console
# Fetch both run logs and console logsAWS_PROFILE=runs-on-admin roc logs https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --include=run,console --watch
# Export a complete diagnostic archiveAWS_PROFILE=runs-on-admin roc logs https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --fullroc cleanup#
Delete re-creatable cache data for the ref used by a GitHub Actions job. roc cleanup removes classic cache objects, isolated cache objects, and sticky-disk EBS snapshots associated with that ref. It never touches repository-wide user caches under cache/repo/<org>/<repo>.
The command shows its deletion plan and asks for confirmation. Pull-request runs target refs/pull/N/merge; other runs target the head ref as both a branch and a tag. Use --include-default-branch only when you also intend to clear the repository’s shared default-branch lineage.
Usage: roc cleanup JOB_URL [flags]
Flags: --dry-run list what would be deleted without deleting anything -h, --help help for cleanup --include-default-branch also delete caches and snapshots for the repository default branch --yes skip the confirmation prompt
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")roc cleanup works with Flex and Fleet. It requires gh authenticated with repository read access, access to the selected stack’s configuration secret and diagnostics resolver, S3 list/delete permission for the cache bucket, and EC2 describe/delete-snapshot permission.
# Inspect the ref-scoped cleanup planAWS_PROFILE=runs-on-admin roc cleanup "$JOB_URL" --dry-run
# Confirm cleanup for the job ref and its default-branch lineageAWS_PROFILE=runs-on-admin roc cleanup "$JOB_URL" --include-default-branch --yesroc interrupt#
Trigger a spot interruption on the instance running a specific job, simulating a spot instance interruption for testing purposes.
This command uses AWS Fault Injection Simulator (FIS) to send a spot interruption notification to the running instance.
Usage: roc interrupt JOB_URL [flags]
Flags: --delay duration Delay before interruption (e.g., 2m, 30s) (default 5s) -h, --help help for interrupt -w, --wait Wait for instance ID if not found
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")Requirements:
- The target instance must be a running spot instance
- AWS FIS service must be available in your region
How it works:
- Validates the instance is a running spot instance
- Uses the
aws-fis-itnIAM role for FIS, creating it the first time - Creates and starts a FIS experiment that sends the interruption notice after
--delay - Prints the experiment’s progress until EC2 interrupts the instance, two minutes after the notice
- Deletes the experiment template when done
Permissions: fis:ListExperimentTemplates, fis:CreateExperimentTemplate, fis:StartExperiment, fis:GetExperiment, fis:DeleteExperimentTemplate, ec2:DescribeInstances, and iam:GetRole and iam:PassRole on the aws-fis-itn role. iam:CreateRole and iam:PutRolePolicy are only needed the first time, while that role does not exist yet.
Stopping: once roc starts creating FIS resources, the first Ctrl-C (or SIGTERM) stops watching and deletes the experiment template; a second one ends roc at once. If the experiment has started, roc prints the aws fis stop-experiment command to stop it. Once the interruption notice is sent, stopping the experiment does not undo it.
Example:
AWS_PROFILE=runs-on-admin roc interrupt https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899# Wait for instance if job hasn't started yetAWS_PROFILE=runs-on-admin roc interrupt https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --wait
# Custom delay before interruption (default is 5 seconds)AWS_PROFILE=runs-on-admin roc interrupt https://github.com/runs-on/runs-on/actions/runs/12415485296/job/34661958899 --delay 30sroc lint#
Validate and lint runs-on.yml configuration files. This command validates your configuration files against the RunsOn schema, checking for syntax errors, invalid values, missing required fields, and schema violations.
When no file path is provided, the command recursively searches for all runs-on.yml files in the current directory and subdirectories. It skips .git and node_modules directories; other hidden directories such as .github are still searched.
Usage: roc lint [flags] [file]
Flags: -f, --format string Output format: text, json, or sarif (default "text") -h, --help help for lint --stdin Read from stdin instead of file
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")What it validates:
- YAML syntax errors
- Schema validation for all top-level fields (
_extends,runners,images,pools,admins) - Required fields and valid value types
- Pool configuration (name pattern, schedule values, runner references)
- Runner specifications (CPU, RAM, family, spot values, etc.)
- Image specifications (AMI IDs, platform, architecture, etc.)
- Custom fields are allowed (e.g.,
x-defaultsfor YAML anchors)
Output formats:
text(default): Human-readable output with file status and diagnosticsjson: Structured JSON output for CI/CD integrationsarif: SARIF format for GitHub Code Scanning and other tools
Examples:
# Lint a specific configuration fileroc lint .github/runs-on.yml
# Lint all runs-on.yml files recursively (no arguments)roc lint
# Lint from stdincat runs-on.yml | roc lint --stdin
# Lint with JSON output for CI/CD pipelinesroc lint config/runs-on.yml --format json
# Lint with SARIF output for GitHub Code Scanningroc lint .github/runs-on.yml --format sarifIntegration with CI/CD:
The command exits with a non-zero status code when validation errors are found:
# Exit code 0 for valid config, 1 for invalidif roc lint runs-on.yml --format json > validation-report.json; then echo "Configuration is valid!"else echo "Configuration validation failed. See validation-report.json for details." exit 1fiPre-commit hook:
You can use roc lint as a pre-commit hook to automatically validate runs-on.yml files before committing. First, install pre-commit ↗:
pip install pre-commitThen add the hook to your .pre-commit-config.yaml:
repos: - repo: https://github.com/runs-on/cli rev: v0.1.13 # Use the latest release tag hooks: - id: roc-lintFinally, install the git hook scripts:
pre-commit installNow roc lint will automatically run on staged runs-on.yml files before each commit. The commit will be blocked if validation errors are found.
roc version#
Display the version of the roc CLI.
roc versionStack Management#
roc stack doctor#
Diagnose RunsOn stack health and export troubleshooting information.
This command performs health checks on your RunsOn stack:
- Checks ECS service health (running vs desired tasks)
- Tests endpoint accessibility (Flex stacks only)
- Validates service readiness via
/readyz, incl. GitHub App configuration (Flex stacks only) - Fetches application logs
Fleet stacks skip the endpoint and readiness checks (no public endpoint).
Results are exported as a timestamped ZIP file containing checks.json and logs. Each check in checks.json has a status of pass, fail or skip. When any check fails, the command still exports the ZIP file, then exits with status 1 and names the failed checks, so you can use it in scripts.
Usage: roc stack doctor [flags]
Flags: -h, --help help for doctor --since string Fetch logs since duration (e.g. 30m, 2h, 24h) (default "24h")
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")Example:
AWS_PROFILE=runs-on-admin roc stack doctor --since 2hOutput:
Checking service (https://us-east-1.console.aws.amazon.com/ecs/v2/clusters/.../services/.../configuration/overview)... ✅ (status: RUNNING (1/1 tasks))Checking service endpoint (https://example.execute-api.us-east-1.amazonaws.com/prod)... ✅Checking service readiness... ✅ (app_tag: v3.1.3)Fetching application logs (since 24h0m0s)... ✅ (5419 lines)
Full results exported to: /Users/crohr/dev/runs-on/cli/roc-doctor-2025-06-20-12-40-29.ziproc stack logs#
Stream all RunsOn application logs from CloudWatch log streams.
This command streams all application logs from the RunsOn service, not filtered by specific jobs. Use this to monitor overall service activity and troubleshoot system-wide issues.
Usage: roc stack logs [flags]
Flags: -f, --format string Output format: long (default) or short (default "long") -h, --help help for logs --no-color Disable color output for streamed logs (also off when stdout is not a terminal or NO_COLOR is set) -s, --since string Show logs since duration (e.g. 30m, 2h) (default "2h") -w, --watch string[="5s"] Watch for new logs with optional interval (e.g. --watch 2s)
Global Flags: -d, --debug Enable debug output --stack string CloudFormation stack name (default "runs-on")Examples:
# Stream last 2 hours of application logs (default)AWS_PROFILE=runs-on-admin roc stack logs
# Stream last 24 hours of logsAWS_PROFILE=runs-on-admin roc stack logs --since 24h
# Stream logs with watch mode (refreshes every 5 seconds)AWS_PROFILE=runs-on-admin roc stack logs --watch
# Stream logs with custom watch intervalAWS_PROFILE=runs-on-admin roc stack logs --watch 10s
# Stream logs in short format without colorAWS_PROFILE=runs-on-admin roc stack logs --format short --no-colorUse Cases#
Debugging Job Failures#
When a GitHub Actions job fails on your RunsOn infrastructure:
- Get logs: Use
roc logs JOB_URLto fetch detailed logs - Connect to instance: Use
roc connect JOB_URLto SSH into the runner - Check stack health: Run
roc stack doctorto ensure infrastructure is healthy - Test interruption handling: Use
roc interrupt JOB_URLto simulate spot interruptions
Real-time Monitoring#
# Monitor logs in real-time for active jobsroc logs JOB_URL --watch
# Monitor all application logsroc stack logs --watchHealth Checks#
# Regular stack health verificationroc stack doctor
# Stream application logs for monitoringroc stack logs --since 1hTroubleshooting#
Stack not found#
If roc reports that a RunsOn stack was not found, check the stack name (--stack or RUNS_ON_STACK_NAME) and the AWS region (AWS profile or AWS_REGION). The error lists the stacks that have a configuration secret in the current region, or suggests the only one (Did you mean --stack ...?). Listing them requires secretsmanager:ListSecrets; without it, the error omits the list.
Connection Issues#
If roc connect fails:
- Ensure AWS Session Manager plugin is installed
- Verify your AWS credentials have appropriate permissions to access the RunsOn stack
- Check that you are targeting the correct stack (
--stackflag) and region (AWS profile orAWS_REGIONenvironment variable) - Use
--debugflag for detailed error information
Log Access Issues#
If roc logs returns no results:
- Read any
Warning: cannot read <source> logslines on stderr: they name the log sourceroccould not read and the AWS error - Verify the job URL is correct
- Check your AWS credentials and permissions
- Check that you are targeting the correct stack (
--stackflag) and region (AWS profile orAWS_REGIONenvironment variable) - Try adjusting the
--sincetimeframe
Interruption Testing Issues#
If roc interrupt fails:
- Ensure the target instance is a running spot instance
- Verify AWS FIS service is available in your region
- Check that your AWS credentials have appropriate FIS permissions
- Use
--debugflag for detailed error information
Support#
For issues with the RunsOn CLI:
- Open an issue on GitHub ↗
- Contact support at ops@runs-on.com
- Include output from
roc stack doctorwhen reporting problems
License#
This project is licensed under the MIT License - see the LICENSE ↗ file for details.