Using Spot Instances for CI Runners on AWS
CI jobs are a good fit for spot capacity: they are short, restartable and usually not urgent. EC2 Spot can cost far less than on-demand, and AWS says discounts can reach up to 90 percent. The catch is that AWS can reclaim the instance with two minutes' notice. This guide shows how to design runners that tolerate that, and how to estimate savings honestly.
What you are buying
Spot instances are spare EC2 capacity sold at a variable price. Facts from AWS's docs:
- You pay the current spot price, which changes slowly, not an auction bid.
- When EC2 needs the capacity back, it sends a Rebalance Recommendation signal (when available) and an interruption notice 2 minutes before the instance is stopped or terminated.
- Interruption frequency varies by instance type and region; the Spot Instance Advisor shows a rate bucket for each type (below 5 percent, 5-10, 10-15, 15-20, above 20).
Use the EC2 cost calculator for on-demand figures and compare them with the current spot price from the CLI:
aws ec2 describe-spot-price-history \
--instance-types c7a.2xlarge c6a.2xlarge m6a.2xlarge \
--product-descriptions "Linux/UNIX" \
--start-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--query 'SpotPriceHistory[].[InstanceType,AvailabilityZone,SpotPrice]' --output table
The architecture
A common design: an Auto Scaling group or EC2 Fleet launches instances; each instance boots, registers as an ephemeral GitHub runner, takes one job, and terminates. A scaler adds instances when jobs queue, using the workflow_job webhook or the runners API.
Key choices:
- Ephemeral, single-job runners. Register with
./config.sh --ephemeral(or just-in-time configuration through the API). A runner that disappears with its instance should not remain registered; ephemeral registration removes it after one job and avoids stale offline runners. - Diversify instance types and zones. Spot pools are per type per zone. Ask for many equivalent types (for example c6a, c7a, c6i, m6a at the same size) across all zones. Use the
price-capacity-optimizedallocation strategy, which AWS recommends for most workloads, to balance low price against interruption risk. - Pre-baked AMI. Boot time is billed and lengthens queue time. Put the runner software, Docker, language toolchains and a warm Docker image cache in the AMI.
- Fast local disk. Instances with NVMe instance storage make builds quicker; gp3 volumes are cheaper and fine for most jobs. Use the EBS gp3 cost calculator to size them.
- Fallback to on-demand. When spot capacity is empty, launch on-demand rather than leave jobs queued. Log how often this happens. If the fallback rate is high, your spot savings shrink.
Handling interruptions
An interrupted job fails with a message like "The runner has received a shutdown signal" or "lost communication with the server". GitHub does not automatically re-run the job. Handle it in layers:
- Make interruptions rare. Prefer instance types with a low interruption rate in the Advisor, and keep jobs short. A job that takes 8 minutes has far less exposure than one that takes 90.
- Watch for the notice. The instance metadata service exposes
spot/instance-actionwhen an interruption is scheduled. A small daemon can poll it and mark the runner as draining; it cannot save a running job, but it can stop accepting new work. - Retry automatically. A scheduler that receives the
workflow_jobcompleted event with a failure caused by runner loss can call the re-run-failed-jobs API:
gh run rerun <run-id> --failed
- Keep jobs idempotent. Deployments, database migrations and releases are poor candidates for spot. Run them on on-demand runners using labels (
runs-on: [self-hosted, ondemand]), and let tests and builds use spot (runs-on: [self-hosted, spot]). - Checkpoint long tasks. For big builds, use remote build caches so a retry resumes from cached results instead of starting from zero.
Savings math
Do the arithmetic with real prices, including everything. Per busy hour:
on-demand cost per hour = instance price
spot cost per hour = spot price (varies)
overhead multiplier = 1 + boot + idle + retries (typically 1.1 to 1.5)
effective spot cost per hour = spot price x overhead
Example only, with illustrative numbers: on-demand at $0.34 an hour, spot at $0.12, and a 1.3 overhead multiplier gives an effective $0.156 an hour, 54 percent below on-demand. If 10 percent of minutes fall back to on-demand, the blended cost rises. Compare the result with GitHub's larger-runner prices (8-core Linux is $0.022 a minute, $1.32 an hour, per its pricing page) and with your operations time, as covered in self-hosted vs GitHub-hosted runners.
Also count:
- EBS volumes and snapshots, public IPv4 addresses (billed per hour), NAT gateway and data transfer. Use the data transfer estimator if jobs pull a lot from outside the region.
- The engineer time to build and run the scaler. Open source options include Actions Runner Controller on Kubernetes, and several AWS-based autoscaler projects.
Security basics
Spot does not change the security model: runners are in your network. Use ephemeral instances, IMDSv2, an instance role with minimum rights, no public-repo jobs, and OIDC for cloud access (see GitHub Actions OIDC to AWS). Tag instances so cost reports can attribute spend.
When not to use spot
- Release and deploy jobs.
- Jobs longer than an hour with no cache or checkpointing.
- Tiny volumes, where a hosted runner's included minutes cost nothing.
- Teams with no capacity to operate autoscaling.
FAQ
How often are CI jobs interrupted?
It depends on instance type, zone and time. Use the Advisor's rate bucket as a guide and measure your own interruption rate for a week before committing.
Can I use Spot with Kubernetes runners?
Yes. Run Actions Runner Controller on a node group of spot instances with several instance types, and set pod disruption handling. Interrupted pods fail their jobs, so the same retry strategy applies.
Is Spot cheaper than a Savings Plan?
For steady 24x7 baseline load, a Compute Savings Plan can beat spot once overhead is included. A mix is common: reserved capacity for the baseline, spot for bursts.
What about GitHub-hosted and third-party runners?
They are billed per minute and hide the capacity management. Some providers use spot internally and pass part of the saving on.
How compiler.dev relates
compiler.dev runs its Linux runners on spot capacity with automatic fallback to on-demand, so you get the savings without running the scaler yourself; the comparison mode shows run time and cost for your own jobs next to GitHub-hosted ones.
Made by compiler.dev. Free tools · Pricing