Test Sharding and Parallel Tests in GitHub Actions
When a test suite takes 20 minutes, you have two ways to make it faster: run tests in parallel inside one job, or split them across several jobs (sharding). Both work. This guide shows when to use each, how to set up a shard matrix for common frameworks, and how to avoid the traps: unbalanced shards, lost coverage and a bigger bill.
Parallel in a job versus sharding across jobs
In-job parallelism uses the cores you already have. A 2-vCPU runner gives you at most about two worker processes, so the win is small unless you move to a larger runner:
pytest -n auto # pytest-xdist, one worker per core
npx jest --maxWorkers=100%
go test ./... -p 4
Sharding runs the suite as N separate jobs, each with its own runner. Wall-clock time drops roughly by N, minus setup time. Billed minutes go up, because every shard pays for checkout, install and any build.
A simple rule: if setup is 1 minute and the suite is 20 minutes, 4 shards give about 6 minutes of wall time and 24 billed minutes instead of 21. Use the test sharding planner to compute the shard count that fits your target time and budget.
A shard matrix
GitHub matrix jobs can generate one job per shard. Pass the shard index to the test runner:
jobs:
test:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3, 4]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- run: npm ci
- run: npx jest --shard=${{ matrix.shard }}/4
fail-fast: false lets the other shards finish, so you see every failure in one run instead of fixing them one at a time. Remember the limit: a matrix can create at most 256 jobs per workflow run, which is not a concern for sharding.
Other frameworks have the same idea:
npx playwright test --shard=2/4 # Playwright
npx vitest run --shard=2/4 # Vitest
pytest --splits 4 --group 2 # pytest-split
For Go, split by package list, for example with go list ./... | split or a tool like gotestsum with a package list per shard. Cypress and RSpec have parallelization tools (Cypress Cloud, parallel_tests, knapsack) that balance by past timing.
Balance shards by time, not by count
Built-in sharding usually splits by file count or alphabetical order. If one file takes 8 minutes, the shard holding it sets your total time. Better balancing uses recorded durations:
- Run the suite once with timing output (JUnit XML or the framework's JSON reporter).
- Save the durations as an artifact or commit a timings file.
- Assign files to shards with a greedy algorithm: sort by duration, put each file on the least loaded shard.
pytest-split stores durations in .test_durations and uses them. Playwright shards by test file, so splitting big spec files into smaller ones helps more than any algorithm. A good target is for the slowest shard to be within 10 to 15 percent of the average.
Merge reports and coverage
Each shard produces a partial report. Upload them as artifacts with unique names and merge in a final job:
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with:
name: junit-${{ matrix.shard }}
path: reports/
report:
needs: test
if: ${{ !cancelled() }}
runs-on: ubuntu-latest
steps:
- uses: actions/download-artifact@v4
with:
pattern: junit-*
merge-multiple: true
Playwright has merge-reports for its blob reporter. For coverage, upload each shard's lcov or coverage XML and merge before sending to your coverage service; otherwise each shard reports only its slice and the number looks low.
One required check, not N
Branch protection needs stable check names. A matrix produces names like test (1), test (2), and changing the shard count changes the names, which breaks required checks. Add one final job that depends on all shards and require only that one:
tests-passed:
needs: test
if: ${{ always() }}
runs-on: ubuntu-latest
steps:
- run: |
if [ "${{ needs.test.result }}" != "success" ]; then exit 1; fi
The always() matters: if a shard fails, a skipped dependent job reports as skipped, and a skipped required check can count as passing.
Keep setup small
Sharding multiplies setup cost, so trim it:
- Cache dependencies (guide).
- Build once in a prior job and share build output as an artifact, then have shards download it, if compilation is slow.
- Skip browsers you do not use.
npx playwright install --with-deps chromiumis faster than installing all three.
Watch the cost
Each shard bills per started minute, rounded up per job. Many tiny shards waste money on rounding and setup. With 4 shards of 5 minutes each you pay for 20 minutes, the same as one 20-minute job, and developers wait one quarter as long. With 20 shards of 1 minute each plus 1 minute setup, you pay 40. Try the forecaster to model the monthly effect.
FAQ
How many shards should I use?
Start with the number that reaches your time target using the planner: suite time divided by target, plus a margin for setup. Then check the slowest shard and rebalance.
Does sharding find more flaky tests?
It can expose order dependence, since tests no longer run in one process. Fixing that is useful. See dealing with flaky tests.
Is a bigger runner better than more shards?
If the framework scales across cores, a larger runner avoids repeated setup. If tests are I/O-bound or single-process, shards scale better. Try both on one workflow.
How do I keep the same tests on every PR?
Pin the shard count in the matrix and keep a single required summary job, as shown above.
Try it on your jobs
Compare the shard approach with a faster single machine: compiler.dev's comparison mode runs the same workflow on larger runners so you can see time and cost side by side before you restructure your tests.
Made by compiler.dev. Free tools · Pricing