How to Find, Fix and Contain Flaky Tests in CI
A flaky test passes and fails on the same code. Each one trains developers to click "re-run" and ignore red builds, which is how real failures get merged. Flakiness also costs money, since every rerun is billed. This guide covers how to measure it, how to contain it so it stops blocking people, and how to fix the usual causes.
How common is it
Google's testing team reported in 2016 that almost 16 percent of their tests had some level of flakiness, and that about 1.5 percent of all test runs reported a flaky result. Their engineering blog post "Flaky Tests at Google and How We Mitigate Them" has the details. Your numbers will differ, but the lesson holds: in a large suite, some flakiness is the default, and you need a process, not heroics.
The cost is easy to compute. If 10 percent of PR runs fail for flaky reasons and each triggers one rerun of a 15-minute, 2-core Linux job, that is about 1.5 extra minutes per PR on average, billed at $0.006 a minute, plus the developer's wait. The money is small; the lost trust and context switching are the real cost. Use the cost of slow CI calculator to put a number on the wait.
Step 1: Measure the flake rate
You cannot manage what you do not count. Define a flaky failure as: a test that failed and then passed on a retry of the same commit. Track:
- Flake rate per test (flaky failures divided by runs).
- Share of workflow runs that needed a rerun.
- Time lost to reruns.
Get the data from test reports. Upload JUnit XML for each run and store results in a simple table keyed by test name and commit SHA. If you have no tooling yet, a weekly script over gh run list that finds runs with run_attempt > 1 that eventually succeeded gives a first estimate:
gh run list --workflow ci.yml --limit 200 --json databaseId,attempt,conclusion \
--jq '[.[] | select(.attempt > 1 and .conclusion == "success")] | length'
Step 2: Detect flakes on purpose
Instead of waiting for random failures, rerun suspect tests many times in a scheduled job:
on:
schedule:
- cron: "0 3 * * *"
jobs:
flake-hunt:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: go test ./... -count=20 -shuffle=on
Equivalents: pytest --count=20 with pytest-repeat, pytest -p randomly for order shuffling, jest --runInBand plus repeated runs, and Playwright's --repeat-each=10. Shuffling test order and running tests in parallel are the two fastest ways to reveal hidden dependencies.
For a pull request that adds or changes tests, run only those tests repeatedly, 20 to 50 times, before merge.
Step 3: Contain with quarantine, not blanket retries
Retries are tempting and dangerous. A global "retry the whole job 3 times" hides real intermittent bugs, such as race conditions in production code, and makes everything slower. Use these in order:
- Quarantine known-flaky tests. Tag them (
@flaky,pytest.mark.flaky,t.Skipbehind an env var) and run them in a separate non-blocking job. They still produce data, but they do not block merges. - Limited retries on specific tests, with a record. Examples:
pytest --reruns 2 --only-rerun ConnectionError, Playwright'sretries: 2in config (it marks tests that passed on retry as "flaky" in the report), Jest'sjest.retryTimes(2). - Open a ticket with an owner and a deadline for every quarantined test. A quarantine with no owner is just deletion with extra steps. Use CODEOWNERS to route tickets to the owning team.
Fail the build if the quarantine list grows past a threshold. Review it weekly.
Step 4: Fix the usual causes
Most flakes fall into a short list.
| Cause | Symptom | Fix |
|---|---|---|
| Time | Fails near midnight, DST, month end | Inject a clock; freeze time in tests |
| Async waits | sleep(1) or timing assertions | Wait on a condition or event, with a timeout |
| Shared state | Passes alone, fails in suite | Isolate data per test; reset databases and globals |
| Test order | Fails when shuffled | Remove dependencies between tests |
| Ports and files | "address already in use" | Use ephemeral ports and temp directories |
| External services | Network errors | Use fakes or recorded responses; keep real-service tests in a separate job |
| Resource limits | Passes locally, fails on 2-core CI | Lower parallelism, or use a larger runner |
| Randomness | Seed-dependent | Log the seed, fix it on failure for reproduction |
| UI timing | Element not yet visible | Use the framework's auto-waiting and web-first assertions |
Log the seed and the test order for every run, so a failure can be reproduced. Keep artifacts like screenshots, traces and server logs for failed tests (if: failure()).
Step 5: Reproduce
To reproduce a flaky test in CI conditions, run it in a loop on the same runner type, in the same container:
for i in $(seq 1 100); do pytest tests/test_checkout.py -x -q || break; done
If it only fails under load, run it alongside CPU stress, or restrict the container: docker run --cpus=1 --memory=2g. Many CI-only flakes are caused by the smaller machine, and a larger runner removes them without fixing the cause. That can be a valid short-term move, but write the real fix down.
Prevent new flakes
- Run new tests 20 times in the PR that adds them.
- Run the full suite shuffled in a nightly job.
- Review test code for sleeps, shared state and real network calls.
- Treat a flaky test as a bug with the same priority as a failing one, because to a developer it is.
If you shard tests, see how to do it without hiding order dependence in test sharding and parallel tests.
FAQ
Should I auto-retry failed tests in CI?
Only for known, narrow causes, and record every retry so passes-on-retry show up in a report. Blanket retries hide real problems.
Is a flaky test worse than no test?
Often yes, if people ignore it. A quarantined test that is tracked beats a blocking test that everyone reruns.
How do I handle flaky end-to-end tests?
Keep them few, test through the API where possible, use auto-waiting locators, isolate test data and keep traces for failures. Run them after cheaper checks pass.
Do flaky tests affect DORA metrics?
Yes: reruns lengthen lead time and a high rerun rate hides your true change failure rate. See DORA metrics.
Measure the effect
After a cleanup, compare rerun rates and median PR time before and after. If you also want to know whether a failure is a CPU-starvation flake, compiler.dev's comparison mode runs the same job on a larger machine, so you can see if the failure disappears when the resources do.
Made by compiler.dev. Free tools · Pricing