Skip to content
compiler.dev

How to Find, Fix and Contain Flaky Tests in CI

A flaky test passes and fails on the same code. Each one trains developers to click "re-run" and ignore red builds, which is how real failures get merged. Flakiness also costs money, since every rerun is billed. This guide covers how to measure it, how to contain it so it stops blocking people, and how to fix the usual causes.

How common is it

Google's testing team reported in 2016 that almost 16 percent of their tests had some level of flakiness, and that about 1.5 percent of all test runs reported a flaky result. Their engineering blog post "Flaky Tests at Google and How We Mitigate Them" has the details. Your numbers will differ, but the lesson holds: in a large suite, some flakiness is the default, and you need a process, not heroics.

The cost is easy to compute. If 10 percent of PR runs fail for flaky reasons and each triggers one rerun of a 15-minute, 2-core Linux job, that is about 1.5 extra minutes per PR on average, billed at $0.006 a minute, plus the developer's wait. The money is small; the lost trust and context switching are the real cost. Use the cost of slow CI calculator to put a number on the wait.

Step 1: Measure the flake rate

You cannot manage what you do not count. Define a flaky failure as: a test that failed and then passed on a retry of the same commit. Track:

  • Flake rate per test (flaky failures divided by runs).
  • Share of workflow runs that needed a rerun.
  • Time lost to reruns.

Get the data from test reports. Upload JUnit XML for each run and store results in a simple table keyed by test name and commit SHA. If you have no tooling yet, a weekly script over gh run list that finds runs with run_attempt > 1 that eventually succeeded gives a first estimate:

gh run list --workflow ci.yml --limit 200 --json databaseId,attempt,conclusion \
  --jq '[.[] | select(.attempt > 1 and .conclusion == "success")] | length'

Step 2: Detect flakes on purpose

Instead of waiting for random failures, rerun suspect tests many times in a scheduled job:

on:
  schedule:
    - cron: "0 3 * * *"
jobs:
  flake-hunt:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: go test ./... -count=20 -shuffle=on

Equivalents: pytest --count=20 with pytest-repeat, pytest -p randomly for order shuffling, jest --runInBand plus repeated runs, and Playwright's --repeat-each=10. Shuffling test order and running tests in parallel are the two fastest ways to reveal hidden dependencies.

For a pull request that adds or changes tests, run only those tests repeatedly, 20 to 50 times, before merge.

Step 3: Contain with quarantine, not blanket retries

Retries are tempting and dangerous. A global "retry the whole job 3 times" hides real intermittent bugs, such as race conditions in production code, and makes everything slower. Use these in order:

  1. Quarantine known-flaky tests. Tag them (@flaky, pytest.mark.flaky, t.Skip behind an env var) and run them in a separate non-blocking job. They still produce data, but they do not block merges.
  2. Limited retries on specific tests, with a record. Examples: pytest --reruns 2 --only-rerun ConnectionError, Playwright's retries: 2 in config (it marks tests that passed on retry as "flaky" in the report), Jest's jest.retryTimes(2).
  3. Open a ticket with an owner and a deadline for every quarantined test. A quarantine with no owner is just deletion with extra steps. Use CODEOWNERS to route tickets to the owning team.

Fail the build if the quarantine list grows past a threshold. Review it weekly.

Step 4: Fix the usual causes

Most flakes fall into a short list.

CauseSymptomFix
TimeFails near midnight, DST, month endInject a clock; freeze time in tests
Async waitssleep(1) or timing assertionsWait on a condition or event, with a timeout
Shared statePasses alone, fails in suiteIsolate data per test; reset databases and globals
Test orderFails when shuffledRemove dependencies between tests
Ports and files"address already in use"Use ephemeral ports and temp directories
External servicesNetwork errorsUse fakes or recorded responses; keep real-service tests in a separate job
Resource limitsPasses locally, fails on 2-core CILower parallelism, or use a larger runner
RandomnessSeed-dependentLog the seed, fix it on failure for reproduction
UI timingElement not yet visibleUse the framework's auto-waiting and web-first assertions

Log the seed and the test order for every run, so a failure can be reproduced. Keep artifacts like screenshots, traces and server logs for failed tests (if: failure()).

Step 5: Reproduce

To reproduce a flaky test in CI conditions, run it in a loop on the same runner type, in the same container:

for i in $(seq 1 100); do pytest tests/test_checkout.py -x -q || break; done

If it only fails under load, run it alongside CPU stress, or restrict the container: docker run --cpus=1 --memory=2g. Many CI-only flakes are caused by the smaller machine, and a larger runner removes them without fixing the cause. That can be a valid short-term move, but write the real fix down.

Prevent new flakes

  • Run new tests 20 times in the PR that adds them.
  • Run the full suite shuffled in a nightly job.
  • Review test code for sleeps, shared state and real network calls.
  • Treat a flaky test as a bug with the same priority as a failing one, because to a developer it is.

If you shard tests, see how to do it without hiding order dependence in test sharding and parallel tests.

FAQ

Should I auto-retry failed tests in CI?

Only for known, narrow causes, and record every retry so passes-on-retry show up in a report. Blanket retries hide real problems.

Is a flaky test worse than no test?

Often yes, if people ignore it. A quarantined test that is tracked beats a blocking test that everyone reruns.

How do I handle flaky end-to-end tests?

Keep them few, test through the API where possible, use auto-waiting locators, isolate test data and keep traces for failures. Run them after cheaper checks pass.

Do flaky tests affect DORA metrics?

Yes: reruns lengthen lead time and a high rerun rate hides your true change failure rate. See DORA metrics.

Measure the effect

After a cleanup, compare rerun rates and median PR time before and after. If you also want to know whether a failure is a CPU-starvation flake, compiler.dev's comparison mode runs the same job on a larger machine, so you can see if the failure disappears when the resources do.

Made by compiler.dev. Free tools · Pricing