Skip to content
compiler.dev

Flaky E2E Tests: Causes and Fixes That Last

Flaky e2e tests pass and fail on the same code, and flaky end-to-end tests are worse than flaky unit tests because they take minutes to run and touch everything: browser, network, server, database and test data. This guide is specific to browser and system-level suites. It covers the four causes that explain most failures (waits, isolation, database state and environment), then how to use retries and quarantine without hiding real bugs. For the general method of measuring and tracking flakes, read the flaky tests guide; for speed, see slow e2e tests.

Why flaky e2e tests also make CI slow

A test that fails 5 percent of the time forces a rerun. If the suite takes 20 minutes and you have 30 such tests, a large share of PRs need at least one rerun, so the effective CI time is closer to 30 minutes plus the wait for someone to notice. Flakiness is a speed problem as well as a trust problem. See the slow CI hub.

Cause 1: timing and waits

The top cause. The test acts before the app is ready, or waits a fixed time that is sometimes too short.

  • Replace sleep and fixed timeouts with auto-waiting assertions: wait for an element to be visible and enabled, a URL to change, or a specific network response.
  • Do not assert on animations or transitions mid-flight. Disable animations in test mode.
  • Wait for the thing your next action depends on (the request that populates a table) rather than for a spinner to vanish.
  • Avoid selecting by position or text that changes. Use stable test ids or roles.
// flaky: fixed wait
await page.waitForTimeout(2000);
await page.click("#save");

// stable: wait for the condition
const save = page.getByRole("button", { name: "Save" });
await expect(save).toBeEnabled();
await save.click();
await expect(page.getByText("Saved")).toBeVisible();

CPU-starved runners make timing bugs show up: a 2-vCPU runner running several browsers can slow page loads past a timeout. If failures cluster in CI and disappear locally with fewer workers, check CPU before blaming the test.

Cause 2: test isolation

Tests that depend on order, or on a previous test's leftovers, pass or fail depending on scheduling, especially when run in parallel or in shards.

  • Each test creates the data it needs, with unique names (a random suffix or test id), and does not assume an empty system.
  • Each test logs in as its own user, or a user it owns exclusively, when it changes account state.
  • Reset browser state between tests: fresh context, cleared storage, no shared tabs.
  • Run the suite in random order and in parallel at least once a day. Hidden dependencies show up quickly.
  • Never let one test's cleanup be required for the next test's success.

Cause 3: database and shared state

Shared databases are the second biggest source after waits.

  • Parallel workers on one database collide on unique constraints and counts. Give each worker its own schema or database, or its own tenant.
  • Leftover rows from failed tests make later runs fail. Seed in a transaction and roll back, or reset the database per run in CI.
  • Assertions on global counts ("there are 3 orders") break when anything else creates an order. Assert on the records the test created.
  • Time-dependent data such as "created today" breaks at midnight UTC. Freeze or inject the clock in the app under test.
  • External services (email, payment sandboxes, third-party APIs) have rate limits and outages. Stub them at the network boundary for most tests and keep a few real contract checks outside the PR path.

Create data through the API or database, not through the UI. It is faster, and it removes UI flakiness from setup steps.

Cause 4: environment and infrastructure

  • Race between the app server and the test runner starting: wait on a health URL before running.
  • Browser, driver or OS version drift: pin versions or use a container image.
  • Memory pressure on small runners: look for browser crashes and out-of-memory kills.
  • Network flakiness in downloads: cache dependencies, mirror registries (caching guide).

Retries: use them carefully

Retries are a safety net, not a fix. Rules that work:

  1. Allow 1 to 2 retries in CI only, and none locally, so developers see real failures.
  2. Report a retried pass as a flake, not a pass. Playwright marks such tests as flaky in reports; Cypress and others have equivalents or plugins.
  3. Track flake counts per test. A test that needed retries 3 times this week needs an owner.
  4. Capture evidence on the first retry: trace, video, screenshot, console and server logs.
// playwright.config.ts
retries: process.env.CI ? 1 : 0,
use: { trace: "on-first-retry", screenshot: "only-on-failure" },

Quarantine without losing coverage

When a test is flaky and cannot be fixed today:

  1. Move it to a quarantine group that still runs but does not block merges (a tag, a separate job with continue-on-error, or a separate project).
  2. Open an issue with an owner and a deadline, and link it in the test code.
  3. Review the quarantine list weekly. Fix, delete or promote each test. A quarantine that only grows is a deleted test suite.

Do not simply mark tests as skipped: skipped tests disappear from view.

Prevent new flakes

  • Run new tests 20 to 50 times in CI before merging, to catch obvious flakiness.
  • Lint for sleep and fixed waits.
  • Keep e2e tests few and focused on critical journeys (unit tests vs e2e tests); fewer tests means fewer chances to flake.
  • Give runs enough CPU. If a runner is the bottleneck, faster GitHub Actions runners compares options; compiler.dev was faster than GitHub's 2-core runner on 15 of 19 benchmarked stacks, by 1.1 to 3 times, and its comparison mode measures your own jobs. More CPU removes timing flakes but does not cure race conditions.

Framework specifics: Playwright tests slow and Cypress tests slow.

FAQ

What causes flaky e2e tests?

Mostly bad waits, shared or leftover data, order dependence between tests, and resource-starved CI machines.

Should I retry flaky end-to-end tests?

Retry once or twice in CI, but record every retry as a flake and fix the test. Never treat retried passes as clean.

How do I find which e2e tests are flaky?

Rerun the suite on unchanged code, or analyse CI history for tests that both passed and failed on the same commit. Report retried passes in your test reports.

Is it better to delete a flaky test?

Sometimes. If it covers a low-value path, or the same behaviour is tested lower in the pyramid, deleting it is better than carrying noise.

Made by compiler.dev. Free tools · Pricing