Slow E2E Tests: Why They Drag and How to Fix Them
Slow e2e tests are the most common reason a pull request takes 30 minutes to go green. If your e2e tests are slow, so are your slow integration tests and probably your slow tests in general: the causes overlap. This guide is the hub for the topic. It shows how to find where the time goes, then points to the fix for each framework.
Why are e2e tests slow?
End-to-end tests drive a real browser against a real app, so every test pays for work that a unit test never does. Time typically goes to:
- Starting things: the app server, the database, the browser, and logging in. Done per test instead of once, this dominates.
- Fixed waits:
sleep(5000)orwaitForTimeoutwhere an event would do. Each one is paid on every run, even when the page was ready in 200 ms. - Serial execution: one worker on a 2-vCPU runner, so tests queue behind each other.
- Retries: flaky tests rerun, and a retry doubles that test's cost. See flaky e2e tests.
- Too many e2e tests: checks that a unit or integration test could do, run through the UI. See unit tests vs e2e tests.
- Cold CI: browser downloads, dependency installs and image pulls on every run.
Step 1: measure per test, not per run
Every framework can print per-test durations. Sort descending and look at the top 20: usually 20 percent of tests account for 60 percent of the time.
npx playwright test --reporter=list,json # durations in the JSON output
npx jest --verbose # per-test times
npx vitest run --reporter=verbose
pytest --durations=20
Also record the non-test time: checkout, install, browser install, server start. If that is more than 20 percent of the job, fix setup first with caching before touching the tests.
Step 2: remove fixed waits and shared setup cost
Replace sleeps with condition-based waits (auto-waiting locators, expect(...).toBeVisible(), network idle on a specific request). Log in once and reuse the session instead of through the UI in every test. Seed data through an API or the database, not by clicking. Details for each tool:
- Playwright tests slow:
storageState, workers,--shard - Cypress tests slow:
cy.session, parallel runs - Pytest slow tests: markers, xdist, splitting
Step 3: run in parallel
Two kinds of parallelism, used together:
- Workers inside one job. Limited by the cores on the runner. A 2-vCPU machine gives you two or three browser workers at most.
- Shards across jobs. Each shard gets its own runner, so wall-clock time falls roughly by the shard count. The test sharding guide has matrix examples, and the test sharding planner estimates shard count against your target time and cost.
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3, 4]
steps:
- run: npx playwright test --shard=${{ matrix.shard }}/4
Parallel tests need isolation: separate data per worker, no shared mutable state, unique accounts or tenants. If parallelism makes tests fail, the tests had hidden dependencies, which is worth knowing.
Step 4: slow unit and integration tests
Not every slow suite is browser-driven. For slow integration and unit tests in JavaScript and Python:
- Jest slow: transforms, workers,
--shard - Vitest slow: isolation, pools, import time
- Pytest slow tests:
-n auto,--durations, markers
Integration tests are often slow because each test creates a database schema. Create it once per worker, wrap each test in a transaction and roll back, and truncate between files only when you must.
Step 5: give the work more CPU
Browser tests are CPU-hungry. If workers are queueing behind a 2-core machine, a larger runner can be the cheapest fix, since per-minute price rises roughly with cores while time falls. The faster GitHub Actions runners guide compares options. compiler.dev is one of them: it was faster than GitHub's 2-core runner on 15 of 19 benchmarked stacks, by 1.1 to 3 times, Linux 8 vCPU is $0.009 per minute, and its comparison mode measures your own jobs before you commit. Check the result on your suite, because test suites bound by external services will not speed up.
Step 6: decide what must block the PR
Not every e2e test needs to run on every push. Run a small smoke set on PRs (5 to 10 minutes), the full suite on merge to main or on a schedule, and use a merge queue for the final gate. See the slow CI hub for the rest of the pipeline.
Quick checklist
- Per-test durations recorded; top 20 reviewed
- No fixed sleeps; login done once and reused
- Data seeded via API or database
- Browsers and dependencies cached, only needed browsers installed
- Workers match cores; shards chosen from the planner
- Retries capped at 1 or 2, with a flaky list reviewed weekly
- Smoke set on PRs, full suite after merge
FAQ
Why are my e2e tests slow in CI but fast locally?
CI runners usually have fewer cores, cold caches and no warm browser or server. Your laptop may have 8 to 16 fast cores. Run with the same worker count and measure setup time separately.
How many e2e tests should I have?
Few, covering critical user journeys. Test business rules at the unit and integration level, where tests are fast and precise. The pyramid is explained in the unit vs e2e guide.
Should I retry slow tests?
Retries hide flakiness and add time. Allow one or two retries in CI, record each retry as a flake, and fix the tests that need them.
Does sharding make slow tests cheaper?
No, it makes them faster in wall-clock time. Billed minutes go up slightly because each shard repeats setup. Cache setup so the extra cost stays small.
Made by compiler.dev. Free tools · Pricing