Unit Tests vs E2E Tests: Speed, Cost and the Pyramid
The difference between unit tests and e2e tests is scope, and scope drives speed, cost and flakiness. A unit test checks one function or class in memory. An end-to-end test drives the whole system the way a user would. Teams get into trouble when they use e2e tests for things unit tests do better, and end up with a slow, flaky suite. This guide compares unit tests vs e2e tests on speed and cost, shows the test pyramid, and gives a way to choose. The CI-speed consequences are covered in the slow e2e tests guide.
The comparison
| Unit | Integration | End-to-end | |
|---|---|---|---|
| Scope | One function or class | Several components, often a real database | Whole app through the UI or public API |
| Typical time per test | Milliseconds | Tens to hundreds of milliseconds | Seconds to tens of seconds |
| Setup | None or mocks | Database, services | Browser, app, database, seed data |
| Failure points | Few | Some | Many (every layer) |
| Flakiness | Rare | Occasional | Common (flaky e2e tests) |
| Failure diagnosis | Points at the line | Points at a component | Points at a journey, then you dig |
| Confidence it works together | Low | Medium | High |
The times are typical orders of magnitude, not benchmarks: your suite will vary. The pattern is consistent, though: each step up the ladder costs roughly 10 to 100 times more per test, in time and in maintenance.
The test pyramid
The test pyramid says to have many unit tests, fewer integration tests, and a small number of e2e tests:
/\ few e2e tests: critical user journeys
/ \
/----\ some integration tests: DB, API, services
/ \
/--------\ many unit tests: logic, edge cases
The reason is cost per unit of confidence. A pricing rule with 20 edge cases can be tested in 20 unit tests that run in 50 ms in total. The same 20 cases through a browser take minutes and fail for reasons unrelated to pricing. Use e2e tests to prove the pieces are wired together (sign up, pay, see the result), not to enumerate cases.
Some teams prefer a "trophy" or "honeycomb" model with more integration tests, especially for service-heavy back ends. The principle is the same: put most checks where they are fast and precise.
What each should test
- Unit: business rules, parsing, formatting, calculations, state machines, error handling and edge cases.
- Integration: database queries and migrations, API handlers with a real database, message handling, authentication and permission rules.
- E2E: a handful of critical journeys: login, checkout, the core workflow your product exists for, plus a smoke test that the deployed app starts.
Avoid e2e tests for validation messages, formatting, permissions matrices and anything with many combinations.
The cost angle: what a PR actually pays
Take a service where a PR runs 2,000 unit tests, 300 integration tests and 120 e2e tests, on a runner billed per minute. Illustrative figures, not measured data:
- Unit: 2,000 tests at 5 ms is about 10 s.
- Integration: 300 tests at 150 ms is about 45 s.
- E2E: 120 tests at 15 s is 30 minutes serial, or about 8 minutes across 4 shards (setup time included).
The e2e layer is over 90 percent of the runtime from 4 percent of the tests. In dollars, using GitHub's Linux standard rate of $0.006 per minute (pricing): 32 serial minutes is about $0.19 per run, and 4 shards at 8 minutes each is 32 minutes billed. The money is small per run; the cost that matters is developer waiting time, which the cost of slow CI calculator estimates: 30 minutes of waiting on every PR for a team of 10 adds up quickly.
Moving 40 of those e2e checks down to integration tests would cut about 10 minutes of serial time at the cost of a few hours of work once. That trade is nearly always worth it.
How to keep CI fast with all three
- Order by speed. Run unit tests first and fail fast; start e2e only if they pass, or run them in parallel jobs if you want the lowest wall-clock time.
- Run the right tests per event. Smoke e2e and unit on every PR; full e2e on merge to main or nightly; see merge queues.
- Parallelise the slow layer. Shard e2e tests; the planner picks the count.
- Cache dependencies and browsers (caching guide).
- Give e2e enough CPU. Browser tests suffer on 2 vCPUs. Faster GitHub Actions runners compares options; compiler.dev was faster than GitHub's 2-core runner on 15 of 19 benchmarked stacks, by 1.1 to 3 times, and its comparison mode measures your own jobs.
- Keep test counts honest. When an e2e test fails for the third time for non-product reasons, push its checks down or delete it.
Framework guides: Jest, Vitest, pytest, Playwright, Cypress. The whole pipeline is in the slow CI hub.
A rule of thumb for choosing
Ask: what is the lowest level at which this bug could be caught? If a unit test can catch it, write a unit test. If it needs a real database, write an integration test. If it can only appear when the browser, network and server meet, write an e2e test.
FAQ
What is the difference between unit tests and e2e tests?
Unit tests check small pieces in isolation, quickly. E2E tests check the whole system through the real interface, slowly. You need both, in different proportions.
Are e2e tests better than unit tests?
Neither is better. E2E tests give more confidence per test but cost much more in time and maintenance, so use few of them for the most important journeys.
What is the test pyramid?
A model with many fast unit tests at the base, fewer integration tests in the middle and a small number of e2e tests at the top.
How many e2e tests should I have?
Enough to cover your critical journeys, often tens rather than hundreds. If your e2e run exceeds your PR time budget, move checks down a level.
Made by compiler.dev. Free tools · Pricing