Hire Software Developers 7
Back to blogs

Your regression suite is mostly noise

A large grid of identical pale tiles with one vivid blue tile lifted above the rest, representing a regression suite where almost every test passes and only one finds a real breakage

Your regression suite is mostly noise

You inherit the regression suite in week two. Four thousand tests, forty minutes a run, red a couple of times a week. The engineer who wrote most of it left in the spring. When the build goes red, somebody reruns it; when it goes green, the PR merges. Nobody can tell you which tests have ever caught anything, and nobody is willing to delete one, because nobody knows which one is holding the building up.

Here's the uncomfortable part. The best-measured suites in the industry are nearly all quiet. At Google, "99% of all test executions pass" [R3], and in a one-month sample "Only 1.23% of tests ever found a breakage" [R5]. There's little reason to think your inherited regression suite has more signal per run.

And here's the part the "cut your test suite" advice leaves out. The teams that measured this and published what they could cut got to about half, not all, published a non-zero miss rate, and kept a full run of everything as a backstop.

A regression suite is the set of automated tests re-run after each change to confirm that behaviour which used to work still works. Its signal is the share of executions that report a real, change-induced defect; everything else (passes, flaky failures, and failures caused by the tests themselves) is cost.

Key takeaways

  • At Google, "99% of all test executions pass", and in a one-month sample "Only 1.23% of tests ever found a breakage" [R3] [R5]. Across 5,562,881 test targets studied over one month in 2016, 91.3% passed at least once and never failed once [R8] [R9].
  • In 61 open-source Java projects, 18% of test suite executions failed; 13% of those failures were flaky, and of the non-flaky failures 26% were caused by incorrect or obsolete tests rather than a bug [R11] [R12].
  • Microsoft's THEO, replayed over three products, would have skipped between 34.9% and 50.36% of test executions while letting between 0.2% and 13.4% of defects escape to a later stage [R23] [R24].
  • Facebook's deployed predictive test selection cut the infrastructure cost of testing code changes "by a factor of two" while still reporting "over 95% of individual test failures and over 99.9% of faulty changes" [R28].
  • None of the three industrial teams removed the noise for free, and neither THEO nor Facebook deleted the suite: both reduced how often tests run and kept a full run of everything as the backstop [R22] [R30].

How much of a regression suite is actually signal?

Very little. At Google, 99% of test executions pass [R3], and in a one-month sample of tests only 1.23% ever found a breakage [R5]. Of the Pass → Fail transitions in that sample, 84% came from flaky tests [R4].

Those three numbers come from John Micco's deck "The State of Continuous Integration Testing @Google", which is in Google's own public research archive [R1]. Keep the units attached, because each one counts something different. The 99% is test executions, out of roughly 150 million a day across 4.2 million tests [R2] [R3]. The 1.23% is tests, over one month [R5]. The 84% is transitions from pass to fail [R4]. Three denominators, one direction.

The peer-reviewed version is sharper. Memon and colleagues, including Micco, studied over 500,000 changelists from 11 February to 11 March 2016, touching more than 5.5 million test targets and over 4 billion test outcomes [R8]. Their Table II: 91.3% of targets passed at least once and never failed even once. Only 2.07% ever both passed and failed. After filtering out the flaky ones, 1.23% of targets "actually found a test breakage (or a code fix) being introduced by a developer" [R9]. Put in their own words: "more than 99% of all tests run by the CI system pass or flake" [R10].

Now the detail that makes this worse, not better. Those tests were already selected. Google's postsubmit system only runs a test when a changed file sits in "the transitive closure of the test dependencies", which is regression test selection by definition [R6]. The 99% pass rate is what's left after the obvious pruning. Dependency-based selection removes the tests that can't be affected; it does nothing about the ones that can be affected and never fail.

You are not Google. Your suite is smaller and your change rate lower. None of that argues your suite has more signal per execution, and the next study is much closer to your situation.

When the build goes red, a quarter of real failures are the suite's own fault

Labuschagne, Inozemtseva and Holmes (Waterloo and UBC) did what the repository-mining studies before them couldn't: they looked at real failures and worked out what each one had actually caught. Their ESEC/FSE 2017 paper covers 61 open-source Java projects on Travis CI and 106,738 build executions [R11].

The headline, verbatim: "18% of test suite executions fail and ... 13% of these failures are flaky. Of the non-flaky failures, only 74% were caused by a bug in the system under test; the remaining 26% were due to incorrect or obsolete tests." [R12]

Take the parts one at a time, with denominators.

18.4% of builds failed: 19,640 of 106,738. Add errors and a third of builds, 33%, were in a state other than pass. The authors' framing: at one build a day, "their products would spend more than 17 weeks in non-passing states" [R13].

12.8% of failures were flaky. They took 935 failures that sat between two passing builds, reran each failed build three times, and 120 came back inconsistent [R14]. Three reruns only catches the flakes that show up in three tries, and the authors call it "only a lower bound on flakiness" [R14]. What causes flakes, and why the rerun reflex is both right and corrosive, is its own subject; I've covered it in Flaky tests: what actually causes them, and I won't repeat it here.

Of the 815 real failures, 25.9% were the suite breaking itself. The method is simple and worth stealing. When a failing build was fixed by changing only production code, the test was right and caught a bug. When it was fixed by changing only test code, the test was the problem. Mixed fixes were split by lines changed. Result: 74.1% of deterministic failures detected real faults, 25.9% were test maintenance [R15]. The authors call their approach "a reasonable lower bound on test maintenance" [R15], so the real share could be higher.

So a red build in that dataset had three possible meanings, and "you broke something" was only one of them.

It gets worse at the project level. Of the projects they could plot, "the tests of nine projects require maintenance more often than they find bugs. It is therefore possible that these test suites add very little value or even represent a loss for the projects." [R16] That's the paper's language, not mine.

Meanwhile the suites kept growing. Test code in these projects grew 267.8% over the study period against 59.1% for production code, from 10.8% of all code to 21.9% [R20]. Every one of those lines has to be kept true as the product changes.

The maintenance failures weren't exotic, either. One test depended on MySQL being case-insensitive and broke when run on Postgres. Another broke because someone removed the context reset between tests [R19]. If you inherited your suite, you inherited assumptions like those, and the person who knew about them may be the one who left. That's key-person risk wearing a test-suite costume.

Inside a failing build, 99.6% of test executions still passed

This is the number that should reset your intuition about what a red build is telling you. In the failing builds where the authors could parse individual test results (40 projects, 586 failures), the average test case failure rate was 0.38% [R17]. On average, 99.62% of the tests in a red build were green.

Be exact about the denominator, because the paper is: "This is not the global test case failure rate, but the test case failure rate within the builds that had at least one test failure." [R17] Across all builds, passing ones included, it is necessarily lower.

And 64% of failed builds had more than one failing test, which "usually have the same root cause" [R17]. One bug, several alarms.

From that, the paper derives a bound: "a perfect reduction technique could reduce the number of test executions by over 99% and still detect the same number of faults" [R18]. Read the qualifier. The abstract says it plainly: "over 99% of the test case executions could have been eliminated with a perfect oracle" [R18]. Nobody has a perfect oracle. You only know which 0.38% would fail after you've run the other 99.62%.

That gap, between "almost everything is noise" and "you can't know in advance which part", is the entire problem. The industrial teams that took it on tell you how big the real saving is.

What industry actually cut: about half, never for free

Three companies with very large test estates have published on this. The two that reported a saving both reported what it cost, and only one of them published a deployed result.

Microsoft THEO (ICSE 2015)Facebook predictive test selection (2019)Google (ICSE-SEIP 2017)
What was reducedTest executionsInfrastructure cost of testing code changesNothing reported
How much34.9% to 50.36% of executions [R24]"a factor of two" in cost; executions by "a factor of three" [R29]Stated as a plan [R10]
What it cost0.2% to 13.4% of defects caught later [R24]< 5% of individual test failures and < 0.1% of faulty changes not reported [R29]—
Evidence typeSimulation, replaying history [R23]Deployed in production [R28]Analysis

Microsoft. Herzig, Greiler, Czerwonka and Murphy built THEO, a cost model that "dynamically skips tests when the expected cost of running the test exceeds the expected cost of removing it" [R21]. They replayed it over Windows, Office and Dynamics: more than 26 months of development and more than 37 million test executions [R23]. Per product, from Table III [R24]:

  • Windows: 40.58% of test executions skipped, 0.20% of defects escaped.
  • Office: 34.9% skipped, 8.7% escaped (a three-month period on one branch).
  • Dynamics: 50.36% skipped, 13.40% escaped.

The abstract rounds that to "a reduction of 50% of test executions ... while maintaining product quality" [R21]. The conclusion is more careful: "Removing tests would result in between 0.2% and 13% of defects being caught later in the development process, thus increasing the cost of fixing those defects" [R26]. The net was still positive, "up to $2 million per development year, per product" [R26], and about a third of test-result inspections on Windows and Dynamics turned out to be unnecessary false alarms [R27].

Facebook. Machalica, Samylkin, Porth and Chandra learned a selection strategy "from a large dataset of historical test outcomes using basic machine learning techniques" [R28]. Deployed, it "reduces the total infrastructure cost of testing code changes by a factor of two, while guaranteeing that over 95% of individual test failures and over 99.9% of faulty changes are still reported back to developers" [R28]. The flip side is in the body: "we fail to report only < 5% of individual test failures and < 0.1% of faulty changes" [R29]. The miss rate is a calibrated setting, chosen on purpose.

Google measured the noise in more depth than anyone, and framed the reduction as a goal: to "schedule fewer tests while retaining high probability of detecting real faults" [R10]. No reduction percentage appears in either Google document I read [R1] [R8].

So the pattern: all three found the suite overwhelmingly noise, and the two that published a saving cut about half and accepted a measured miss rate to do it. If someone quotes you a bigger number with a zero next to it, ask for the paper.

The part everyone leaves out: they kept a full run

This is the detail that changes what you do on Monday. Neither THEO nor Facebook deleted tests. Both reduced how often tests run, and both kept a full run as a backstop.

THEO was built that way from the start: "We designed THEO to ensure that all tests will execute on all code changes at least once before shipping the software product ... THEO does not sacrifice product quality but may delay defect detection to later development phases." [R22] That's why its "escaped" defects weren't lost. On Windows, 71% of them were caught one branch later, and "none of the defects escaped into the trunk branch" [R25]. On Dynamics, 97% were caught on the next merge branch [R25].

Facebook's selection decides what runs before a change lands. Behind it: "Once every few hours, all tests are exercised on the most recent version of master branch", and "any faulty change that makes it into the master branch will be detected in the stabilization stage" [R30].

The saving came from running the quiet tests less often, not from deleting them. That's my reading across the two papers rather than a claim either makes about the other, but both describe the same architecture: a selective run on the path a change takes into the main line, and a complete run before anything ships.

That changes the question you ask about a test that has never failed. Not "can I delete this?" but "does this need to run on every pull request?"

The case against cutting anything

The strongest argument for leaving your suite alone comes from the Microsoft paper itself. Large-system tests "rarely find defects", and yet "they act as an insurance process" [R21]. A 1.23% hit rate is exactly what insurance looks like [R5]. You don't cancel fire cover because the house didn't burn down this year.

That argument is right about the suite and wrong about the schedule. Insurance priced by the minute should be bought where it pays. THEO didn't drop the cover; it moved it later in the pipeline and measured what that cost [R22] [R24].

The second argument is scale, and it's a fair one. Every industrial result here is from Google, Microsoft or Facebook, where the savings are counted in millions of dollars a year per product [R26] and a model-training pipeline is a rounding error. If your whole suite runs in eight minutes and your CI bill is a line item nobody reads, predictive selection is not your problem. The noise still is, because the cost that matters at your size isn't machines. It's the engineer reading a red build that means nothing, and learning to stop reading.

The third argument is the one that should actually slow you down. Pruning a suite using its own history only works if the history is honest. Facebook's team tried training on data where flaky failures hadn't been separated out. At one setting, the model looked like it caught about 90% of failures and actually caught about 70%: "Had we deployed such a model in production, we would fail to report three times as many test failures as expected from the evaluation." [R32] In Facebook's sampled full runs, flaky failures outnumbered real ones about four to one [R31]; at Google, 2 to 16% of compute went on re-running flakes [R7]. If you rank your inherited tests by "has it ever failed?" without separating flakes first, you'll keep the liars and demote the tests that were right.

And one honest limit on the open-source evidence: Labuschagne's classification is a heuristic based on which files a fix touched, on open-source Java projects selected for their 2012–2014 activity [R11] [R15]. It is the best measurement of test maintenance cost I found. It is not a law.

What to do with the regression suite you inherited

None of the published distributions is about your suite, so measure yours. Everything below uses data your CI already has.

  1. Pull the history. Export every test result you can get, per test, for the last three to six months. Count failures per test. Expect a long list of tests that have never failed once: at Google it was 91.3% of targets over a month [R9].
  2. Separate flakes before you judge anything. Rerun each historical failure against its original commit a few times. Anything inconsistent is a flake, not a verdict on the test's value, and it has to come out of the data before step 4 or you repeat Facebook's three-times error [R32].
  3. Classify the real failures the way Labuschagne did. For each one, look at the commit that turned the build green again. Production code only: the test caught a bug. Test code only: the test was the cost. Tally per test [R15]. A test that has only ever produced maintenance is a candidate for deletion; a test that has caught a real bug stays, however rarely it fires.
  4. Split the schedule, not the suite. Tests with a record of catching bugs, plus anything guarding the flows your revenue depends on, run on every pull request. The tests that have never failed move to a nightly or pre-release full run. That's THEO's and Facebook's architecture at a size that needs a CI configuration change, not a machine-learning model [R22] [R30].
  5. Measure your own escape rate. Every time the full run catches something the pull-request run missed, log it. That's your version of THEO's 0.2% to 13% [R24], and it tells you whether the split is too aggressive. Production is the final backstop, and you should be watching it anyway.

Two things not to do. Don't use code coverage to decide which tests matter; coverage as a target is its own failure mode, argued in the code audit checklist. And before you spend a month stabilising an automated test, ask whether the check belongs in automation at all. That's a different decision, and it's worth making first.

The suite nobody trusts

A suite nobody trusts gets rerun until it's green, then ignored, and meanwhile nobody is checking whether the release works. QA On Demand works inside the suite you already have, same framework and conventions, and leads with hands-on functional and regression testing, so a person is verifying your riskiest flows while the suite gets measured, split and repaired. Automation gets added as the product stabilises, and every suite and fixture lands in your repo, owned by you. It's $3,495/mo for a single stream on a 3-day task cycle, cancel any time.

Before any of that, run one query: which of your tests has failed, for a reason that turned out to be a real bug, in the last six months? The length of that list is how much of your suite you actually know.

Frequently asked questions

What percentage of regression tests actually find bugs?

Very few, in every dataset cited here. At Google, "99% of all test executions pass" and in a one-month sample "Only 1.23% of tests ever found a breakage" [R3] [R5]. In the peer-reviewed Google study of 5,562,881 test targets over one month in 2016, 91.3% passed at least once and never failed, and after removing flaky targets 1.23% found a breakage or a fix [R9]. In 61 open-source Java projects, only 0.38% of test case executions failed even inside builds that failed [R17].

Can you delete most of a regression test suite without missing bugs?

No published industrial result supports that. Microsoft's THEO, replayed over Windows, Office and Dynamics, would have skipped 34.9% to 50.36% of test executions with 0.2% to 13.4% of defects escaping to a later stage [R24]. Facebook's deployed system halved infrastructure cost while missing under 5% of individual test failures and under 0.1% of faulty changes [R28] [R29]. Both kept a full run of every test as a backstop [R22] [R30]. The "over 99%" reduction figure in the literature assumes "a perfect oracle" and is a theoretical bound, not a result [R18].

What is regression test selection?

Regression test selection is choosing which tests to run for a given change instead of running the whole suite. The basic form runs only tests that depend on the changed files; Google's postsubmit system runs a test only when a changed file is "present in the transitive closure of the test dependencies" [R6]. Predictive selection goes further and uses historical outcomes to guess which affected tests are likely to fail; Facebook's version selects fewer than a third of the dependency-selected tests and cut test executions by a factor of three [R29].

Why do regression tests fail when nothing is broken?

Two reasons dominate the evidence. Flaky tests: in 61 Java projects, 12.8% of failures that sat between two passing builds were inconsistent on rerun, and that is a lower bound [R14]; at Google, 84% of Pass → Fail transitions came from flaky tests [R4]. And the tests themselves: of the non-flaky failures, 25.9% were classed as test maintenance, resolved by changing the test rather than the product, meaning the test was what was wrong [R15]. Causes included tests that depended on database behaviour and on other tests' state [R19].

Should I move slow tests out of the pull-request pipeline?

Move the ones with no record of catching anything, but keep running them somewhere. That is the design both published industrial systems use: THEO guarantees every test runs on every change at least once before shipping [R22], and Facebook runs all tests on master every few hours behind its selective pre-land stage [R30]. Separate flaky failures from the history first: in Facebook's experiment, a model trained on un-deflaked data would have missed three times as many failures as its evaluation predicted [R32].

Sources

back to top

Related Articles

Book 30 min with Albert
Smiling man with short dark hair and glasses wearing a black suit, white shirt, and black tie against blue background.
Tell Albert what you're shipping.
He'll read this before joining the call. Phone number comes next, on the calendar step.
↳ info@you-source.com
↳ 4-hour response
Please wait while we retrieve meeting schedules.
Oops! There's a problem with your request. We're working on fixing it. Please try again later.