Hire Software Developers 7
Back to blogs

Flaky tests: what actually causes them

A row of pale blocks each in its own groove, with one vivid blue block caught between two positions and doubled in orange, representing a flaky test that both passes and fails on the same code

Flaky tests: what actually causes them

The build goes red at 16:40 on a Thursday. The failure is in an integration test that touches stock reservation. Nobody opens the stack trace. Someone clicks retry, the second run comes back green, the PR merges, and the ticket closes without a comment.

Taken one incident at a time, that's the correct decision. The engineer had no evidence the failure was real, the rerun costs a fraction of a cent, and an investigation costs an afternoon. Do it two hundred times over a quarter and every instance is still individually defensible.

It's also, exactly, how flaky tests get a real defect waved through. The retry button can't tell the difference between a test that's lying and a test that's right.

A flaky test is a test that both passes and fails on the same code. That is Google's own working definition, from its internal CI deck: "Flakiness is a test that is observed to both Pass and Fail with the same code". [R59]

Key takeaways

  • Google's three flakiness figures answer three different questions: about 16% of 4.2M tests have shown some level of flakiness [R58], about 1.5% of executions report a flaky result at any given time [R60], and 46,694 of 5,562,881 targets were classified as flakes in one dataset partition [R65].
  • The study everyone quotes, Luo et al. FSE 2014, inspected 201 Apache commits but classified only 161. Every percentage in it, the famous 45% for Async Wait included, is out of 161, not 201. [R5] [R6]
  • Its strongest finding is the one nobody repeats: 126 of those 161 tests were flaky from the revision that introduced them, and the average one survived 388 days. [R11] [R13]
  • Of the 20 fixes in that study that added or changed a sleep, zero removed the flakiness. Of the 42 that used a waitFor, 23 did. [R17]
  • The distribution does not replicate: test order dependency is 12% in that Java study, 50.5% in a rerun-based one, and 59% across 876,186 Python tests. [R8] [R46] [R36]

The rerun reflex: how much flakiness is actually out there

Before asking what causes flakiness, work out how much of it there is. Scale changes what kind of problem it is.

The rerun is not a bad habit somebody picked up. At Google it is documented policy: "We re-run test failure transitions (10x) to verify flakiness" and "If we observe a pass the test was flaky" [R63]. Your engineer hitting retry is doing an unautomated version of what Google does on purpose.

Google has also published three prevalence numbers, which get quoted against each other as if one must be wrong. None is. Each counts a different object, and the unit is the entire explanation.

The first counts tests. From John Micco's deck "The State of Continuous Integration Testing @Google", 31 slides in Google's own research archive [R57], the sentence in full: "Almost 16% of our 4.2M tests have some level of flakiness." [R58] Read the qualifier, not the percentage. The unit is individual tests; the criterion is some level of flakiness, ever observed. A lifetime-ever measure across 4.2 million tests, and not the same claim as "16% of Google's tests are flaky".

The second counts targets, and it's peer-reviewed. Google's ICSE-SEIP 2017 paper "Taming Google-Scale Continuous Testing" partitions its dataset in Table II: 5,562,881 test targets, of which 5,082,803 never failed once and 115,160 both passed at least once and failed at least once. Of that last group: "Flakes 46,694 of 115,160". [R65] Against all 5,562,881 targets that is 0.84%. Against the 115,160 targets that ever produced a mixed signal, about 41%.

The third counts executions. Back in the deck: "Continual rate of 1.5% of test executions reporting a 'flaky' result" [R60]. Point in time, not lifetime-ever.

Three consistent answers to three different questions, and the unit tells you which question you're reading. A target is not a test, and the same paper uses both terms, describing milestones "as large as 4.2 million tests as selected using reverse dependencies on changed source files" while Table II partitions 5,562,881 targets [R66]. And "ever observed to be flaky at all" is not the predicate "classified as a flake in this dataset partition".

Both documents are Google's own, and both are free in Google's public research archive: no paywall, no third-party mirror [R57] [R65]. Check the wording yourself. The compressed version of the first number has travelled the industry for a decade with its qualifier stripped off.

They differ because of method, not disagreement. That is this post's argument in miniature: the same thing has happened to the cause distribution, on a much larger scale, with nobody reconciling it.

Microsoft measured a normal-sized version: across five projects in CloudBuild over one month, "4.6% of all individual test cases are flaky" [R32], but the share of builds with at least one flaky failure ranged from 17% to 52% [R33]. Flakiness concentrates by build, not by test, and that's how the reflex forms. When half your builds go red for no reason, the team learns that red means nothing.

What causes flaky tests? What the most-cited study actually measured

In 161 classified fix commits from 51 Apache projects, Async Wait accounted for 45%, Concurrency 20% and Test Order Dependency 12% [R4] [R6] [R7] [R8]. That's a fact about those 161 commits, in Java, in 2014 — not a distribution you can assume holds in your suite.

Almost every article about flaky-test causes traces back to one paper: Luo, Hariri, Eloussi and Marinov, "An Empirical Analysis of Flaky Tests", FSE 2014, out of the University of Illinois at Urbana-Champaign [R1]. It is a good paper, quoted with the denominator removed, which is how a careful 45% becomes a careless one.

Here's what it did. The authors keyword-searched the whole Apache Software Foundation commit history for "intermit" and "flak", then sampled 201 of the resulting commits for line-by-line inspection [R3]. Those 201 span 51 Apache projects, and across the 486 likely flaky-test fixes the search turned up, the heaviest contributors are HBase, ActiveMQ, Hadoop and Derby [R4] [R2].

Now the part that gets dropped. Of the 201 commits inspected, the authors could precisely classify the root cause of 161. The other 40 were "hard to classify for various reasons" and are excluded from every percentage in the paper [R5]. So the denominator behind "45% Async Wait" is 161. Divide by 201 instead and you get 37%, a number that appears nowhere in the paper.

With that established, the breakdown [R6] [R7] [R8] [R10]:

Category (the paper's own name)CountShare of the 161 classified
Async Wait7445%
Concurrency3220%
Test Order Dependency1912%
Resource Leak117%
Network106%
Time53%
IO42%
Randomness42%
Floating Point Operations32%
Unordered Collections11%

The top three account for "77% of the 161 studied commits" [R9]. One caution if you recompute from the paper: its Table 2 counts sum to 163 rather than 161. Use the prose figures — 74, 32 and 19 out of 161 [R6] [R7] [R8].

Async Wait is defined narrowly: a commit lands there "when the test execution makes an asynchronous call and does not properly wait for the result of the call to become available before using it" [R6]. Not "the test is slow". A wrong assumption about when a result is ready.

Concurrency, at 32, decomposes into the shapes of ordinary concurrency bugs: 9 data races, 10 atomicity violations, 2 deadlocks, and a subcategory the authors had to invent and named "bug in condition", at 6 [R21]. If your flakiness lives here the defect is in the code under test, and the mechanics are the ones that produce the race you only see in production.

Test Order Dependency, at 19, is the most interesting because of where it comes from: six are a static field in the code under test, three a static field in the test itself, and ten are an external dependency — a shared file or a shared network port [R20]. More than half are not in-memory state at all, so resetting your test fixtures does not catch them. That comes back in the last section. And for anyone proposing faster CI hardware as the answer: "154 out of 161 (96%) have outcome that does not depend on the platform" [R16].

78% were born flaky — and lived 388 days

The best-evidenced finding in the paper is also the least-quoted one. It's the one that should change what a CTO does on Monday.

Finding F.2, verbatim: "Most flaky tests (78%) are flaky the first time they are written." [R11] The arithmetic behind it: "From the 161 tests we categorized, 126 are flaky from the first time they were written, 23 became flaky at a later revision, and 12 others are hard to determine" [R11].

Be precise about what "the first time they are written" means, because the phrasing invites a stronger reading than the method supports. The authors established it by tracing each test's revision history back to the revision that introduced it, not by rerunning anything [R12]. The 23 that became flaky later did so because a new test violated another's isolation, or because the test code changed: bug patches, refactors, incomplete fixes to earlier flakiness [R12].

Then the second number. Across the 152 flaky tests where evolution history was available, Luo et al. report "the average number of days it takes to fix a test to be 388.46" [R13].

Put those together and the standard mental model collapses. That model is decay: a healthy suite accumulates rot as the system grows and CI gets busier, so the intervention is maintenance. A quarterly flake hunt. A dashboard. Somebody's 20% time.

What the data describes is not decay. It's intake. The defect is installed the day the test is written, it passes review because a test that passes is not something a reviewer interrogates, and then it sits in the suite for thirteen months while everybody reruns it. Your review process admitted it, one test at a time, and your rerun policy paid the interest.

Be careful how far that generalises. This is a study of flaky tests developers eventually fixed and wrote a commit message about (more on why that matters shortly). It does not say 78% of the flaky tests currently in your suite were born that way; it says that among flaky tests diagnosed and repaired across 51 large Apache projects, 78% had been flaky since the day they landed. The reading that follows — intake rather than maintenance — is mine rather than the paper's.

The remedy is cheap. A test that asserts on something asynchronous, ordered or shared is assuming something about timing or state, and that assumption either holds always or holds usually. "Holds usually" is the entire category. The review question is not "does this pass" — it passed, that is why you are looking at it. It is "what does this assume about when things happen and what else has run?" Nobody asks, because a green check mark reads as an answer. Which is why chasing a coverage number makes it worse: tests written to hit a target get written fast, against whatever the code happens to do, and that is a failure mode of its own.

Every sleep() fix in the study failed

The paper's Table 4 is the most immediately actionable thing in the flaky-test literature, and it's brutal.

For Async Wait, the largest category, the authors classified how each fix was implemented and whether it removed the flakiness or only decreased it. Of the 20 Async Wait fixes that added or modified a sleep: zero removed the flakiness, and all 20 merely decreased it. Of the 42 Async Wait fixes that added or modified a waitFor-style call, 23 removed it completely. [R17]

Twenty attempts. Twenty failures. Not "sleep is a code smell" — a measured zero.

The reason sits in the numbers beside it. "The average waiting time for waitFor calls in our cases is 13.04 seconds, while the average waiting time for sleep calls is 1.52 seconds" [R18]. A waitFor is a condition with a generous ceiling: it returns the moment the thing is true and will hang around for thirteen seconds if it has to. A sleep is a guess with a tight budget. Somebody watched the operation take 400ms on their laptop, wrote a 1.5-second sleep, saw green and shipped it. On a loaded runner with a cold cache, 1.5 seconds is not enough.

waitFor is the majority strategy here, not a niche one: 42 of the 74 Async Wait cases, 57%, were fixed that way [R6] [R17]. The correct fix isn't exotic, and people on your team already know it. The 20 sleeps are the ones written under time pressure by whoever was on the hook to get CI green before a release.

Nor is this only a 2014 finding. Google's own "Sources of Flakiness" slide — a qualitative list of factors that cause flakes, with no percentages attached — names sleep() explicitly under test case factors, alongside waits for a resource, Webdriver tests and UI tests [R64]. Two independent parties, a decade and several orders of magnitude apart, name the same culprit.

Async Wait, as practised, does not get removed. It gets reduced, until the failure rate falls under the team's annoyance threshold and the ticket closes.

So here's a diagnostic that costs nothing. Grep your test directories for Thread.sleep, time.sleep, setTimeout, cy.wait( with a number in it. Every hit is somebody encoding a guess about duration as if it were a fact about the system, and in the only study that measured what happens next, that guess was never once the thing that removed the flakiness [R17].

The distribution doesn't replicate — and the method explains why

Now the part that undercuts every listicle built on that paper. Six studies have since measured flaky-test causes in different languages and by different methods, and each produces a different top cause.

StudyPopulation and methodTop causeSecond
Luo et al. 2014, Apache/Java [R1] [R6] [R7]161 classified fix-commits, keyword-minedAsync Wait 45%Concurrency 20%
Eck et al. 2019, Mozilla [R40]200 fixed tests, classified by the devs who fixed themConcurrency, 61 of 234 labelsAsync Wait, 52
Lam et al. 2019, iDFlakies, Java [R46]422 flaky tests, found by rerunning with reorderingOrder-dependent 50.5%Non-order-dependent 49.5%
Gruber et al. 2021, Python [R36] [R37]876,186 tests, 22,352 PyPI projects, rerunTest Order Dependency 59%Test infrastructure 28%
Hashemi et al. 2022, JavaScript [R42]452 commits analysed, 358 classifiedConcurrency 20.7%Async Wait 19.6%
Schroeder et al. 2025, Rust [R43]53 fixed flaky tests, work in progressAsync Wait 33.9%Concurrency 24.5%
Berndt et al. 2026, SAP HANA [R44]559 fixed-flakiness issue reportsConcurrency 23%Timeout 16%

Follow test order dependency across that table. Luo puts it at 12% [R8]. iDFlakies, which found its flaky tests by rerunning suites in randomised orders rather than by reading commit logs, puts order-dependent tests at 50.5% of 422 [R46]. Gruber's Python study, the largest in this set at 876,186 test cases, reports that "Order dependency is a much more dominant problem in Python, causing 59 % of the 7 571 flaky tests in our dataset" [R36]. From 12% to 59%.

It gets sharper. Gruber's team went looking for Luo's top two categories and could not find them: "whereas they reported Async Wait and Concurrency to be the most common root causes of NOD flakiness, we found only little evidence of these categories" [R38]. Those two scored three cases each in their sample, while 28% of the flakiness came from test infrastructure: "a previously undocumented cause of flakiness", a category Luo's taxonomy had no slot for [R37]. Mozilla shows the same gap: four of the eleven categories Eck's developers needed were new, "not included in the taxonomy proposed by Luo et al." [R40] [R41].

None of this means anybody measured wrong. It means the method decides the answer.

Luo mined commit messages for two word stems, and the paper concedes the consequence in its own Threats to Validity: "there is no guarantee on the recall of our search. In fact, we believe our search could miss many flaky tests whose fixes could use words like 'concurrency', 'race', 'stall', 'fail', etc." [R3] Commit archaeology is a census of successful diagnoses: flakiness a developer noticed, named and repaired. Detection-based studies do the opposite: iDFlakies and Gruber's team reran suites in original and shuffled order and recorded what came out inconsistent [R46] [R36], which finds flakiness nobody noticed and nobody fixed.

The two methods fail in opposite directions. Order dependency is trivially exposed by a tool that shuffles test order and nearly invisible to a developer who always runs the suite the same way. Async Wait is the reverse: it fails with a stack trace and a timeout in it, precisely the kind of failure a developer diagnoses and writes a commit message about. Archaeology is biased towards the diagnosable, reruns towards the exposable. (That reconciliation is my reading — the studies do not claim to explain each other's results.)

Developers agree with neither. Parry's survey of 170 developers put setup and teardown first, network second and unknown reasons third: "the causes of many flaky tests go undiagnosed by developers" [R55]. When the practitioner's third-ranked answer is "I don't know", the tidy percentages are describing the diagnosable minority.

So "45% of flaky tests are Async Wait" is not a fact about software. It's a fact about 161 classified Apache commits, found by searching for two word stems, in Java, in 2014. A fine starting hypothesis for a Java codebase with heavy async integration tests, a bad thing to put in a deck about your Python service.

The case for just retrying, taken seriously

The moralising version of this post would stop here and say "quit rerunning and fix your tests". That version is wrong, and the only team that has actually costed the alternatives says so.

Leinen and colleagues (ICST 2024) instrumented a commercial project — roughly 30 developers, about a million lines of code, five years of history, combining CI logs, version control, tickets and tracked work time. Their headline: "the time spent dealing with flaky tests in the studied project represents at least 2.5% of the productive developer time", split as 1.1% investigating flaky failures, 1.3% repairing them, 0.1% building tools to monitor them [R48].

Then the sentence that belongs in every argument about this: "Contrary to most other studies, we find the cost for rerunning tests to be negligible and inexpensive. Automatically rerunning a test costs 0.02~cents, while not rerunning and thus letting the pipeline fail results in a manual investigation costing $5.67 in our context." [R49] Their response was not to lecture anyone: "The insights gained from our case study have led to the decision to shift effort from investigation and repair to automatically rerunning tests" [R50]. (Only that study's abstract was reachable, so those figures are all there is.)

Google arrives at the same position from the other end of the scale. Under the heading "Flakes are Inevitable", the deck records that the "Observed insertion rate is about the same as fix rate" [R61]. The conclusion drawn is not eradication: "Testing systems must be able to deal with a certain level of flakiness. Preferably minimizing the cost to developers." [R62] Which is why the rerun is institutionalised at 10x [R63].

If insertion and removal balance, the backlog is constant and the only lever left is cost per incident, which reruns push towards zero. Here's what that model doesn't price.

Reruns are also how real bugs get normalised. Parry's survey work names the mechanism: "developers who experience flaky tests more often may be more likely to ignore potentially genuine test failures" [R56]. The rerun is cheap per incident; the habit is not, because the habit is indiscriminate. A team that has learned red means nothing has lost the ability to act on red meaning something.

And flaky tests are finding real bugs right now. In Luo's study, 38 of the 161 commits — 24% — fixed the flakiness by changing both the test and the code under test, and 94% of the fixes that touched production code fixed a real bug in it [R14]. Roughly one flaky-test fix in five turned out to be a real product bug, in software as heavily exercised as HBase and Hadoop. The implication is stated as bluntly as a paper can: "Flaky tests should not simply be removed or disabled because they can help uncover bugs in the CUT." [R15] And the failure does not always announce itself: in Parry's survey of the literature, Zhang et al. are reported as having examined 96 order-dependent tests, two of which caused a missed alarm: a bug with no test failure at all [R47]. Second-hand, so treat it as indicative, but no rerun policy is designed to catch a test that is silently green while something is broken.

So hold both. Rerunning is the right call at the moment of failure and a corrosive policy over a quarter, and no amount of arguing resolves that into one answer. Stop treating them as one decision. Rerun automatically, because Leinen's arithmetic is not beatable on cost. But make the rerun record something: which test, which build, how many attempts, on what branch. A rerun that leaves a row in a table is a measurement; a rerun that leaves nothing is a decision delegated to whoever was watching CI that afternoon. Google's insertion rate is a known quantity only because somebody counted it [R61].

Flaky test detection: stop borrowing someone else's percentages

Which leaves one defensible move. None of those distributions is about your suite, so measure your suite.

Measuring is not archaeology. You are not reading commit messages for "flaky" and "intermittent": that is the method that produced the distribution which does not replicate [R3]. Detection means running the suite repeatedly against an unchanged commit and recording which tests come out inconsistent, then running it again in a randomised order. That is how iDFlakies found 50.5% of its 422 flaky tests to be order-dependent [R46], and the shuffled run is the cheap half that archaeology systematically misses.

Two constraints, both measured. Run it in CI, not on a developer machine: when Microsoft's team re-ran known flaky tests locally 100 times, 86% were only flaky in the CI pipeline [R34]. And do not expect certainty: "A 95 % confidence that a passing test case is not flaky on average would require 170 reruns" [R39]. That is the cost of proving a negative. Aim at finding flaky tests, not at certifying the rest as clean.

Then the finding that changes the shape of the work rather than just its target. Parry, Kapfhammer, Hilton and McMinn re-analysed an existing dataset of 10,000 test-suite runs across 24 Java projects containing 810 flaky tests, and found they do not arrive one at a time: "75% of flaky tests across all projects belong to a cluster, with a mean cluster size of 13.5 flaky tests" [R53], sharing a driver: "we identified intermittent networking issues and instabilities in external dependencies as the predominant causes of systemic flakiness" [R54].

So the standard flaky-test backlog is misconceived. A ticket per flaky test, triaged individually by whoever has capacity, is a treadmill by construction: thirteen and a half tickets that all resolve to one unstable dependency, each producing its own local workaround, and the dependency untouched at the end.

The unit of repair is the dependency boundary. That's consistent with everything else here: ten of Luo's 19 order-dependency cases were an external dependency rather than in-memory state [R20], and 28% of Gruber's Python flakiness was test infrastructure [R37]. Cluster your flaky tests by what they touch before you triage them by name, and most of the backlog collapses into a handful of boundaries: the inconsistently stubbed service, the shared database tests do not reset, the port two suites both want.

So, concretely, for a suite you have never measured:

  1. Run the full suite ten times against one unchanged commit, in CI. Record every test not consistent across all ten runs.
  2. Run it ten more times with the order randomised. Anything that only fails here is order-dependent, and those fixes stick, because all of Luo's order-dependency fixes removed the flakiness completely [R19].
  3. Group what you found by the external thing it touches, not the file it lives in [R53] [R54].
  4. Grep the same suite for hardcoded sleeps. Those are unfixed flakiness whether or not they surfaced in step 1 [R17].
  5. Add one question to review for tests touching async, ordering or shared state: what does this assume about timing, and about what else has run? That is the intake control, and 126 of the 161 studied flaky tests were in its scope [R11].

Whether a given check belongs in an automated suite at all is a separate question, worth asking before you spend a month stabilising a test that should have been manual.

The flaky backlog nobody owns

Run that exercise and you end up with a list nobody disputes is broken and nobody is assigned to. It's not feature work, so it never survives a sprint boundary, and every quarter somebody new rediscovers it. The average flaky test in the only study that tracked it lasted 388 days [R13]: not a statement about difficulty, but one about ownership.

Getting that list owned is what QA On Demand is staffed for: single stream $3,495/mo, dual stream $6,795/mo for two streams in parallel. A 3-day task cycle, daily async updates, first ship in 5 days, a replacement guarantee within 5 business days, cancel any time with no lock-in, and you own the IP and the test assets. It's manual-first by design: hands-on functional and regression testing first, automation added as your product stabilises, with the automated suites landing in your repo. That order matters. Nobody arrives on Monday and repairs your async integration tests. What gets covered first is the job your broken suite has been failing to do — a person checking whether the thing actually works.

Pick the ten tests your team reruns most often this week, run each of them ten times against an unchanged commit, and write down which ones you cannot make fail twice the same way.

Frequently asked questions

What causes flaky tests?

There is no single answer: the measured distribution changes with the language, the test types, and the method used to find the flaky tests. The most-cited study, Luo et al. FSE 2014, classified 161 flaky-test fix commits from 51 Apache projects and found Async Wait at 45%, Concurrency at 20% and Test Order Dependency at 12%, together 77% of those 161 classified commits [R6] [R7] [R8] [R9]. That distribution has not replicated. A study of 876,186 Python tests found test order dependency at 59% and test infrastructure at 28%, and reported "only little evidence" of Luo's top two categories [R36] [R37] [R38]. Mozilla developers classifying their own fixes put Concurrency first and needed four categories Luo's taxonomy did not have [R40] [R41]. What survives across all of them is the mechanism rather than the percentages: flakiness is a wrong or missing assumption about ordering, waiting, or shared state.

Are most flaky tests flaky from the day they were written?

In the one study that traced it, yes: 126 of the 161 flaky tests Luo et al. classified were flaky from the revision that introduced them, 23 became flaky at a later revision, and 12 were hard to determine, which the paper states as "Most flaky tests (78%) are flaky the first time they are written" [R11]. The mechanism is that a test that passes does not get interrogated in code review. They established the finding by tracing each test back to its introducing revision, not by rerunning anything [R12]. Read it as a result about flaky tests that were eventually diagnosed and fixed across 51 Apache projects, not as a rate for the flaky tests sitting in your suite today. The practical reading is that flakiness is an intake problem rather than a decay problem: the assumption about timing or shared state is baked in at authoring time, and then survives on average 388 days before anyone fixes it [R13].

Does adding a sleep() fix a flaky test?

No, on the only evidence available. In Luo et al.'s Table 4, of 20 Async Wait fixes that added or modified a sleep, zero removed the flakiness and all 20 merely decreased it; of the 42 fixes that used a waitFor-style call, 23 removed it completely [R17]. The reason shows up in the wait times: "the average waiting time for waitFor calls in our cases is 13.04 seconds, while the average waiting time for sleep calls is 1.52 seconds" [R18]. A sleep encodes a guess about duration; a waitFor encodes the condition you actually care about and tolerates a slow run. Google's own "Sources of Flakiness" slide names sleep() as a cause of flakes [R64].

Is it acceptable to just re-run flaky tests?

Per incident, yes, and the arithmetic is not close. Leinen et al. measured an automatic rerun at $0.0002 against $5.67 for the manual investigation it replaces, found that dealing with flaky tests consumed at least 2.5% of productive developer time, and decided to "shift effort from investigation and repair to automatically rerunning tests" [R48] [R49] [R50]. Google reaches the same conclusion at its own scale, recording that "Testing systems must be able to deal with a certain level of flakiness" and rerunning failure transitions 10x as policy [R62] [R63]. The cost the rerun does not price is behavioural, because "developers who experience flaky tests more often may be more likely to ignore potentially genuine test failures" [R56], and because flaky tests find real defects: 24% of the fixes in Luo's study changed the code under test as well as the test, and 94% of the fixes that touched production code fixed a real bug in it [R14] [R15]. Rerun, but record every rerun, so the insertion rate is a number you know rather than a habit you have.

How do I find the flaky tests in my own suite?

Detect them, don't research them. Run the full suite repeatedly against one unchanged commit and record every test that is not consistent, then run it again with the test order randomised. That reordering step is what iDFlakies used to find 50.5% of its 422 flaky tests were order-dependent, a category commit-log analysis systematically under-counts [R46]. Run it in CI rather than locally, because when Microsoft re-ran flaky tests locally 100 times, "86% of them are only flaky in the CI pipeline" [R34]. Aim at finding flaky tests rather than certifying the rest as clean, because "A 95 % confidence that a passing test case is not flaky on average would require 170 reruns" [R39]. Then group what you find by the external dependency it touches rather than by test file: 75% of flaky tests belong to clusters averaging 13.5 tests, driven predominantly by "intermittent networking issues and instabilities in external dependencies" [R53] [R54].

Sources

  • [R1] Luo, Hariri, Eloussi, Marinov, "An Empirical Analysis of Flaky Tests", FSE’14, 16–22 Nov 2014, Hong Kong; Dept of Computer Science, University of Illinois at Urbana-Champaign. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R2] The sample: "We study in detail a total of 201 commits that likely fix flaky tests in 51 open-source projects" — all from the Apache Software Foundation's central SVN repository. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R3] Selection method: keyword search on "intermit" and "flak" over all Apache commit history yielded 1,129 commits to 855 "likely about flaky tests" to 486 "likely distinct fixed flaky tests" (LDFFT) to 201 sampled for deep inspection. Threats to Validity concedes: "there is no guarantee on the recall of our search. In fact, we believe our search could miss many flaky tests whose fixes could use words like 'concurrency', 'race', 'stall', 'fail', etc." — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R4] The 51 projects are where the 486 LDFFT commits live: "At least 51 projects out of the 153 projects in Apache likely have at least one flaky test." Top four by commit volume: HBase, ActiveMQ, Hadoop, Derby. 20,654,322 LOC total. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R5] The real denominator for every percentage is 161, not 201: "We precisely classify the root causes for 161 commits, while the remaining 40 commits are hard to classify for various reasons". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R6] Async Wait is the top cause at 45%: "74 out of 161 (45%) commits are from the Async Wait category. We classify a commit into the Async Wait category when the test execution makes an asynchronous call and does not properly wait for the result of the call to become available before using it." (§3.1.1) — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R7] Concurrency is second: "32 out of 161 (20%) commits are from the Concurrency category". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R8] Test Order Dependency is third: "19 out of 161 (12%) commits are from the Test Order Dependency category". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R9] The three together: "the top three categories that represent 77% of the 161 studied commits". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R10] The tail, from Table 2 "Total" column (raw counts across the 201 inspected commits): Resource Leak 11, Network 10, Time 5, IO 4, Randomness 4, Floating Point Operations 3, Unordered Collections 1, Hard to classify 40. Note that Table 2's ten category counts sum to 163 rather than the 161 the prose states — use the prose figures. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R11] Finding F.2 reads "Most flaky tests (78%) are flaky the first time they are written." The body gives the arithmetic: "From the 161 tests we categorized, 126 are flaky from the first time they were written, 23 became flaky at a later revision, and 12 others are hard to determine". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R12] "When first written" means the test was flaky from its introducing revision onward, established by tracing the test's own revision history — not by rerunning it. The 23 later-flaky cases are attributed to new tests violating isolation, or to test-code changes (bug patches, refactors, incomplete flakiness patches). — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R13] Flaky tests survive a long time: "For the 152 flaky tests for which we have evolution information, we calculate the average number of days it takes to fix a test to be 388.46". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R14] The 94% claim, with its denominator: "38 out of 161 (24%) of the analyzed commits fix the flakiness by changing both the tests and the CUT." Finding F.12: "Some fixes to flaky tests (24%) modify the CUT, and most of these cases (94%) fix a bug in the CUT." So roughly 36 of 161 fixes (about 22%) revealed a real product bug, not 94% of fixes. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R15] The paper's own implication drawn from that, I.12: "Flaky tests should not simply be removed or disabled because they can help uncover bugs in the CUT." — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R16] Flakiness is almost never a platform problem: "we find that 154 out of 161 (96%) have outcome that does not depend on the platform." Of the 7 that do, 4 need a specific OS, 2 a specific browser, 1 a buggy JRE. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R17] A sleep does not fix flakiness, it hides it: of 20 Async Wait fixes that add or modify a sleep, Table 4 shows 0 removed the flakiness and 20 merely decreased it. Of 42 waitFor fixes, 23 (55%) removed it completely. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R18] Average waits tell the same story: "the average waiting time for waitFor calls in our cases is 13.04 seconds, while the average waiting time for sleep calls is 1.52 seconds". — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R19] Test order dependency is the one category that fixes cleanly: "All the fixes we find in our study for Test Order Dependency flaky tests completely remove the flakiness," 14 of 19 (74%) by setting up or cleaning up shared state. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R20] Where order dependency comes from: "Static field in TEST" (3 of 19), "Static field in CUT" (6 of 19), "External dependency" (10 of 19, a shared file or network port). More than half are external, so they cannot be caught by comparing in-memory state. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R21] Concurrency flakiness maps onto ordinary concurrency bugs: data races 9 of 32, atomicity violations 10 of 32, deadlocks 2 of 32, plus a new subcategory the authors name "bug in condition" 6 of 32. — https://mir.cs.illinois.edu/lamyaa/publications/fse14.pdf
  • [R32] Microsoft (Lam, Godefroid, Nath, Santhiar, Thummalapenta, ISSTA 2019, CloudBuild): "over a one-month period monitoring of five software projects, we observed that 4.6% of all individual test cases are flaky". — https://www.microsoft.com/en-us/research/wp-content/uploads/2019/11/LamETAL19RootFinder.pdf
  • [R33] Same paper, the asymmetry that matters to a CTO: "although the number of distinct flaky tests is low, the percentage of builds that include flaky test failures is substantial." Per-project share of builds containing at least one flaky failure ranged 17%–52% (ProjB 52%, ProjC 30%, ProjE 24%, ProjD 17%). — https://www.microsoft.com/en-us/research/wp-content/uploads/2019/11/LamETAL19RootFinder.pdf
  • [R34] Same paper, why "just run it locally" fails: "When we re-run flaky tests locally 100 times, we find that 86% of them are only flaky in the CI pipeline". — https://www.microsoft.com/en-us/research/wp-content/uploads/2019/11/LamETAL19RootFinder.pdf
  • [R36] Python (Gruber, Lukasczyk, Kroiß, Fraser; 22,352 PyPI projects, 876,186 test cases, 7,571 flaky tests found by rerunning): "Order dependency is a much more dominant problem in Python, causing 59 % of the 7 571 flaky tests in our dataset". — https://arxiv.org/pdf/2101.09077
  • [R37] Same study: "Another 28 % were caused by test infrastructure problems, which represent a previously undocumented cause of flakiness." Async Wait and Concurrency scored 3 cases each in its 100-test non-order-dependent sample. — https://arxiv.org/pdf/2101.09077
  • [R38] Same study, head-to-head: its Table IX prints Luo's raw counts beside its own and concludes, "whereas they reported Async Wait and Concurrency to be the most common root causes of NOD flakiness, we found only little evidence of these categories". — https://arxiv.org/pdf/2101.09077
  • [R39] Same study, on how hard detection really is: "A 95 % confidence that a passing test case is not flaky on average would require 170 reruns". — https://arxiv.org/pdf/2101.09077
  • [R40] Mozilla, classified by the developers who fixed them (Eck, Palomba, Castelluccio, Bacchelli, ESEC/FSE 2019; 21 professional developers labelling 200 flaky tests, N=234 labels): Concurrency 61, Async Wait 52, *Too Restrictive Range 40, Test Order Dependency 22, *Test Case Timeout 18, Resource Leak 14, *Platform Dependency 10, Float Precision 6, *Test Suite Timeout 4, Time 4, Randomness 3. Starred categories are new, "not included in the taxonomy proposed by Luo et al." — https://arxiv.org/pdf/1907.01466
  • [R41] Same study, the new categories are the expensive ones: "four of which have never been reported before, despite being the most costly to fix". — https://arxiv.org/pdf/1907.01466
  • [R42] JavaScript (Hashemi, Tahir, Rasheed; 452 commits from top-scoring GitHub JS projects): Concurrency 74 (20.7%), Async Wait 70 (19.6%), OS 66 (18.4%), Network 45 (12.6%), Platform 37 (10.3%), UI 21 (5.9%), Hardware 17 (4.7%), Time 12 (3.4%), Resource Leak 10 (2.8%), Other 13 (3.6%). — https://arxiv.org/pdf/2207.01047
  • [R43] Rust (Schroeder, Phan, Chen et al., UIUC; 53 fixed flaky tests inspected so far): "asynchronous wait (33.9%), concurrency issues (24.5%), logic errors (9.4%), and network-related problems (9.4%)". — https://arxiv.org/pdf/2502.02760
  • [R44] SAP HANA (Berndt, Bach, Baltes; 559 fixed-flakiness issue reports, FTW ’26): "SAP HANA's tests most commonly suffer from issues related to concurrency (23%, 130 of 559 analyzed issue reports)". Async Wait is 52 of 559 (9%), behind Timeout (16%), Oracle-Brittleness (10%) and Configuration (10%). — https://arxiv.org/pdf/2602.03556
  • [R46] Test order dependency measured by detection rather than commit archaeology (Lam, Oei, Shi, Marinov, Xie, iDFlakies, ICST 2019): "Using iDFlakies, we build a dataset of 422 flaky tests, with 50.5% order-dependent and 49.5% not". — https://taoxie.cs.illinois.edu/publications/icst19-idflakies.pdf
  • [R47] Order dependency mostly produces false alarms, but not always: of 96 order-dependent tests found across five Java issue trackers, "94 of which caused a false alarm, a test failure in absence of a bug, and two caused a missed alarm", a bug with no test failure. Secondary source: Parry et al. summarising Zhang et al. ISSTA 2014; the Zhang primary was not fetched. — https://eprints.whiterose.ac.uk/id/eprint/230095/1/parry2021.pdf
  • [R48] The one real cost measurement (Leinen, Elsner, Pretschner, Stahlbauer, Sailer, Juergens, ICST 2024; commercial project, ~30 developers, ~1M SLoC, five years of history, CI logs plus version control, tickets and tracked work time): "the time spent dealing with flaky tests in the studied project represents at least 2.5% of the productive developer time … investigating potentially flaky test failures, which accounts for 1.1% of the total time spent, repairing flaky tests adds another 1.3%, and developing tools to monitor flaky tests adds 0.1%". Official ICST 2024 programme page; the full PDF was unreachable, so nothing beyond the abstract is attributed to this study. — https://conf.researchr.org/details/icst-2024/icst-2024-industry/1/Cost-of-Flaky-Tests-in-CI-An-Industrial-Case-Study
  • [R49] Same study, and it cuts against the moralising: "Contrary to most other studies, we find the cost for rerunning tests to be negligible and inexpensive. Automatically rerunning a test costs 0.02~cents, while not rerunning and thus letting the pipeline fail results in a manual investigation costing $5.67 in our context". — https://conf.researchr.org/details/icst-2024/icst-2024-industry/1/Cost-of-Flaky-Tests-in-CI-An-Industrial-Case-Study
  • [R50] Same study's decision: "The insights gained from our case study have led to the decision to shift effort from investigation and repair to automatically rerunning tests". — https://conf.researchr.org/details/icst-2024/icst-2024-industry/1/Cost-of-Flaky-Tests-in-CI-An-Industrial-Case-Study
  • [R53] Flaky tests come in clusters, not singles (Parry, Kapfhammer, Hilton, McMinn, "Systemic Flakiness", EASE 2025; 10,000 test-suite runs, 24 Java GitHub projects, 810 flaky tests): "75% of flaky tests across all projects belong to a cluster, with a mean cluster size of 13.5 flaky tests". — https://arxiv.org/pdf/2504.16777
  • [R54] Same paper on what drives the clusters: "we identified intermittent networking issues and instabilities in external dependencies as the predominant causes of systemic flakiness", and its framing, "This study represents an inflection point by challenging the deep-seated assumption that flaky test failures are isolated occurrences". — https://arxiv.org/pdf/2504.16777
  • [R55] Developers' own ranking disagrees with the literature (Parry, Kapfhammer, Hilton, McMinn, ICSE-SEIP 2022; 170 survey responses plus 38 StackOverflow threads): "developers rate issues in setup and teardown to be the most common causes of flaky tests"; network issues second; "unknown reasons" third, so "This indicates that the causes of many flaky tests go undiagnosed by developers". — https://eprints.whiterose.ac.uk/id/eprint/230090/1/parry2022a.pdf
  • [R56] Same survey, the mechanism by which flakiness turns into escaped defects: "developers who experience flaky tests more often may be more likely to ignore potentially genuine test failures". — https://eprints.whiterose.ac.uk/id/eprint/230090/1/parry2022a.pdf
  • [R57] The deck is "The State of Continuous Integration Testing @Google" by John Micco (jmicco@google.com), 31 slides, hosted in Google's own research archive. — https://research.google.com/pubs/archive/45880.pdf
  • [R58] The tests-ever-observed measure, verbatim: "Almost 16% of our 4.2M tests have some level of flakiness". Unit = individual tests; criterion = some level of flakiness, ever observed. — https://research.google.com/pubs/archive/45880.pdf
  • [R59] The deck's own definition: "Flakiness is a test that is observed to both Pass and Fail with the same code". — https://research.google.com/pubs/archive/45880.pdf
  • [R60] A third Google number, different unit again: "Continual rate of 1.5% of test executions reporting a 'flaky' result". Unit = test executions. — https://research.google.com/pubs/archive/45880.pdf
  • [R61] Why it never goes away: "Observed insertion rate is about the same as fix rate", under the slide heading "Flakes are Inevitable". — https://research.google.com/pubs/archive/45880.pdf
  • [R62] Google's stated conclusion: "Conclusion: Testing systems must be able to deal with a certain level of flakiness." followed by "Preferably minimizing the cost to developers". — https://research.google.com/pubs/archive/45880.pdf
  • [R63] Google's rerun policy, in its own words: "We re-run test failure transitions (10x) to verify flakiness" and "If we observe a pass the test was flaky". — https://research.google.com/pubs/archive/45880.pdf
  • [R64] Google's own cause taxonomy, from the "Sources of Flakiness" slide, "Factors that cause flakes": Test case factors (Waits for resource, sleep(), Webdriver test, UI test); Code being tested (Multi-threaded); Execution environment/flags (Chrome, Android). A qualitative list with no percentages attached, in which sleep() is named by Google itself. — https://research.google.com/pubs/archive/45880.pdf
  • [R65] The peer-reviewed targets measure. Memon, Gao, Nguyen, Dhanda, Nickell, Siemborski, Micco, "Taming Google-Scale Continuous Testing" (ICSE-SEIP 2017), Table II "PARTITIONING OUR DATASET BY PASSED/FAILED HISTORY", verbatim rows: Total Targets 5,562,881; Never FAILED, PASSED at least once 5,082,803; Never PASSED, FAILED at least once 15,893; Never PASSED/FAILED, most likely SKIPPED 349,025; PASSED at least once AND FAILED at least once 115,160; Flakes 46,694 of 115,160; Remaining 68,466. Unit = test targets. — https://research.google.com/pubs/archive/45861.pdf
  • [R66] A target is not a test: the same paper describes milestones "as large as 4.2 million tests as selected using reverse dependencies on changed source files", while Table II partitions 5,562,881 test targets. The two Google figures therefore count different objects. — https://research.google.com/pubs/archive/45861.pdf
back to top

Related Articles

Book 30 min with Albert
Smiling man with short dark hair and glasses wearing a black suit, white shirt, and black tie against blue background.
Tell Albert what you're shipping.
He'll read this before joining the call. Phone number comes next, on the calendar step.
↳ info@you-source.com
↳ 4-hour response
Please wait while we retrieve meeting schedules.
Oops! There's a problem with your request. We're working on fixing it. Please try again later.