
The build goes red at 16:40 on a Thursday. The failure is in an integration test that touches stock reservation. Nobody opens the stack trace. Someone clicks retry, the second run comes back green, the PR merges, and the ticket closes without a comment.
Taken one incident at a time, that's the correct decision. The engineer had no evidence the failure was real, the rerun costs a fraction of a cent, and an investigation costs an afternoon. Do it two hundred times over a quarter and every instance is still individually defensible.
It's also, exactly, how flaky tests get a real defect waved through. The retry button can't tell the difference between a test that's lying and a test that's right.
A flaky test is a test that both passes and fails on the same code. That is Google's own working definition, from its internal CI deck: "Flakiness is a test that is observed to both Pass and Fail with the same code". [R59]
Key takeaways
sleep, zero removed the flakiness. Of the 42 that used a waitFor, 23 did. [R17]Before asking what causes flakiness, work out how much of it there is. Scale changes what kind of problem it is.
The rerun is not a bad habit somebody picked up. At Google it is documented policy: "We re-run test failure transitions (10x) to verify flakiness" and "If we observe a pass the test was flaky" [R63]. Your engineer hitting retry is doing an unautomated version of what Google does on purpose.
Google has also published three prevalence numbers, which get quoted against each other as if one must be wrong. None is. Each counts a different object, and the unit is the entire explanation.
The first counts tests. From John Micco's deck "The State of Continuous Integration Testing @Google", 31 slides in Google's own research archive [R57], the sentence in full: "Almost 16% of our 4.2M tests have some level of flakiness." [R58] Read the qualifier, not the percentage. The unit is individual tests; the criterion is some level of flakiness, ever observed. A lifetime-ever measure across 4.2 million tests, and not the same claim as "16% of Google's tests are flaky".
The second counts targets, and it's peer-reviewed. Google's ICSE-SEIP 2017 paper "Taming Google-Scale Continuous Testing" partitions its dataset in Table II: 5,562,881 test targets, of which 5,082,803 never failed once and 115,160 both passed at least once and failed at least once. Of that last group: "Flakes 46,694 of 115,160". [R65] Against all 5,562,881 targets that is 0.84%. Against the 115,160 targets that ever produced a mixed signal, about 41%.
The third counts executions. Back in the deck: "Continual rate of 1.5% of test executions reporting a 'flaky' result" [R60]. Point in time, not lifetime-ever.
Three consistent answers to three different questions, and the unit tells you which question you're reading. A target is not a test, and the same paper uses both terms, describing milestones "as large as 4.2 million tests as selected using reverse dependencies on changed source files" while Table II partitions 5,562,881 targets [R66]. And "ever observed to be flaky at all" is not the predicate "classified as a flake in this dataset partition".
Both documents are Google's own, and both are free in Google's public research archive: no paywall, no third-party mirror [R57] [R65]. Check the wording yourself. The compressed version of the first number has travelled the industry for a decade with its qualifier stripped off.
They differ because of method, not disagreement. That is this post's argument in miniature: the same thing has happened to the cause distribution, on a much larger scale, with nobody reconciling it.
Microsoft measured a normal-sized version: across five projects in CloudBuild over one month, "4.6% of all individual test cases are flaky" [R32], but the share of builds with at least one flaky failure ranged from 17% to 52% [R33]. Flakiness concentrates by build, not by test, and that's how the reflex forms. When half your builds go red for no reason, the team learns that red means nothing.
In 161 classified fix commits from 51 Apache projects, Async Wait accounted for 45%, Concurrency 20% and Test Order Dependency 12% [R4] [R6] [R7] [R8]. That's a fact about those 161 commits, in Java, in 2014 — not a distribution you can assume holds in your suite.
Almost every article about flaky-test causes traces back to one paper: Luo, Hariri, Eloussi and Marinov, "An Empirical Analysis of Flaky Tests", FSE 2014, out of the University of Illinois at Urbana-Champaign [R1]. It is a good paper, quoted with the denominator removed, which is how a careful 45% becomes a careless one.
Here's what it did. The authors keyword-searched the whole Apache Software Foundation commit history for "intermit" and "flak", then sampled 201 of the resulting commits for line-by-line inspection [R3]. Those 201 span 51 Apache projects, and across the 486 likely flaky-test fixes the search turned up, the heaviest contributors are HBase, ActiveMQ, Hadoop and Derby [R4] [R2].
Now the part that gets dropped. Of the 201 commits inspected, the authors could precisely classify the root cause of 161. The other 40 were "hard to classify for various reasons" and are excluded from every percentage in the paper [R5]. So the denominator behind "45% Async Wait" is 161. Divide by 201 instead and you get 37%, a number that appears nowhere in the paper.
With that established, the breakdown [R6] [R7] [R8] [R10]:
| Category (the paper's own name) | Count | Share of the 161 classified |
|---|---|---|
| Async Wait | 74 | 45% |
| Concurrency | 32 | 20% |
| Test Order Dependency | 19 | 12% |
| Resource Leak | 11 | 7% |
| Network | 10 | 6% |
| Time | 5 | 3% |
| IO | 4 | 2% |
| Randomness | 4 | 2% |
| Floating Point Operations | 3 | 2% |
| Unordered Collections | 1 | 1% |
The top three account for "77% of the 161 studied commits" [R9]. One caution if you recompute from the paper: its Table 2 counts sum to 163 rather than 161. Use the prose figures — 74, 32 and 19 out of 161 [R6] [R7] [R8].
Async Wait is defined narrowly: a commit lands there "when the test execution makes an asynchronous call and does not properly wait for the result of the call to become available before using it" [R6]. Not "the test is slow". A wrong assumption about when a result is ready.
Concurrency, at 32, decomposes into the shapes of ordinary concurrency bugs: 9 data races, 10 atomicity violations, 2 deadlocks, and a subcategory the authors had to invent and named "bug in condition", at 6 [R21]. If your flakiness lives here the defect is in the code under test, and the mechanics are the ones that produce the race you only see in production.
Test Order Dependency, at 19, is the most interesting because of where it comes from: six are a static field in the code under test, three a static field in the test itself, and ten are an external dependency — a shared file or a shared network port [R20]. More than half are not in-memory state at all, so resetting your test fixtures does not catch them. That comes back in the last section. And for anyone proposing faster CI hardware as the answer: "154 out of 161 (96%) have outcome that does not depend on the platform" [R16].
The best-evidenced finding in the paper is also the least-quoted one. It's the one that should change what a CTO does on Monday.
Finding F.2, verbatim: "Most flaky tests (78%) are flaky the first time they are written." [R11] The arithmetic behind it: "From the 161 tests we categorized, 126 are flaky from the first time they were written, 23 became flaky at a later revision, and 12 others are hard to determine" [R11].
Be precise about what "the first time they are written" means, because the phrasing invites a stronger reading than the method supports. The authors established it by tracing each test's revision history back to the revision that introduced it, not by rerunning anything [R12]. The 23 that became flaky later did so because a new test violated another's isolation, or because the test code changed: bug patches, refactors, incomplete fixes to earlier flakiness [R12].
Then the second number. Across the 152 flaky tests where evolution history was available, Luo et al. report "the average number of days it takes to fix a test to be 388.46" [R13].
Put those together and the standard mental model collapses. That model is decay: a healthy suite accumulates rot as the system grows and CI gets busier, so the intervention is maintenance. A quarterly flake hunt. A dashboard. Somebody's 20% time.
What the data describes is not decay. It's intake. The defect is installed the day the test is written, it passes review because a test that passes is not something a reviewer interrogates, and then it sits in the suite for thirteen months while everybody reruns it. Your review process admitted it, one test at a time, and your rerun policy paid the interest.
Be careful how far that generalises. This is a study of flaky tests developers eventually fixed and wrote a commit message about (more on why that matters shortly). It does not say 78% of the flaky tests currently in your suite were born that way; it says that among flaky tests diagnosed and repaired across 51 large Apache projects, 78% had been flaky since the day they landed. The reading that follows — intake rather than maintenance — is mine rather than the paper's.
The remedy is cheap. A test that asserts on something asynchronous, ordered or shared is assuming something about timing or state, and that assumption either holds always or holds usually. "Holds usually" is the entire category. The review question is not "does this pass" — it passed, that is why you are looking at it. It is "what does this assume about when things happen and what else has run?" Nobody asks, because a green check mark reads as an answer. Which is why chasing a coverage number makes it worse: tests written to hit a target get written fast, against whatever the code happens to do, and that is a failure mode of its own.
sleep() fix in the study failedThe paper's Table 4 is the most immediately actionable thing in the flaky-test literature, and it's brutal.
For Async Wait, the largest category, the authors classified how each fix was implemented and whether it removed the flakiness or only decreased it. Of the 20 Async Wait fixes that added or modified a sleep: zero removed the flakiness, and all 20 merely decreased it. Of the 42 Async Wait fixes that added or modified a waitFor-style call, 23 removed it completely. [R17]
Twenty attempts. Twenty failures. Not "sleep is a code smell" — a measured zero.
The reason sits in the numbers beside it. "The average waiting time for waitFor calls in our cases is 13.04 seconds, while the average waiting time for sleep calls is 1.52 seconds" [R18]. A waitFor is a condition with a generous ceiling: it returns the moment the thing is true and will hang around for thirteen seconds if it has to. A sleep is a guess with a tight budget. Somebody watched the operation take 400ms on their laptop, wrote a 1.5-second sleep, saw green and shipped it. On a loaded runner with a cold cache, 1.5 seconds is not enough.
waitFor is the majority strategy here, not a niche one: 42 of the 74 Async Wait cases, 57%, were fixed that way [R6] [R17]. The correct fix isn't exotic, and people on your team already know it. The 20 sleeps are the ones written under time pressure by whoever was on the hook to get CI green before a release.
Nor is this only a 2014 finding. Google's own "Sources of Flakiness" slide — a qualitative list of factors that cause flakes, with no percentages attached — names sleep() explicitly under test case factors, alongside waits for a resource, Webdriver tests and UI tests [R64]. Two independent parties, a decade and several orders of magnitude apart, name the same culprit.
Async Wait, as practised, does not get removed. It gets reduced, until the failure rate falls under the team's annoyance threshold and the ticket closes.
So here's a diagnostic that costs nothing. Grep your test directories for Thread.sleep, time.sleep, setTimeout, cy.wait( with a number in it. Every hit is somebody encoding a guess about duration as if it were a fact about the system, and in the only study that measured what happens next, that guess was never once the thing that removed the flakiness [R17].
Now the part that undercuts every listicle built on that paper. Six studies have since measured flaky-test causes in different languages and by different methods, and each produces a different top cause.
| Study | Population and method | Top cause | Second |
|---|---|---|---|
| Luo et al. 2014, Apache/Java [R1] [R6] [R7] | 161 classified fix-commits, keyword-mined | Async Wait 45% | Concurrency 20% |
| Eck et al. 2019, Mozilla [R40] | 200 fixed tests, classified by the devs who fixed them | Concurrency, 61 of 234 labels | Async Wait, 52 |
| Lam et al. 2019, iDFlakies, Java [R46] | 422 flaky tests, found by rerunning with reordering | Order-dependent 50.5% | Non-order-dependent 49.5% |
| Gruber et al. 2021, Python [R36] [R37] | 876,186 tests, 22,352 PyPI projects, rerun | Test Order Dependency 59% | Test infrastructure 28% |
| Hashemi et al. 2022, JavaScript [R42] | 452 commits analysed, 358 classified | Concurrency 20.7% | Async Wait 19.6% |
| Schroeder et al. 2025, Rust [R43] | 53 fixed flaky tests, work in progress | Async Wait 33.9% | Concurrency 24.5% |
| Berndt et al. 2026, SAP HANA [R44] | 559 fixed-flakiness issue reports | Concurrency 23% | Timeout 16% |
Follow test order dependency across that table. Luo puts it at 12% [R8]. iDFlakies, which found its flaky tests by rerunning suites in randomised orders rather than by reading commit logs, puts order-dependent tests at 50.5% of 422 [R46]. Gruber's Python study, the largest in this set at 876,186 test cases, reports that "Order dependency is a much more dominant problem in Python, causing 59 % of the 7 571 flaky tests in our dataset" [R36]. From 12% to 59%.
It gets sharper. Gruber's team went looking for Luo's top two categories and could not find them: "whereas they reported Async Wait and Concurrency to be the most common root causes of NOD flakiness, we found only little evidence of these categories" [R38]. Those two scored three cases each in their sample, while 28% of the flakiness came from test infrastructure: "a previously undocumented cause of flakiness", a category Luo's taxonomy had no slot for [R37]. Mozilla shows the same gap: four of the eleven categories Eck's developers needed were new, "not included in the taxonomy proposed by Luo et al." [R40] [R41].
None of this means anybody measured wrong. It means the method decides the answer.
Luo mined commit messages for two word stems, and the paper concedes the consequence in its own Threats to Validity: "there is no guarantee on the recall of our search. In fact, we believe our search could miss many flaky tests whose fixes could use words like 'concurrency', 'race', 'stall', 'fail', etc." [R3] Commit archaeology is a census of successful diagnoses: flakiness a developer noticed, named and repaired. Detection-based studies do the opposite: iDFlakies and Gruber's team reran suites in original and shuffled order and recorded what came out inconsistent [R46] [R36], which finds flakiness nobody noticed and nobody fixed.
The two methods fail in opposite directions. Order dependency is trivially exposed by a tool that shuffles test order and nearly invisible to a developer who always runs the suite the same way. Async Wait is the reverse: it fails with a stack trace and a timeout in it, precisely the kind of failure a developer diagnoses and writes a commit message about. Archaeology is biased towards the diagnosable, reruns towards the exposable. (That reconciliation is my reading — the studies do not claim to explain each other's results.)
Developers agree with neither. Parry's survey of 170 developers put setup and teardown first, network second and unknown reasons third: "the causes of many flaky tests go undiagnosed by developers" [R55]. When the practitioner's third-ranked answer is "I don't know", the tidy percentages are describing the diagnosable minority.
So "45% of flaky tests are Async Wait" is not a fact about software. It's a fact about 161 classified Apache commits, found by searching for two word stems, in Java, in 2014. A fine starting hypothesis for a Java codebase with heavy async integration tests, a bad thing to put in a deck about your Python service.
The moralising version of this post would stop here and say "quit rerunning and fix your tests". That version is wrong, and the only team that has actually costed the alternatives says so.
Leinen and colleagues (ICST 2024) instrumented a commercial project — roughly 30 developers, about a million lines of code, five years of history, combining CI logs, version control, tickets and tracked work time. Their headline: "the time spent dealing with flaky tests in the studied project represents at least 2.5% of the productive developer time", split as 1.1% investigating flaky failures, 1.3% repairing them, 0.1% building tools to monitor them [R48].
Then the sentence that belongs in every argument about this: "Contrary to most other studies, we find the cost for rerunning tests to be negligible and inexpensive. Automatically rerunning a test costs 0.02~cents, while not rerunning and thus letting the pipeline fail results in a manual investigation costing $5.67 in our context." [R49] Their response was not to lecture anyone: "The insights gained from our case study have led to the decision to shift effort from investigation and repair to automatically rerunning tests" [R50]. (Only that study's abstract was reachable, so those figures are all there is.)
Google arrives at the same position from the other end of the scale. Under the heading "Flakes are Inevitable", the deck records that the "Observed insertion rate is about the same as fix rate" [R61]. The conclusion drawn is not eradication: "Testing systems must be able to deal with a certain level of flakiness. Preferably minimizing the cost to developers." [R62] Which is why the rerun is institutionalised at 10x [R63].
If insertion and removal balance, the backlog is constant and the only lever left is cost per incident, which reruns push towards zero. Here's what that model doesn't price.
Reruns are also how real bugs get normalised. Parry's survey work names the mechanism: "developers who experience flaky tests more often may be more likely to ignore potentially genuine test failures" [R56]. The rerun is cheap per incident; the habit is not, because the habit is indiscriminate. A team that has learned red means nothing has lost the ability to act on red meaning something.
And flaky tests are finding real bugs right now. In Luo's study, 38 of the 161 commits — 24% — fixed the flakiness by changing both the test and the code under test, and 94% of the fixes that touched production code fixed a real bug in it [R14]. Roughly one flaky-test fix in five turned out to be a real product bug, in software as heavily exercised as HBase and Hadoop. The implication is stated as bluntly as a paper can: "Flaky tests should not simply be removed or disabled because they can help uncover bugs in the CUT." [R15] And the failure does not always announce itself: in Parry's survey of the literature, Zhang et al. are reported as having examined 96 order-dependent tests, two of which caused a missed alarm: a bug with no test failure at all [R47]. Second-hand, so treat it as indicative, but no rerun policy is designed to catch a test that is silently green while something is broken.
So hold both. Rerunning is the right call at the moment of failure and a corrosive policy over a quarter, and no amount of arguing resolves that into one answer. Stop treating them as one decision. Rerun automatically, because Leinen's arithmetic is not beatable on cost. But make the rerun record something: which test, which build, how many attempts, on what branch. A rerun that leaves a row in a table is a measurement; a rerun that leaves nothing is a decision delegated to whoever was watching CI that afternoon. Google's insertion rate is a known quantity only because somebody counted it [R61].
Which leaves one defensible move. None of those distributions is about your suite, so measure your suite.
Measuring is not archaeology. You are not reading commit messages for "flaky" and "intermittent": that is the method that produced the distribution which does not replicate [R3]. Detection means running the suite repeatedly against an unchanged commit and recording which tests come out inconsistent, then running it again in a randomised order. That is how iDFlakies found 50.5% of its 422 flaky tests to be order-dependent [R46], and the shuffled run is the cheap half that archaeology systematically misses.
Two constraints, both measured. Run it in CI, not on a developer machine: when Microsoft's team re-ran known flaky tests locally 100 times, 86% were only flaky in the CI pipeline [R34]. And do not expect certainty: "A 95 % confidence that a passing test case is not flaky on average would require 170 reruns" [R39]. That is the cost of proving a negative. Aim at finding flaky tests, not at certifying the rest as clean.
Then the finding that changes the shape of the work rather than just its target. Parry, Kapfhammer, Hilton and McMinn re-analysed an existing dataset of 10,000 test-suite runs across 24 Java projects containing 810 flaky tests, and found they do not arrive one at a time: "75% of flaky tests across all projects belong to a cluster, with a mean cluster size of 13.5 flaky tests" [R53], sharing a driver: "we identified intermittent networking issues and instabilities in external dependencies as the predominant causes of systemic flakiness" [R54].
So the standard flaky-test backlog is misconceived. A ticket per flaky test, triaged individually by whoever has capacity, is a treadmill by construction: thirteen and a half tickets that all resolve to one unstable dependency, each producing its own local workaround, and the dependency untouched at the end.
The unit of repair is the dependency boundary. That's consistent with everything else here: ten of Luo's 19 order-dependency cases were an external dependency rather than in-memory state [R20], and 28% of Gruber's Python flakiness was test infrastructure [R37]. Cluster your flaky tests by what they touch before you triage them by name, and most of the backlog collapses into a handful of boundaries: the inconsistently stubbed service, the shared database tests do not reset, the port two suites both want.
So, concretely, for a suite you have never measured:
Whether a given check belongs in an automated suite at all is a separate question, worth asking before you spend a month stabilising a test that should have been manual.
Run that exercise and you end up with a list nobody disputes is broken and nobody is assigned to. It's not feature work, so it never survives a sprint boundary, and every quarter somebody new rediscovers it. The average flaky test in the only study that tracked it lasted 388 days [R13]: not a statement about difficulty, but one about ownership.
Getting that list owned is what QA On Demand is staffed for: single stream $3,495/mo, dual stream $6,795/mo for two streams in parallel. A 3-day task cycle, daily async updates, first ship in 5 days, a replacement guarantee within 5 business days, cancel any time with no lock-in, and you own the IP and the test assets. It's manual-first by design: hands-on functional and regression testing first, automation added as your product stabilises, with the automated suites landing in your repo. That order matters. Nobody arrives on Monday and repairs your async integration tests. What gets covered first is the job your broken suite has been failing to do — a person checking whether the thing actually works.
Pick the ten tests your team reruns most often this week, run each of them ten times against an unchanged commit, and write down which ones you cannot make fail twice the same way.
There is no single answer: the measured distribution changes with the language, the test types, and the method used to find the flaky tests. The most-cited study, Luo et al. FSE 2014, classified 161 flaky-test fix commits from 51 Apache projects and found Async Wait at 45%, Concurrency at 20% and Test Order Dependency at 12%, together 77% of those 161 classified commits [R6] [R7] [R8] [R9]. That distribution has not replicated. A study of 876,186 Python tests found test order dependency at 59% and test infrastructure at 28%, and reported "only little evidence" of Luo's top two categories [R36] [R37] [R38]. Mozilla developers classifying their own fixes put Concurrency first and needed four categories Luo's taxonomy did not have [R40] [R41]. What survives across all of them is the mechanism rather than the percentages: flakiness is a wrong or missing assumption about ordering, waiting, or shared state.
In the one study that traced it, yes: 126 of the 161 flaky tests Luo et al. classified were flaky from the revision that introduced them, 23 became flaky at a later revision, and 12 were hard to determine, which the paper states as "Most flaky tests (78%) are flaky the first time they are written" [R11]. The mechanism is that a test that passes does not get interrogated in code review. They established the finding by tracing each test back to its introducing revision, not by rerunning anything [R12]. Read it as a result about flaky tests that were eventually diagnosed and fixed across 51 Apache projects, not as a rate for the flaky tests sitting in your suite today. The practical reading is that flakiness is an intake problem rather than a decay problem: the assumption about timing or shared state is baked in at authoring time, and then survives on average 388 days before anyone fixes it [R13].
sleep() fix a flaky test?No, on the only evidence available. In Luo et al.'s Table 4, of 20 Async Wait fixes that added or modified a sleep, zero removed the flakiness and all 20 merely decreased it; of the 42 fixes that used a waitFor-style call, 23 removed it completely [R17]. The reason shows up in the wait times: "the average waiting time for waitFor calls in our cases is 13.04 seconds, while the average waiting time for sleep calls is 1.52 seconds" [R18]. A sleep encodes a guess about duration; a waitFor encodes the condition you actually care about and tolerates a slow run. Google's own "Sources of Flakiness" slide names sleep() as a cause of flakes [R64].
Per incident, yes, and the arithmetic is not close. Leinen et al. measured an automatic rerun at $0.0002 against $5.67 for the manual investigation it replaces, found that dealing with flaky tests consumed at least 2.5% of productive developer time, and decided to "shift effort from investigation and repair to automatically rerunning tests" [R48] [R49] [R50]. Google reaches the same conclusion at its own scale, recording that "Testing systems must be able to deal with a certain level of flakiness" and rerunning failure transitions 10x as policy [R62] [R63]. The cost the rerun does not price is behavioural, because "developers who experience flaky tests more often may be more likely to ignore potentially genuine test failures" [R56], and because flaky tests find real defects: 24% of the fixes in Luo's study changed the code under test as well as the test, and 94% of the fixes that touched production code fixed a real bug in it [R14] [R15]. Rerun, but record every rerun, so the insertion rate is a number you know rather than a habit you have.
Detect them, don't research them. Run the full suite repeatedly against one unchanged commit and record every test that is not consistent, then run it again with the test order randomised. That reordering step is what iDFlakies used to find 50.5% of its 422 flaky tests were order-dependent, a category commit-log analysis systematically under-counts [R46]. Run it in CI rather than locally, because when Microsoft re-ran flaky tests locally 100 times, "86% of them are only flaky in the CI pipeline" [R34]. Aim at finding flaky tests rather than certifying the rest as clean, because "A 95 % confidence that a passing test case is not flaky on average would require 170 reruns" [R39]. Then group what you find by the external dependency it touches rather than by test file: 75% of flaky tests belong to clusters averaging 13.5 tests, driven predominantly by "intermittent networking issues and instabilities in external dependencies" [R53] [R54].