
The key goes in on a Friday afternoon. Someone needs the payments integration working on their laptop, pastes the production credential into a settings file "just for now", and pushes. On Monday a reviewer spots it. The author pushes a commit that deletes the line, the PR description says "removed secret", and the ticket closes.
Nothing in that sequence fixed anything. The credential still works. It's still in the repository's history. And if the repository was public, the peer-reviewed study that measured it found a freshly pushed string became searchable on GitHub in a median of 20 seconds [R6].
That's the gap this post is about: the distance between the configuration your team runs in development and the configuration that ships to production, and the secrets that leak across it. The research on it is better than the folklore, and it says something uncomfortable. The mistakes are ordinary, the clean-up (when it happens at all) is the wrong one [R7] [R8], and the configuration change that takes production down is more likely to be a valid change than a mistake [R18].
Secrets management is the discipline of creating, storing, distributing, rotating and revoking the credentials software uses (API keys, database passwords, private keys, tokens) so that each environment has its own, none of them live in source code, and any one of them can be killed without a deploy. The Twelve-Factor App defines the wider category, config, as "everything that is likely to vary between deploys (staging, production, developer environments, etc)" [R39].
Key takeaways
Twelve-Factor's litmus test is the cleanest statement of the goal: could the codebase "be made open source at any moment, without compromising any credentials"? [R39] If the answer is no, a credential is living in the code, and the dev/prod boundary is a convention rather than a control.
There are two separate rules hiding in that test, and it's easy to honour one and miss the other.
The first is separation of config from code. Credentials don't go in source files, committed config files, Dockerfiles or CI definitions. OWASP Top 10:2025 now says it directly under Security Misconfiguration: "Use identity federation, short-lived credentials, or role-based access mechanisms provided by the underlying platform instead of embedding static keys or secrets in code, configuration files, or pipelines." [R31]
The second is separation of environments from each other. Development, staging and production each get their own credentials, so that a leaked development key is a development problem. The same OWASP category puts it in one sentence: "Development, QA, and production environments should all be configured identically, with different credentials used in each environment." [R30] OWASP's Secrets Management Cheat Sheet goes one step further and suggests "separating the production and development secrets by having separate secret management solutions. Then, reduce access to the production secrets management solution." [R34]
Note what that A02 sentence asks for: identical configuration, different credentials. The leaks below break the second half. The outages further down break the first.
The study to anchor on is Meli, McNiece and Reaves, "How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories", published at NDSS 2019 [R1]. They ran two collections side by side: "a nearly six-month scan of real-time public GitHub commits and a public snapshot covering 13% of open-source repositories", looking for private key files and "11 high-impact platforms with distinctive API key formats" [R1]. The snapshot alone covered 2,312,763,353 files in 3,374,973 repositories [R3].
The headline: "We find that not only is secret leakage pervasive -- affecting over 100,000 repositories -- but that thousands of new, unique secrets are leaked every day." [R2] The real-time collection, which ran from 31 October 2017 to 20 April 2018, recorded a median of 1,793 unique candidate secrets per day [R3]. When the authors manually reviewed a random sample of 240, they estimated "89.10% of all discovered secrets are sensitive" [R5], meaning real credentials whose exposure was a risk to the owner, rather than test keys.
Then the number that matters for the Friday-afternoon scene. The authors pushed a known string to a known repository, started a timer, and polled GitHub's search API until the string appeared, once a minute for 24 hours: "the median time to discovery was 20 seconds, with times ranging from half a second to over 4 minutes" [R6]. Twenty seconds is the window between git push and "anyone watching can find it." That was measured on 2018 infrastructure.
Two scope notes, because they make the finding stronger rather than weaker. The study is explicitly a floor: "our work is not exhaustive but rather demonstrates a lower bound on the problem" [R4]. And it deliberately left out the secret type with no fixed pattern: "we did not attempt to examine passwords as they can be virtually any string" [R4].
A newer, vendor-produced number points the same way. GitGuardian's State of Secrets Sprawl 2026, which is marketing research from a company that sells secret scanning and should be read as such, reports that "28.65 million new hardcoded secrets were added to public GitHub commits in 2025 alone, a 34% increase year over year" [R42]. It also reports that "Internal repos are roughly 6× more likely than public ones to contain hardcoded secrets" [R44]. Private is not the same as safe; it's the same leak with a smaller audience, until someone gets inside.
Here's where the scene goes wrong, and where the study is at its most useful.
From 4 April 2018 the authors watched every secret their real-time collection found, hourly for the first day and daily after that, to see whether it disappeared from the head of the default branch [R7]. About 6% were removed within the first hour. "At the end of the first day, over 12% of secrets were gone, while only 19% were gone after 16 days." Put the other way: "81% of the secrets we discover were not removed." [R7]
The 19% who did act did the thing from the opening scene. Removals of secrets and files far outpaced removals of repositories, which the authors read as users "simply creating new commits that removed the file or secret" [R7]. The authors checked whether any of them had rewritten history to get rid of the original commit: "none of the monitored repos had their history rewritten, meaning the secrets were trivially accessible via Git history." [R8]
And rewriting history wouldn't have saved them. Using their own repositories, the authors confirmed that deleted commits could still be recovered "with only the commit's SHA-1 ID", for both of the removal methods GitHub recommended at the time, git filter-branch and BFG [R9]. Their conclusion is the sentence to put in your incident runbook: the consequences of "even rapidly detected secret disclosure is severe and difficult to mitigate short of deleting a repository or reissuing credentials." [R9]
Reissuing is the fix. OWASP's Secrets Management Cheat Sheet orders remediation exactly that way: first "Revocation: Keys that were exposed should undergo immediate revocation", then rotation, and only then deletion, with a warning that squashing history "may introduce other problems as it rewrites git history" [R37]. The commit that removes the line is housekeeping. The revocation is the security control.
Two more findings make "we'll just remove it" worse:
If you sell software into the EU, whether an exploited leaked credential becomes a reportable event is now a question with a legal answer attached; the Cyber Resilience Act's Article 14 reporting duty is worth reading with this scenario in mind.
The tempting response to a leak is to find out who did it and send them on a course. The data doesn't support it.
Meli et al. compared the repositories and contributors that leaked against a random control group of roughly 100,000 repositories, on forks, watchers, contributors, and each contributor's public repositories and contribution counts: "neither repo activity nor developer experience are strongly correlated with leakage." [R11] One of their case studies is AWS credentials for "a major government agency" website, committed by a developer who "claims in their online presence to have nearly 10 years of development experience." [R11]
What they did find is a root cause you can act on: "most secret leakage is likely caused by committed cryptographic key files and API keys embedded in code", and "it is clear that the poor practice of embedding secrets in code is a major root cause." [R10] The mechanism doesn't depend on seniority. It depends on whether the codebase makes the wrong thing the easy thing.
The same paper has a small, telling case study: collecting .gitignore files for three weeks, the authors "identified 58 additional secrets" inside them, in the one file whose job is to keep secrets out, which they read as "fundamental misunderstandings of features like .gitignore" [R14].
Tooling helps, but check what it actually covers. GitHub's push protection for users "Is enabled by default" and "Stops you from pushing secrets to public repositories on GitHub". Push protection for repositories, which is what would cover your private ones, "Requires GitHub Secret Protection to be enabled" and "Is disabled by default" [R46]. So the default protects the repositories the vendor data says are less likely to hold secrets [R44], and leaves the others to you. Meli et al.'s verdict on scanners in general still holds: they are "mitigations, taking action late in the secret's lifetime after it has already been exposed" [R15].
Code isn't the only place a key lands, either. GitGuardian puts "About 28% of incidents" entirely outside repositories, "in places like Slack, Jira, and Confluence" [R44], which is the same habit that pastes credentials and customer data into an LLM prompt; the controls overlap with keeping PII out of the LLM.
If your security slide still cites "A05:2021 Security Misconfiguration", it's a version out of date. OWASP Top 10:2025 is released, and "A02:2025 - Security Misconfiguration moved up from #5 in 2021 to #2 in 2025." [R25] The data behind it came from contributors who "donated data for over 2.8 millions applications" [R26].
The prevalence figure to quote is from the edition's introduction: "3.00% of the applications tested had one or more of the 16 CWEs in this category." [R25]
The figure not to quote is on the category page itself, which opens: "100% of the applications tested were found to have some form of misconfiguration, with an average incidence rate of 3.00%" [R27]. Both can't be prevalence. The score table resolves it: 100.00% sits in the Max Coverage column, next to an average incidence rate of 3.00% [R27], and OWASP 2025 defines coverage as "The percentage of applications tested by all organizations for a given CWE" [R28]. The 2021 page used the unambiguous wording, "90% of applications were tested for some form of misconfiguration" [R29]. My reading, not OWASP's statement: the 2025 sentence is a rewording slip, and the 100% means every application in the data was tested for misconfiguration, not that every application had one. If someone quotes "100% of apps are misconfigured" at you, that's where it came from.
Two more details matter for secrets specifically:
What OWASP doesn't give you is a breach rate. Its figures measure how often a weakness appears in tested applications [R28], not how often it caused an incident. The pages cited here give no figure of the form "N% of breaches are caused by misconfiguration", and I didn't find one elsewhere with a disclosed method.
Secrets are the security half of the dev/prod gap. The reliability half is configuration that differs between environments, and it has better evidence than you'd expect.
Yuan and colleagues at the University of Toronto studied "198 randomly selected, user-reported failures that occurred on Cassandra, HBase, Hadoop Distributed File System (HDFS), Hadoop MapReduce, and Redis", limited to tickets marked Blocker, Critical or Major from 2010 onwards (OSDI 2014) [R16]. For each failure they recorded which input events were required to trigger it. A configuration change was required in 23% of them [R17]. Their summary: "Configuration changes: 23% of the failures are caused by configuration changes. Of those, 30% involve misconfigurations." [R18] The remaining majority were valid changes that enabled features which may be rarely used [R18].
Read that denominator carefully. It's 23% of 198 severe failures in five distributed data systems, and the event categories overlap because a failure can need several (the column sums past 100%) [R17]. The authors are also explicit that their method under-samples this exact problem: misconfigurations "are more likely to be reported in user discussion forums, which we chose not to study", so "we do not draw any conclusions on the distribution of faults" [R19].
The finding that survives those caveats is the one that matters here: 70% of the configuration changes that triggered failures were correct. Someone turned on a feature, legitimately, and the system fell over. Yuan et al.'s recommendation is that testing should "combine (both valid and invalid) configuration changes with other operations" [R18].
Yin and colleagues came at the same problem from the support desk (SOSP 2011), studying "546 real world misconfigurations, including 309 misconfigurations from a commercial storage system deployed at thousands of customers, and 237 from four widely used open source systems" [R20]. At the commercial vendor, 27% of customer cases were related to configuration issues, and "Configuration issues cause the largest percentage (31%) of high-severity support requests." [R21] In the more complex systems studied, between 16.7% and 32.4% of the misconfigurations were introduced into systems that had previously worked, and 46% of the vendor's 100 used-to-work cases came from configuration parameter changes "due to routine maintenance, configuring for new functionality, system outages, etc" [R23]. When they broke things, between 16.1% and 47.3% of misconfigurations made the system fully unavailable or severely degraded it [R24], and only 7.2% to 15.5% came with an error message that pinpointed the configuration problem [R22].
Both teams say plainly that their samples are specific: the Yin paper does "not intend to draw any general conclusions about all applications" [R24]. The bridge to your stack is mine, not theirs. A production-only setting, flag or credential is configuration that has, by construction, never run anywhere else. The closest thing in the data is what Yuan et al. found failing, valid changes that enable features [R18], and what Yin et al. found rarely announces itself with a pinpointing error [R22].
This is where Twelve-Factor and OWASP disagree, and it's worth taking both seriously rather than picking a side by habit.
Twelve-Factor's case is about accidents. Config files kept out of version control are "a huge improvement over using constants which are checked into the code repo, but still has weaknesses: it's easy to mistakenly check in a config file to the repo". So "The twelve-factor app stores config in environment variables", because "unlike config files, there is little chance of them being checked into the code repo accidentally". [R40] The .gitignore finding above is a small piece of evidence on its side [R14].
OWASP's Secrets Management Cheat Sheet makes the opposite call for secrets specifically: "environment variables are generally accessible to all processes and may be included in logs or system dumps. Using environment variables is therefore not recommended unless the other methods are not possible." [R35] It's blunter still about baking them into images: secrets "should never be hardcoded using docker ENV or docker ARG commands, as these can easily leak with the container definitions" [R35]. And A02:2025 maps "Exposure of Sensitive Information Through Environmental Variables" as a misconfiguration weakness in its own right [R32].
Both are right about the threat they're addressing. Twelve-Factor is defending against the mistake Meli et al. measured, credentials committed to repositories [R10]. OWASP is defending against what happens at runtime, once the process, its crash dumps and its logs exist.
My resolution, which is a judgement rather than a finding: treat the environment variable as an interface, not a store. The application reads its database password from DATABASE_PASSWORD because that's portable and keeps it out of code. The value is injected at start-up by the orchestrator or a secrets-manager sidecar, from a store that's separate per environment [R34], never written into a Dockerfile, a committed .env or a CI YAML file [R35]. Where the platform offers it, skip the static secret entirely in favour of "identity federation, short-lived credentials, or role-based access mechanisms" [R31]. A credential that expires in an hour has a much shorter leak than one that never does.
Twelve-Factor also has the best line on the other half of the problem. Named environments, it warns, invite developers to "add their own special environments like joes-staging, resulting in a combinatorial explosion of config which makes managing deploys of the app very brittle." [R41] Every one of those environments is a place a production credential can end up.
None of these needs a new tool. Each needs someone to look. This is also the ground a senior reviewer covers in a code audit, and it's worth doing before one.
HEAD, because a deleted line is still a leaked line [R8]. As a quick first pass, git log -p --all -S AKIA shows every commit that added or removed the prefix Meli et al. used to identify AWS access key IDs [R48]. Then apply Twelve-Factor's test: could this repository go public today [R39]?ENV and ARG lines carrying credentials, and CI definitions for literal values [R35]. OWASP's A02:2025 names "code, configuration files, or pipelines" together [R31].On rotation schedules, the standards are less prescriptive than the folklore. OWASP says "You should regularly rotate secrets" and gives no interval [R36]. NIST's current SP 800-63B says the opposite for human passwords, "Verifiers and CSPs SHALL NOT require subscribers to change passwords periodically", while requiring a change when there's evidence of compromise [R45]. The cheat sheet follows NIST and excludes user credentials from regular rotation [R36]. For machine secrets, the priority the sources above actually support is being able to revoke and reissue quickly [R37]; a fixed calendar interval is a policy choice, not a standard.
Run those five checks and you end up with a list: keys to revoke, a production credential shared with staging, a Dockerfile to rewrite, a production-only flag nobody has tested, and a rotation procedure that exists in one engineer's head, which is a key-person risk in its own right. None of it is feature work, so it keeps losing at sprint planning, and 81% of the leaked secrets in the NDSS study were still sitting there after 16 days of monitoring [R7].
That backlog is what Dev On Demand is for: one dedicated AI-augmented engineer for $3,495/mo, or two in parallel for $6,795/mo, on a 3-day task cycle with daily async updates and task-by-task approval. First ship in 5 days, a 5-day replacement guarantee, 100% IP ownership, and no lock-in: cancel any time [R47]. Hand it over one item at a time, revocations first.
Before any of that, run git log -p --all -S AKIA on your main repository and count how many of the keys it finds you can prove were revoked.
No. Deleting the line removes the secret from the current version, not from history, and in the NDSS 2019 study none of the repositories that removed a secret rewrote their history, leaving it "trivially accessible via Git history" [R8]. Even rewriting history with git filter-branch or BFG didn't protect it, because deleted commits could still be recovered from GitHub with the commit's SHA-1 ID [R9]. The fix is to revoke the credential at the provider immediately, then rotate, then clean up the repository, which is the order OWASP's Secrets Management Cheat Sheet gives [R37].
In the NDSS 2019 study that measured it, a string pushed to a public GitHub repository became findable through GitHub's search API in a median of 20 seconds, with a range from half a second to over 4 minutes [R6]. That measurement was taken in 2018 by Meli, McNiece and Reaves for NDSS 2019, whose real-time collection recorded a median of 1,793 unique candidate secrets per day [R3]. Treat any secret pushed to a public repository as compromised the moment the push completes.
As an interface, yes; as the store, no. The Twelve-Factor App recommends environment variables for config because, unlike config files, "there is little chance of them being checked into the code repo accidentally" [R40]. OWASP's Secrets Management Cheat Sheet warns that environment variables "may be included in logs or system dumps" and are "not recommended unless the other methods are not possible", and says secrets should never be hardcoded with Docker ENV or ARG [R35]. A workable middle is to read secrets from environment variables in the application while an orchestrator or secrets-manager sidecar injects the values at runtime from a separate store per environment [R34], and to prefer short-lived credentials where the platform supports them [R31].
Security Misconfiguration is A02:2025, up from #5 in the 2021 edition, and OWASP reports that 3.00% of the applications tested had one or more of its 16 CWEs [R25]. The category page's statement that "100% of the applications tested were found to have some form of misconfiguration" refers to the category's maximum testing coverage (the share of applications tested for it), not to how many were vulnerable [R27] [R28]. Its prevention guidance includes "different credentials used in each environment" and avoiding "static keys or secrets in code, configuration files, or pipelines" [R30] [R31]. Hard-coded credentials themselves (CWE-798) are listed under A07:2025 Authentication Failures [R33].
No standard sets a universal interval. OWASP's Secrets Management Cheat Sheet says to "regularly rotate secrets so that any stolen credentials will only work for a short time", without a number, and excludes user credentials from regular rotation [R36]. NIST SP 800-63B states that verifiers "SHALL NOT require subscribers to change passwords periodically" but must force a change on evidence of compromise [R45]. For API keys and service credentials, the measurable priority is being able to revoke and reissue quickly, because in vendor data more than 64% of credentials confirmed valid in 2022 were still valid in January 2026 [R43].
Yes, measurably, and in the study that split them, mostly through correct changes rather than typos. In Yuan et al.'s OSDI 2014 study of 198 severe user-reported failures in Cassandra, HBase, HDFS, Hadoop MapReduce and Redis, a configuration change was required to trigger 23% of the failures, and only 30% of those configuration changes were misconfigurations; the rest were valid changes enabling features [R17] [R18]. In Yin et al.'s SOSP 2011 study, configuration issues were 27% of a storage vendor's customer cases and the largest share, 31%, of its high-severity support requests [R21].