PII in LLMs: What Actually Leaks, and What to Build
For about $200 in API queries, a research team extracted more than 10,000 unique verbatim training examples out of production ChatGPT [A15]. When they hand-labelled 15,000 generations from that attack, 16.9% contained memorised personal information, and 85.8% of the generations that contained potential PII contained real PII [A16].
That's PII in LLMs measured against a live, aligned, commercially deployed model. It's not the story anyone tells. The two that circulate in every "AI and customer data" conversation are a Korean press report nobody has confirmed and a fine an Italian court threw out in March. The measured evidence is stronger than either, and it points at your architecture rather than your vendor's press release.
PII in LLMs is personal data about identifiable people — names, contact details, records — that a language model has absorbed from its training or input data and can reproduce, or infer, in its outputs. Memorisation is the specific case where the model emits a training example verbatim instead of generating something new. Because a model can do that, a trained model is not automatically anonymous.
Key takeaways
- A 2023 divergence attack on production ChatGPT extracted more than 10,000 unique verbatim training examples for about $200 in API queries [A15]; 16.9% of the hand-labelled generations contained memorised personal information [A16]. The authors conclude that current alignment techniques do not eliminate memorisation [A17].
- A separate 2021 study of GPT-2 confirmed 604 memorised training examples out of 1,800 candidates [A2], including 46 containing individual people's names [A4] and 32 containing contact information [A5]. The authors call that 0.1% rate "likely an extremely loose lower bound" [A10], and found the 1.5-billion-parameter GPT-2 memorised over 18 times as much as the 124-million-parameter one [A8]. GPT-2 was trained on public web data, so these figures describe that model and say nothing about a modern model trained on your prompts.
- EDPB Opinion 28/2024 states that AI models trained on personal data "cannot, in all cases, be considered anonymous" and must be assessed case by case [A28]; its paragraph 43 sets a two-limb test covering both extraction from the model and disclosure through queries [A29].
- PII redaction reduces exposure but does not eliminate it. Scrubbing techniques "reduce but do not prevent the risk of PII leakage" [A19], and a membership-inference attack still scores AUC 0.82 against a scrubbed model versus 0.505 against a scrubbed-and-DP model [A24].
- A provider's retention window is a dated vendor self-statement, not a property of the system. OpenAI's published 30-day API retention carries an "unless we are legally required to retain them" exception [A65], a US court order made that exception live in 2025 [A68], and among API customers only data that is never stored was unaffected [A69][A70].
What PII in LLMs actually leaks, measured
Start with the paper everyone cites on training data extraction. The chain of numbers in it matters more than the headline. Carlini and eleven co-authors from Google, Stanford, UC Berkeley, Northeastern, OpenAI, Harvard and Apple published "Extracting Training Data from Large Language Models" at USENIX Security in 2021 [A13]. The attack used only black-box query access; they treated the model as a generative function and never looked inside it [A6]. They generated 600,000 samples from GPT-2 and selected 1,800 candidates: 100 under each of 18 attack configurations, being three text-generation strategies crossed with six membership-inference strategies [A1]. Four of the authors then manually adjudicated all 1,800, and the GPT-2 authors fuzzy-matched the set against the original training corpus [A3].
What survived was 604 unique memorised training examples [A2]. Inside those 604: 46 examples containing individual people's names, explicitly excluding anything about famous politicians or national news [A4], and 32 containing some form of contact information — a phone number, a social media handle — of which 16 were businesses and 16 were private individuals [A5].
Most summaries flatten that chain. 600,000 generations, 1,800 candidates, 604 confirmed, 46 names, 32 contact records. The 604 works out to 0.1% of the generations, and the authors call that "likely an extremely loose lower bound" [A10]. They weren't trying to find everything, only to prove the attack works.
Two findings have aged into the load-bearing ones. Scale makes it worse: in one setting the 1.5-billion-parameter GPT-2 memorised over 18 times as much content as the 124-million-parameter model, and the abstract states flatly that larger models are more vulnerable [A8]. And this isn't overfitting. Large language models show no significant train-test gap and still emit training data verbatim, because some individual training examples have anomalously low loss [A11]. The comfortable mitigation — regularise harder, don't overfit — doesn't reach the failure mode.
Now the caveat the authors insist on. GPT-2 was trained on public data, and Carlini et al. say their attacks "are not particularly severe" for that reason: everything extracted could also be found via internet search. The attack was indiscriminate, not targeted. Fair objection, and it's why those numbers should never be quoted as though they describe GPT-5 or Claude.
The 2023 sequel is what closes the gap. Attacking production gpt-3.5-turbo, the same research line found a "divergence attack" that pushes the model out of its chatbot behaviour and into emitting training data at a rate 150 times higher than normal operation [A14]. Then the $200 and the 10,000 examples [A15], and the 16.9% PII rate [A16]. Their conclusion runs one sentence: practical attacks recover far more data than previously thought, and current alignment techniques do not eliminate memorisation [A17]. RLHF is a behaviour layer, not a containment layer.
The trend runs the wrong way too. A separate quantification paper finds memorisation growing log-linearly with model capacity, with duplication in training and with prompt context length, and expects it to "likely get worse as models continue to scale, at least without active mitigations" [A18].
The two PII-in-LLMs stories everyone tells
Neither of the two anecdotes that dominate this topic survives being checked.
The first is Samsung. You know the version: engineers pasted proprietary source code into ChatGPT, the company banned it, everyone learned a lesson. Every English-language account traces back to a single report in The Economist Korea. Cybersecurity Dive, among the more careful outlets to cover it, attributes the claim to that report and adds the sentence everyone else dropped: Samsung Electronics did not respond to requests for comment. No Samsung statement, no incident report, no regulator filing [A81]. It may well be true. It isn't documented.
The second is the Italian fine. The Garante's provvedimento n. 755 of 2 November 2024 imposed EUR 15 million on OpenAI and ordered a six-month public information campaign across radio, TV, newspapers and the internet, using the Authority's Article 166(7) powers for the first time [A76]. The findings were substantive: failure to notify the March 2023 breach, training on users' personal data without first identifying an adequate legal basis, breach of the transparency principle and the related information obligations, and no age verification [A77].
That fine is not live. A footnote on the Garante's own press release records that provision 755 was removed from the Authority's website following the judgment of the Tribunale di Roma n. 4153/2026, published 18 March 2026, which upheld the opposition brought against it [A78]. A regulator publishing the record of its own defeat is about as strong as secondary evidence gets.
Be precise about what that doesn't mean. The judgment text itself was not reachable during this research; the evidence here is the Garante's footnote recording the outcome, not the ruling. The grounds are unknown, and so is whether the Garante has taken it further. Nobody should characterise why the court decided as it did on the strength of a footnote.
What survives is the shape of the objection, not the penalty. Enforcement attention didn't go away either. The case file went to the Irish DPC under the one-stop-shop rule once OpenAI established its European headquarters in Ireland [A79], and the Garante's press room since lists a fine against Character.AI, a block on DeepSeek, and an investigation into OpenAI's Sora [A80]. The regulator lost a case. It didn't lose interest.
What the regulators actually said about GDPR and AI models
The most useful regulatory document here isn't a fine. It's EDPB Opinion 28/2024, adopted 17 December 2024 at the request of the Irish supervisory authority under Article 64(2) GDPR, which asks when an AI model can be considered anonymous [A27]. This is information for engineering planning, not legal advice; your DPO owns the conclusion.
Paragraph 34 is the headline, and it kills the most popular engineering defence: "AI models trained on personal data cannot, in all cases, be considered anonymous. Instead, the determination of whether an AI model is anonymous should be assessed, based on specific criteria, on a case-by-case basis" [A28]. Paragraph 31 explains why, aimed squarely at "it's just weights": information from the training dataset, including personal data, "may still remain 'absorbed' in the parameters of the model, namely represented through mathematical objects" — objects that may differ from the original data points but may still retain the original information, which may ultimately be extractable directly or indirectly from the model [A30]. The mathematical-abstraction argument was considered. It didn't land.
Paragraph 43 is the part an engineer can work with, because it's a test with two limbs. For a model to be considered anonymous, using reasonable means, both the likelihood of direct or probabilistic extraction of training-set personal data and the likelihood of obtaining such data from queries, intentionally or not, must be "insignificant for any data subject" [A29]. Both limbs fail closed. Some models never pass at all: paragraph 29 says a model designed to reply with personal data from training when prompted about a specific person cannot be considered anonymous [A31].
Paragraph 55 is where this stops being law and starts being a test plan. The EDPB names the attack classes a controller should run structured testing against — attribute and membership inference, exfiltration, regurgitation of training data, model inversion, reconstruction attacks — while warning that passing widely-known state-of-the-art attacks is only evidence of resistance to those attacks [A35]. Paragraph 54 goes further: supervisory authorities should evaluate document-based audits, which "could include the analysis of reports of code reviews" [A36]. Your engineering artefacts, the same ones a code audit would read, are the evidence.
Two more things worth knowing. Paragraph 51 treats pseudonymisation as data preparation, not a route to anonymity [A37], tracking Recital 26: pseudonymised data attributable to a person using additional information "should be considered to be information on an identifiable natural person" [A55]. And paragraph 57 supplies the consequence. A supervisory authority that can't confirm effective anonymisation is in a position to find the controller failed its accountability obligations under Article 5(2) [A40].
NIST arrived at the same place from the security side. AI 600-1, approved by its Editorial Review Board on 25 July 2024 [A43], states in section 2.4 that "models may leak, generate, or correctly infer sensitive information about individuals" [A45]. Note the third verb. NIST warns that models may correctly infer PII that was never in training and never disclosed by the user, by stitching together disparate sources, and that this harms the individual even when the inference is wrong [A46]. Redaction has no answer to inference.
Is the position soft? Somewhat, and honestly so: Opinion 28/2024 hedges throughout, and paragraph 48 says the presence or absence of the listed elements isn't conclusive. The UK picture is thinner still. The ICO's AI guidance was last updated 15 March 2023 and is now under review because of the Data (Use and Access) Act [A58], so it's useful for stable definitions like membership inference — deducing whether an individual was in the training data by exploiting the model's disproportionate confidence about people it has seen [A60] — rather than as the current word.
"Case-by-case, assessed by your supervisory authority" is not an engineering specification. Fine. That's precisely the argument for building controls that hold up whichever way an assessment goes.
Your provider's policy is a policy, not a property
Everything in this section is a vendor's own published statement about its own product. None of it is an independent property of the system, and none of it is permanent.
OpenAI's Enterprise Privacy page, as fetched on 7 September 2026, states: "By default, we do not use your business data for training our models," with an exception where a customer has explicitly opted in through feedback mechanisms [A64]. The same page scopes that default to ChatGPT Business, Enterprise, Healthcare, Edu and Teachers, and the API Platform after 1 March 2023 [A66]. On retention, the page says "Except for certain endpoints and features listed in our platform documentation," OpenAI "may securely retain API inputs and outputs for up to 30 days" to provide the service and identify abuse, after which they are removed "unless we are legally required to retain them", with zero data retention available for eligible endpoints and qualifying use cases [A65]. Fine-tuning gets its own rule: the data submitted is "retained until the customer deletes the files" [A67].
Anthropic's published position, as of writing, is: "By default, we will not use your inputs or outputs from our commercial products to train our models" [A72]. The stated exception is explicit feedback, the thumbs up/down button, and where feedback is submitted the company says it stores the entire related conversation in a secured back-end for up to five years, de-linked from user identifiers [A73]. One caveat: that page carries only a relative date, so read it as the position published at time of writing, not as a policy pinned to a day.
Two providers. Not "the industry". Google's Vertex AI and Microsoft's Azure OpenAI positions were not checked for this piece, so nothing here describes them. If that's your provider, the exercise is the same: find the retention window and the training-use default on their own page, and date it.
Now read that "unless we are legally required to retain them" clause again, then read what happened to it. By June 2025, a US court order in the New York Times litigation required OpenAI to retain consumer ChatGPT and API content indefinitely. OpenAI's own update states: "After months of litigation, we are no longer under a legal order to retain consumer ChatGPT and API content indefinitely. Our obligations under the earlier order ended on September 26, 2025" [A68]. The scope, in OpenAI's words, covered ChatGPT Free, Plus, Pro and Team subscribers "or if you use the OpenAI API (without a Zero Data Retention agreement)" [A69].
One group of API customers was untouched, and the reason is the whole argument: "If you are a business customer that uses our Zero Data Retention (ZDR) API, we never retain the prompts you send or the answers we return. Because it is not stored, this court order doesn't affect that data" [A70].
A retention window is a commitment a court can override. A path that never writes the data has nothing to override. Different kinds of object, and only one of them survives contact with a subpoena. Note too that published product behaviour isn't the same question as what your contract obliges the provider to do: a vendor security questionnaire question, not a docs question.
Does PII redaction stop the leak?
No. PII redaction reduces exposure; it does not eliminate it. The standard answer is to strip the PII before it leaves. Do that. Then read the measured recall, because the gap between "we redact" and "the PII is gone" is where teams get hurt. Lukas and colleagues at Waterloo and Microsoft name the error directly: PII leakage is attributable to "the false assumption that dataset curation techniques such as scrubbing are sufficient to prevent PII leakage. Scrubbing techniques reduce but do not prevent the risk of PII leakage" [A19].
The numbers underneath that sentence:
- Automated scrubbing rests on named-entity recognition, and modern transformer-based NER has mixed recall: around 97% on names but 80% on care unit numbers in clinical health data, "meaning that much PII is retained after scrubbing" [A20].
- A second benchmark in the same paper: a fine-tuned NER model with ground-truth annotations achieves 84–93% recall, dropping to 77% without annotations [A21]. Your recall depends on whether someone labelled your fields, which is an organisational fact rather than a model capability.
- Against a fine-tuned model, extraction isn't marginal. GPT-2-Large recalls 23% of the PII in the ECHR dataset at 30% precision, versus about 9% for GPT-2-Small [A22]. Fine-tuning on your own customer data is where the abstract threat becomes yours.
- Differential privacy narrows the hole without closing it: sentence-level DP "still leaks about 3% of PII sequences" [A23].
- The defence-stacking result is the most useful line in the paper. A membership-inference attack scores AUC 0.96 against an undefended model, 0.82 against a scrubbed model, and 0.505 against a scrubbed-and-DP model [A24]. Scrubbing alone takes the attacker from near-perfect to still-very-good. Only the stack gets to coin-flip.
- Redaction is also more reversible than people assume. Reconstructing PII from masked text, the authors recover 18.27% where a prior method recovers 5.81% [A25].
Their conclusion: "scrubbing cannot fully prevent PII leakage" [A26]. Carlini's team said it two years earlier in plainer words. Sanitising data is imperfect, some private data will always slip through, "and thus it serves as a first line of defence and not an outright prevention against privacy leaks" [A12].
There's a cost, too, and the papers are honest about it: scrubbing and DP protect training-data privacy "at the cost of degrading model utility", aggressive scrubbing "drastically harms utility", and DP lengthens training time. So the statistical defences cost you output quality and compute and still leave a residue, while the architectural ones — don't send it, don't store it — cost neither and are deterministic.
Build the redaction layer. Just don't file it under "solved".
What to do at the code level
Stop treating LLM data privacy as one question. "Where does customer data go?" is four questions with four control points, and teams get hurt by answering the easy one and assuming it covers the rest.
Boundary one: what enters the prompt. The only boundary you fully own, and the cheapest place to win. OWASP lists Sensitive Information Disclosure as LLM02:2025, covering PII, financial details, health records, confidential business data, security credentials and legal documents [A74], with named mitigations starting at data sanitisation, input validation, strict access controls, tokenisation and redaction [A75]. Concretely: a redaction layer in the request path rather than in a runbook, with a recall number you can quote. The research puts it between 77% and 97% depending on field and annotation [A20][A21]. Log what you redacted, not what you redacted from.
Boundary two: what the provider retains. Read your provider's current published position, date it, treat it as revocable. OpenAI's 30-day window carried an explicit "unless we are legally required to retain them" [A65], and for roughly four months in 2025 that exception was live [A68][A69]. Where a no-retention path exists, move the sensitive routes onto it, because unstored data is unaffected by preservation orders [A70]. NIST's action MP-4.1-005 says the same in governance language: policies for collection, retention and minimum quality of data, set explicitly against the risk of leaking PII [A49].
Boundary three: what a fine-tune absorbs permanently. The boundary with no undo. Fine-tuning data is retained until you delete the files [A67], and a fine-tuned model gives up a measurable fraction of the PII it saw: 23% recall at 30% precision in the published experiment [A22]. If a fine-tune touches customer records, the EDPB's paragraph 43 two-limb test applies to the resulting artefact, not to your intentions [A29], and paragraph 52 asks whether you used regularisation and, "crucially", techniques such as differential privacy [A38].
Boundary four: what the model can emit back. The EDPB names output filtering itself: measures "to prevent the storage, regurgitation or generation of personal data, especially in the context of generative AI models (such as output filters)" [A39]. NIST asks for periodic monitoring of AI-generated content for privacy risks, addressing "any possible instances of PII or sensitive data exposure" [A50]. That belongs with everything else you already monitor in production, not on a launch-day checklist.
Then the part almost nobody builds: a test suite. Paragraph 55 hands you the list — attribute and membership inference, exfiltration, regurgitation of training data, model inversion, reconstruction attacks [A35]. Five test classes, not five legal concepts. Run them in CI against your own fine-tuned artefacts and retrieval paths, keep the results, and you have the document-based engineering evidence paragraph 54 describes [A36]. Put it in the pipeline you already have rather than a parallel process. NIST's MP-4.1-003 says exactly that, connecting new policies to existing model, data, software-development and IT governance [A51]. And if the code doing the redacting was written by a model, it needs the same scrutiny as anything else in the request path.
One scoping note: if you ship into the EU, incident and vulnerability reporting duties are a separate regime with separate clocks, not covered here.
None of this needs a research team. It needs someone to do it.
This is task-shaped work
Look at what has to get built: a redaction layer with a measured recall number, a retention policy that maps to code paths rather than a wiki page, sensitive routes on no-retention endpoints, an extraction and membership-inference test suite in CI, and the write-up that makes it demonstrable. Every one is well-specified. None will ever win a sprint against a customer-facing feature, which is why they sit in the backlog until an enterprise deal or a supervisory authority asks the question.
That's the shape of work Dev On Demand exists for: 1 dedicated AI-augmented engineer, a 3-day task cycle with daily async updates, and a task-by-task approval gate so you pick what gets built next. Single stream is $3,495/mo, dual stream $6,795/mo, first ship in 5 days, cancel any time with no notice period required. If you'd rather see the work before subscribing, Proof of Quality is one real task — you judge the engineer first. We build the controls; the compliance judgement stays with you and your counsel.
Start at the cheapest boundary. Take one production endpoint that sends customer data to a model and answer three things in writing this week: what fields leave your network, what your provider's currently published retention position is and the date you read it, and whether anything you've fine-tuned contains a real customer record. If the third answer is yes, fix that one first. It's the only boundary without an undo.
Frequently asked questions
Can an AI model be considered anonymous under GDPR?
Not automatically. EDPB Opinion 28/2024 states that AI models trained on personal data "cannot, in all cases, be considered anonymous" and that anonymity "should be assessed, based on specific criteria, on a case-by-case basis" [A28]. Paragraph 43 sets the test in two limbs: using reasonable means, both the likelihood of direct or probabilistic extraction of training-set personal data and the likelihood of obtaining such data from queries must be "insignificant for any data subject" [A29]. That's a case-by-case assessment your supervisory authority makes, not a blanket ruling that every model is personal data.
Two separate studies, with separate numbers. In the 2021 GPT-2 study, 1,800 candidate generations yielded 604 confirmed memorised training examples [A2], of which 46 contained individual people's names [A4] and 32 contained contact information [A5]; the authors describe the resulting 0.1% rate as "likely an extremely loose lower bound" [A10]. In the separate 2023 attack on production ChatGPT, about $200 of API queries yielded more than 10,000 unique verbatim training examples [A15], and 16.9% of the hand-labelled generations contained memorised personal information [A16]. The GPT-2 figures describe GPT-2 only — a model trained on public web data, not on anyone's prompts — and the ChatGPT figures describe that 2023 attack only.
Did Italy's EUR 15 million fine against OpenAI stick?
No. The Garante's provvedimento n. 755 of 2 November 2024 imposed the fine and a six-month public information campaign [A76], but a footnote on the Garante's own press release records that the provision was removed from its website following the judgment of the Tribunale di Roma n. 4153/2026, published 18 March 2026, which upheld the opposition against it [A78]. The judgment text itself was not reachable during this research, so the grounds are unknown and the outcome rests on the regulator's own footnote. The fine should not be cited as a live penalty.
Does redacting PII before it reaches the model solve the problem?
No, it reduces exposure without eliminating it. Lukas and colleagues state that scrubbing techniques "reduce but do not prevent the risk of PII leakage" [A19] and conclude that "scrubbing cannot fully prevent PII leakage" [A26]. A membership-inference attack scores AUC 0.96 against an undefended model, 0.82 against a scrubbed model, and 0.505 only once scrubbing is combined with differential privacy [A24]. Treat redaction as a first line of defence rather than an outright prevention [A12].
Does my provider's default of not training on my data mean prompts are safe?
It means the provider currently publishes that default, on a page with a date. OpenAI's Enterprise Privacy page states "By default, we do not use your business data for training our models," with an opt-in exception through feedback mechanisms [A64], and says that except for certain endpoints and features listed in its platform documentation, API inputs and outputs may be retained for up to 30 days "unless we are legally required to retain them" [A65]. That exception became live when a US court order required indefinite retention of consumer ChatGPT and API content, an obligation OpenAI says ended on September 26, 2025 [A68]. Customers on the Zero Data Retention API were unaffected because the data was never stored [A70]. Read the position on your own provider's page, record the date you read it, and treat it as revocable.
Sources
- [A1] Carlini et al. generated 1,800 candidate memorised samples — "100 under each of the 3 × 6 attack configurations" (3 text-generation strategies × 6 membership-inference strategies = 18 configurations) against GPT-2. — https://arxiv.org/pdf/2012.07805v2
- [A2] From those 1,800 candidates the authors identified 604 unique memorised training examples, "for an aggregate true positive rate of 33.5% (our best variant has a true positive rate of 67%)". — https://arxiv.org/pdf/2012.07805v2
- [A3] Adjudication was manual then externally confirmed: "For each of the 1,800 selected samples, one of four authors manually determined whether the sample contains memorized text," then all 1,800 sequences were sent to the GPT-2 authors, who ran "a fuzzy 3-gram match" against the original training dataset. — https://arxiv.org/pdf/2012.07805v2
- [A4] Of the 604 memorised examples, "We find 46 examples that contain individual peoples' names" — explicitly excluding memorised samples relating to national and international news, e.g. famous politicians. — https://arxiv.org/pdf/2012.07805v2
- [A5] "We further find 32 examples that contain some form of contact information (e.g., a phone number or social media handle). Of these, 16 contain contact information for businesses, and 16 contain private individuals' contact details." — https://arxiv.org/pdf/2012.07805v2
- [A6] The attack model is black-box: the paper proposes "a simple and efficient method for extracting verbatim sequences from a language model's training set using only black-box query access", and the authors "treat LMs as black-box generative functions". — https://arxiv.org/pdf/2012.07805v2
- [A8] Larger models memorise more: "in one setting the 1.5 billion parameter GPT-2 model memorizes over 18× as much content as the 124 million parameter model", and the abstract states "Worryingly, we find that larger models are more vulnerable than smaller models." — https://arxiv.org/pdf/2012.07805v2
- [A10] The 604 figure is a floor, not a ceiling: "among 600,000 (honestly) generated samples, our attacks find that at least 604 (or 0.1%) contain memorized text… Note that this is likely an extremely loose lower bound." — https://arxiv.org/pdf/2012.07805v2
- [A11] Memorisation does not require overfitting: "large LMs have no significant train-test gap, and yet we still extract numerous examples verbatim from the training set… there are still some training examples that have anomalously low losses." — https://arxiv.org/pdf/2012.07805v2
- [A12] The paper's own verdict on scrubbing: "Overall, sanitizing data is imperfect — some private data will always slip through — and thus it serves as a first line of defense and not an outright prevention against privacy leaks." — https://arxiv.org/pdf/2012.07805v2
- [A13] The paper was published at the 30th USENIX Security Symposium (2021); author affiliations at publication were Google, Stanford, UC Berkeley, Northeastern, OpenAI, Harvard, Apple. — https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- [A14] Carlini et al.'s 2023 follow-up attacked production ChatGPT: a "divergence attack" causes gpt-3.5-turbo "to diverge from its chatbot-style generations and emit training data at a rate 150× higher than when behaving properly". — https://arxiv.org/abs/2311.17035
- [A15] "Using only $200 USD worth of queries to ChatGPT (gpt-3.5-turbo), we are able to extract over 10,000 unique verbatim-memorized training examples", with a Good–Turing extrapolation suggesting "over 10× more data" is extractable at larger budgets. — https://arxiv.org/pdf/2311.17035
- [A16] PII rate in production-model extraction: "We labeled 15,000 generations for substrings that looked like PII… In total, 16.9% of generations we tested contained memorized PII, and 85.8% of generations that contained potential PII were actual PII." — https://arxiv.org/pdf/2311.17035
- [A17] Alignment/RLHF is not a memorisation fix: the paper's conclusion is that "practical attacks can recover far more data than previously thought, and reveal that current alignment techniques do not eliminate memorization." — https://arxiv.org/abs/2311.17035
- [A18] "Quantifying Memorization Across Neural Language Models" describes "three log-linear relationships": memorisation grows with (1) model capacity, (2) how many times an example was duplicated, and (3) the number of tokens of context used to prompt the model — concluding memorisation "will likely get worse as models continue to scale, at least without active mitigations". — https://arxiv.org/abs/2202.07646
- [A19] Lukas et al. (University of Waterloo + Microsoft) state the premise directly: PII leakage "can be attributed to the false assumption that dataset curation techniques such as scrubbing are sufficient to prevent PII leakage. Scrubbing techniques reduce but do not prevent the risk of PII leakage." — https://arxiv.org/pdf/2302.00539
- [A20] Automated PII scrubbing rests on NER, which is imperfect: "Modern NER is based on the Transformer architecture and has mixed recall of 97% (for names) and 80% (for care unit numbers) on clinical health data, meaning that much PII is retained after scrubbing." — https://arxiv.org/pdf/2302.00539
- [A21] Second recall benchmark cited in the same paper: "a fine-tuned NER model with ground-truth PII annotations achieves recall rates between 84-93%, decreasing to 77% without annotations" (Pilán et al.). — https://arxiv.org/pdf/2302.00539
- [A22] Measured extraction against a fine-tuned model: "GPT-2-Large recalls 23% of PII in the ECHR dataset with a precision of 30%"; GPT-2-Small has "only about 9%" recall at similar precision. — https://arxiv.org/pdf/2302.00539
- [A23] Differential privacy helps but does not close the hole: "sentence-level differential privacy reduces the risk of PII disclosure but still leaks about 3% of PII sequences"; on DP-trained ECHR models "we can extract about 3% of PII with a precision of 3%". — https://arxiv.org/pdf/2302.00539
- [A24] Layered defences measured against membership inference: the attack "achieves an AUC score of 0.96, 0.82, and 0.505 against undefended, scrubbed, and DP & scrubbed models respectively" — i.e. scrubbing alone barely dents it, DP plus scrubbing reduces it to near-chance. — https://arxiv.org/pdf/2302.00539
- [A25] Reconstruction from redacted text is real: "On ECHR and GPT-2-Large, TAB correctly reconstructs at least 5.81% of PII whereas our attack achieves 18.27%", with the authors reconstructing "up to 10× more PII" than prior work by using the suffix of a masked query. — https://arxiv.org/pdf/2302.00539
- [A26] The paper's stated limitation and conclusion: "Since scrubbing cannot remove all PII and we show PII leakage empirically, we conclude that scrubbing cannot fully prevent PII leakage." — https://arxiv.org/pdf/2302.00539
- [A27] Opinion 28/2024 "on certain data protection aspects related to the processing of personal data in the context of AI models" was adopted on 17 December 2024, at the request of the Irish supervisory authority under Article 64(2) GDPR, and answers four questions: when an AI model can be considered anonymous, legitimate interest in development and in deployment, and the consequences of unlawful processing in the development phase. — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A28] Headline finding, paragraph 34: "the EDPB considers that AI models trained on personal data cannot, in all cases, be considered anonymous. Instead, the determination of whether an AI model is anonymous should be assessed, based on specific criteria, on a case-by-case basis." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A29] The two-limb anonymity test, paragraph 43: "for an AI model to be considered anonymous, using reasonable means, both (i) the likelihood of direct (including probabilistic) extraction of personal data regarding individuals whose personal data were used to train the model; as well as (ii) the likelihood of obtaining, intentionally or not, such personal data from queries, should be insignificant for any data subject." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A30] The "it's only weights" argument is addressed head-on, paragraph 31: "information from the training dataset, including personal data, may still remain 'absorbed' in the parameters of the model, namely represented through mathematical objects. They may differ from the original training data points, but may still retain the original information of those data, which may ultimately be extractable or otherwise obtained, directly or indirectly, from the model." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A31] Some models are never anonymous, paragraph 29: models "specifically designed to provide personal data regarding individuals whose personal data were used to train the model… cannot be considered anonymous" — the EDPB's examples are a generative model fine-tuned on one person's voice recordings, and "any model designed to reply with personal data from the training when prompted for information regarding a specific person". — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A35] The EDPB names the attack classes a controller should test against, paragraph 55: "structured testing against: (i) attribute and membership inference; (ii) exfiltration; (iii) regurgitation of training data; (iv) model inversion; or (v) reconstruction attacks" — and warns that "successful testing which covers widely known, state-of-the-art attacks can only be evidence for the resistance to those attacks." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A36] Engineering evidence is explicitly in scope, paragraph 54: SAs "should evaluate whether controllers have conducted any document-based audits (internal or external)… This could include the analysis of reports of code reviews, as well as a theoretical analysis documenting the appropriateness of the measures chosen to reduce the likelihood of re-identification." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A37] Data preparation is an assessed area, paragraph 51: SAs should examine "(i) whether the use of anonymous and/or personal data that has undergone pseudonymisation have been considered; and (ii) where it was decided not to use such measures, the reasons for this decision… (iii) the data minimisation strategies and techniques employed… and (iv) any data filtering processes implemented prior to model training intended to remove irrelevant personal data." Pseudonymisation is listed as a preparation measure, not as a route to anonymity. — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A38] Training methodology is assessed, paragraph 52: whether the methodology "uses regularisation methods to improve model generalisation and reduce overfitting; and, crucially, whether the controller implemented appropriate and effective privacy-preserving techniques (e.g. differential privacy)." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A39] Named deployment-phase mitigations, paragraph 107(a): "Technical measures may for instance be put in place to prevent the storage, regurgitation or generation of personal data, especially in the context of generative AI models (such as output filters), and/or to mitigate the risk of unlawful reuse by general purpose AI models (e.g. digital watermarking of AI-generated outputs)." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A40] Failing the anonymity demonstration has a named consequence, paragraph 57: if a supervisory authority "is not able to confirm… that effective measures were taken to anonymise the AI model, the SA would be in a position to consider that the controller has failed to meet its accountability obligations under Article 5(2) GDPR." — https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- [A43] NIST AI 600-1 is dated July 2024 and was "Approved by the NIST Editorial Review Board on 07-25-2024"; it is a cross-sectoral profile of the AI RMF for Generative AI issued under EO 14110. — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A45] Section 2.4 names memorisation explicitly: "Models may leak, generate, or correctly infer sensitive information about individuals. For example, during adversarial attacks, LLMs have revealed sensitive information (from the public domain) that was included in their training data. This problem has been referred to as data memorization, and may pose exacerbated privacy risks even for data present only in a small number of training samples." — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A46] NIST also flags inference of data that was never in training: "GAI models may be able to correctly infer PII or sensitive data that was not in their training data nor disclosed by the user by stitching together information from disparate sources. These inferences can have negative impact on an individual even if the inferences are not accurate." — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A49] Suggested action MP-4.1-005 tags Data Privacy: "Establish policies for collection, retention, and minimum quality of data, in consideration of the following risks: … Leak of personally identifiable information, including facial likenesses of individuals." — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A50] A further Data Privacy action: "Conduct periodic monitoring of AI-generated content for privacy risks; address any possible instances of PII or sensitive data exposure." — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A51] Governance action MP-4.1-003, tagged Information Security and Data Privacy: "Connect new GAI policies, procedures, and processes to existing model, data, software development, and IT governance and to legal, compliance, and risk management activities." — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- [A55] Recital 26 settles the pseudonymisation question: "Personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person." — http://publications.europa.eu/resource/celex/32016R0679
- [A58] The ICO's "Guidance on AI and data protection" was last updated 15 March 2023 and now carries a standing notice: "Due to changes made by the Data (Use and Access) Act, this guidance is under review and may be subject to change." — https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/
- [A60] ICO definition of membership inference: attacks that "allow malicious actors to deduce whether a given individual was present in the training data of a ML model", exploiting the fact that "if an individual was in the training data, then the model will be disproportionately confident in a prediction about that person because it has seen them before". — https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-should-we-assess-security-and-data-minimisation-in-ai/
- [A64] OpenAI's published position, stated on its Enterprise Privacy page: "Does OpenAI train its models on my business data? By default, we do not use your business data for training our models. If you have explicitly opted in to share your data with us (for example, through our opt-in feedback mechanisms) to improve our services, then we may use the shared data to train our models." — https://openai.com/enterprise-privacy/
- [A65] OpenAI's published API retention window: "OpenAI may securely retain API inputs and outputs for up to 30 days to provide the services and to identify abuse. After 30 days, API inputs and outputs are removed from our systems, unless we are legally required to retain them. You can also request zero data retention (ZDR) for eligible endpoints if you have a qualifying use-case." — https://openai.com/enterprise-privacy/
- [A66] OpenAI's scope statement for default no-training: "By default, data from ChatGPT Business, ChatGPT Enterprise, ChatGPT for Healthcare, ChatGPT Edu, ChatGPT for Teachers, and the API Platform (after March 1, 2023) isn't used for training our models, unless you have explicitly opted in to share your data with us to improve the services." — https://openai.com/enterprise-privacy/
- [A67] OpenAI's fine-tuning statement: "Your fine-tuned models are for your use alone and never served to or shared with other customers or used to train other models. Data submitted to fine-tune a model is retained until the customer deletes the files." — https://openai.com/enterprise-privacy/
- [A68] The 30-day window was overridden by a US court for roughly four months. OpenAI's own update: "After months of litigation, we are no longer under a legal order to retain consumer ChatGPT and API content indefinitely. Our obligations under the earlier order ended on September 26, 2025." The original page is dated 5 June 2025; the update is dated 22 October 2025. — https://openai.com/index/response-to-nyt-data-demands/
- [A69] Who was in scope of that order, in OpenAI's words: "Yes, if you have a ChatGPT Free, Plus, Pro, and Team subscription or if you use the OpenAI API (without a Zero Data Retention agreement). This does not impact ChatGPT Enterprise or ChatGPT Edu customers. This does not impact API customers who are using Zero Data Retention endpoints under our ZDR amendment." — https://openai.com/index/response-to-nyt-data-demands/
- [A70] Why ZDR survived the order — OpenAI's own explanation: "If you are a business customer that uses our Zero Data Retention (ZDR) API, we never retain the prompts you send or the answers we return. Because it is not stored, this court order doesn't affect that data." — https://openai.com/index/response-to-nyt-data-demands/
- [A72] Anthropic's published position: "By default, we will not use your inputs or outputs from our commercial products to train our models." — https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training
- [A73] Anthropic's stated exception and retention: "If you explicitly report feedback or bugs to us (e.g. via our thumbs up/down feedback button), or otherwise choose to allow us to use your data, then we may use your chats and coding sessions to train our models," and where feedback is submitted the company "will store the entire related conversation… in our secured back-end for up to 5 years", de-linked from user identifiers. — https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training
- [A74] OWASP lists Sensitive Information Disclosure as LLM02:2025 in the OWASP Top 10 for LLM Applications: "Sensitive information can affect both the LLM and its application context. This includes personal identifiable information (PII), financial details, health records, confidential business data, security credentials, and legal documents." — https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/
- [A75] OWASP's named mitigations for LLM02:2025 are data sanitization and robust input validation; strict access controls and restricting data sources; federated learning and differential privacy; user education and transparency in data usage; concealing the system preamble and following security-misconfiguration best practice; and homomorphic encryption, tokenization and redaction. — https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/
- [A76] The Italian Garante's provision against OpenAI was provvedimento n. 755 of 2 November 2024, imposing a €15 million fine and ordering a six-month public information campaign on radio, TV, newspapers and the internet, using for the first time the Authority's powers under Article 166(7) of the Italian Privacy Code. — https://www.gpdp.it/garante/doc.jsp?ID=10085432
- [A77] The findings were: failure to notify the March 2023 data breach; processing users' personal data to train ChatGPT "senza aver prima individuato un'adeguata base giuridica" (without first identifying an adequate legal basis); breach of the transparency principle and the related information obligations; and absence of age-verification mechanisms. — https://www.gpdp.it/garante/doc.jsp?ID=10085432
- [A78] The fine is not live. A footnote added to the Garante's own press release states: "Il provvedimento n. 755 del 2 novembre 2024 è stato temporaneamente rimosso dal sito web del Garante per la protezione dei dati personali a seguito della sentenza del Tribunale di Roma n. 4153/2026, pubbl. il 18/03/2026, con la quale è stata accolta l'opposizione proposta avverso il provvedimento del Garante." (Provision no. 755 of 2 November 2024 has been temporarily removed from the Garante's website following the Court of Rome judgment no. 4153/2026, published 18 March 2026, which upheld the opposition brought against the Garante's provision.) — https://www.gpdp.it/garante/doc.jsp?ID=10085432
- [A79] Because OpenAI established its European headquarters in Ireland during the investigation, the Garante transmitted the case file to the Irish DPC under the one-stop-shop rule, the DPC having become lead supervisory authority under the GDPR. — https://www.gpdp.it/garante/doc.jsp?ID=10085432
- [A80] The Garante has remained active on AI enforcement since: its English press room lists, among others, "Artificial intelligence: The Italian Data Protection Authority imposes a fine on Character.AI", "Artificial Intelligence: The Italian Data Protection Authority blocks DeepSeek", and "Artificial intelligence: the Italian Data Protection Authority opens an investigation into OpenAI's 'Sora'". — https://www.gpdp.it/web/garante-privacy-en/press-room
- [A81] Every English-language account of the Samsung/ChatGPT story traces to a single report in The Economist Korea. Cybersecurity Dive attributes the claim to "a report from The Economist Korea" and states: "Samsung Electronics did not respond to requests for comment." No Samsung statement, incident report or regulator filing was located. — https://www.cybersecuritydive.com/news/Samsung-Electronics-ChatGPT-leak-data-privacy/647219/