Embedding inversion recovering 50–70% of original input words; near-optimal reconstruction with ~1,000 samples against black-box encoders. https://www.promptfoo.dev/lm-security-db/vuln/embedding-inversion-privacy-leak-bd566a31/ Primary: Transferable Embedding Inversion Attack, arXiv 2406.10280. https://arxiv.org/pdf/2406.10280
OWASP Top 10 for LLM Applications — vector and embedding weaknesses. https://genai.owasp.org/initiatives/top-10-for-llm-and-genai/ · 2025 edition PDF: https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf Numbered LLM08:2025 and renumbered LLM09:2026. The apparent disagreement between sources is an edition change, not an error; the book gives both.
Cross-context retrieval in multi-tenant vector stores. https://www.protecto.ai/blog/how-rag-systems-sliently-expose-pii/ · https://christian-schneider.net/blog/rag-security-forgotten-attack-surface/ (both practitioner/vendor)
Prompt injection leaking personal data observed by agents during task execution, arXiv 2506.01055. https://arxiv.org/pdf/2506.01055 Data exfiltration via backdoored tool use, arXiv 2604.05432. https://arxiv.org/pdf/2604.05432
Guardrail classifiers and malicious-prompt success rates — a ">90% reduction" figure circulates widely in secondary reporting. The primary result could not be located at verification, so no figure is quoted anywhere in this book. Chapter 15 makes the point qualitatively and says why.
EchoLeak, CVE-2025-32711 — zero-click data exfiltration from Microsoft 365 Copilot via crafted email, reference-style markdown bypassing link redaction, and auto-fetched images. https://wraith.sh/learn/markdown-image-exfiltration (narrative account; cite the CVE record itself for the vulnerability)
Extracting memorized training data via decomposition, arXiv 2409.12367. https://arxiv.org/pdf/2409.12367 Towards more realistic extraction attacks: an adversarial perspective, TACL. https://direct.mit.edu/tacl/article/doi/10.1162/TACL.a.62/134536/ Thousands of email addresses, phone numbers and URLs extracted; combining attacks roughly doubles extraction risk, persisting under deduplication. Treated in Chapter 3 as a training-time risk and therefore a contract question for API consumers.
GDPR — Article 5(1)(c) data minimisation; Recital 26 on pseudonymised versus anonymous data. Statutory.
Italian Garante €15M fine against OpenAI (decision November 2024, published December 2024). https://www.lewissilkin.com/en/insights/2025/01/14/openai-faces-15-million-fine-as-the-italian-garante-strikes-again-102jtqc
Court of Rome annulled that fine, 18 March 2026. https://www.crossborderdataforum.org/generative-ai-and-gdpr-enforcement-in-europe-a-lot-of-noise-one-fine-zero-survivors/ Chapter 16 uses this pairing deliberately. The original fine is widely cited; the annulment rarely is.
EDPB opinion on training AI models using personal data. https://www.dataprotectionreport.com/2025/01/the-edpb-opinion-on-training-ai-models-using-personal-data-and-recent-garante-fine-lawful-deployment-of-llms/
EU AI Act — application dates. Article 113: https://artificialintelligenceact.eu/article/113/ Chapters I–II from 2 February 2025. Chapter III §4, Chapter V, Chapter VII, Chapter XII and Article 78 from 2 August 2025, expressly excluding Article 101 (fines on providers of general-purpose AI models), which therefore applies from the general date of 2 August 2026. Article 50 transparency duties applicable 2 August 2026. GPAI models already on the market transition to 2 August 2027.
Digital Omnibus on AI — Regulation (EU) 2026/1744. ADOPTED, not a proposal. European Parliament approved 16 June 2026; Council adopted 29 June 2026; signed 8 July 2026; published in the Official Journal 24 July 2026; in force 27 July 2026. High-risk regime deferred to 2 December 2027 (standalone Annex III) and 2 August 2028 (Annex I, AI embedded in products already covered by EU product-safety law). Article 50 not amended, apart from a four-month grace period to 2 December 2026 on the Article 50(2) content-marking duty for systems placed on the market before 2 August 2026. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/ · https://www.joneswalker.com/en/insights/blogs/ai-law-blog/yes-august-2-still-matters-the-eu-approved-a-high-risk-ai-delay-but-most-trans.html ⚠️ The widely-linked artificialintelligenceact.eu Article 113 page still shows the pre-Omnibus 2 August 2027 date for Annex I. The adopted regulation says 2 August 2028. Verified 14 September 2026.
HIPAA de-identification: Safe Harbor and Expert Determination. HHS, Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule: https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html Regulatory text: 45 CFR §164.514(b) — Expert Determination at (b)(1), Safe Harbor at (b)(2). HHS sets no numeric threshold for the "very small" risk standard; the expert defines it for the dataset and release environment and retains the justification. The 18 identifiers in Appendix A are reproduced from the standard.
Uniqueness of simple demographics. L. Sweeney, Simple Demographics Often Identify People Uniquely, Carnegie Mellon Data Privacy Working Paper 3, 2000: https://dataprivacylab.org/projects/identifiability/paper1.pdf — 87% unique on ZIP + gender + full date of birth, 1990 Census data. P. Golle, Revisiting the Uniqueness of Simple Demographics in the US Population, ACM WPES 2006: https://crypto.stanford.edu/~pgolle/papers/census.pdf — 63% on the same combination using 2000 Census data. Chapter 2 quotes the lower, more recent figure in the prose and names both. The 87% number is one of the most-repeated statistics in privacy and is almost always cited without its 1990 date, which is exactly the failure mode Chapter 4 is about.
Download the full PDF for free?
Free download — no account required