An engagement against Aria produces this: the tester planted a document in the corpus and got the agent to include a customer's email address in a search query sent to the external index.
Sort it. No irreversible action ran. The gate was never involved, because searching is not a consequential action. The capability was correctly scoped. Provenance was tracked correctly throughout.
It is still a finding. Chapter 10 identified the search index as an outbound channel, and the egress budget was set but the provenance check on query content was not applied to that tool. Private data left the perimeter through a tool nobody thinks of as sending anything.
The class matters more than the instance. The question it raises is: which other tools take model-composed arguments that leave our network? The answer will be a list, and the list is the actual fix. Patching the search tool alone means paying for this finding again with a different tool in eighteen months.
This is the step that makes the engagement worth repeating.
Each finding goes into chapter 15's corpus as a case, with the channel it arrived through and the tool it aimed at. It then runs on every build forever. The engagement's value is not the report, which ages in weeks. It is the permanent addition to the suite.
A team that does this has a corpus that grows with each engagement and encodes everything anyone has ever found. A team that does not will pay for the same findings again in eighteen months.
Once a year is a compliance exercise. The useful rhythm is tied to change rather than to the calendar.
Run one after each part of the architecture lands, while the design is fresh and the fixes are cheap. A new untrusted channel deserves another, because that is a surface the corpus does not cover yet. So does an irreversible action arriving in the tool register.
Between those, the two-engineer version above costs an afternoon and catches most of what a scheduled engagement would have found six months later.
Money, and calendar time to arrange access to the channels that matter.
The subtler cost is that a good engagement produces findings faster than you can fix them, and the backlog is demoralising. Prioritise by blast radius rather than by count: a finding that reaches an irreversible action outranks ten that produce wrong text, whatever the severity labels in the report say.
The other risk is buying the wrong thing. "AI red teaming" is sold as a product by vendors whose actual offering is automated jailbreak probing against a model endpoint. That is a useful thing and it is not this. If the engagement does not involve planting content in your corpus, it is testing the model, not your system.
Chapter 17 assumes all of this failed and something happened anyway.
| Claim | Source | Status |
|---|---|---|
| Adversarial testing and attack simulation as an OWASP-listed mitigation | https://genai.owasp.org/llmrisk/llm01-prompt-injection/ | PRIMARY |
| Coding agents as the concentration of agentic advisories | https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/ | SECONDARY |
| Offensive agent tooling as an established practice area | AI Agents for Offensive Security (Manning), https://www.manning.com/books/ai-agents-for-offensive-security | SECONDARY |
The three differences from conventional penetration testing, the recommendation to hand over the architecture while withholding the corpus, the sort-by-primitive method for reading a report, and the warning about vendors selling model probing as agent red teaming are all the author's, drawn from the mechanics of the architecture in this book rather than from published engagement data. No measured claims appear in this chapter.
Download the full PDF for free?
Free download — no account required