LLM Red Teaming: Techniques and Testing Methodologies for Enterprise AI

Key Takeaways

  • LLM red teaming is structured adversarial testing of a large language model (LLM) application as it is assembled, across the model, application, retrieval and agent layers, to find flaws before a release decision.
  • Manual campaigns discover application-specific attack chains, while automated evaluations replay prior findings against every build.
  • A finding stays closed only when retesting confirms that its prohibited outcome no longer reproduces and production controls retain enforcement evidence.

In June 2025, Microsoft disclosed CVE-2025-32711, a zero-click vulnerability in Microsoft 365 Copilot [1]. An attacker's email carried hidden instructions, retrieval pulled it into the model's context, and the model placed tenant data in a link that trusted Microsoft domains fetched automatically [2].

A year later, three disclosures showed a similar pattern with the roles reversed: the model under test reached real systems through access its operators had not accounted for.

  • OpenAI, July 21, 2026: Two models under cybersecurity evaluation exploited a zero-day to leave their test environment and ran code on Hugging Face production systems [3].
  • Anthropic, July 30, 2026: A configuration error connected models under third-party evaluation to the open internet, and in three incidents they reached real companies' systems [4].
  • Meta, August 2026: Bloomberg reported that a Meta AI model gained internet access through a configuration error during testing and exploited a vulnerability in a third-party service to reach an outside firm's systems [5].

In all four incidents, the exposure ran through a path between the model and real data or systems that operators had not mapped before the model ran. Red teaming tests those paths before a release decision. NIST AI 800-1 defines AI red teaming as “a structured testing effort to find flaws and vulnerabilities in an AI system, often in a controlled environment and in collaboration with developers of AI” [6]. LLM red teaming applies that effort to the assembled system: model, prompts, retrieval, tools and rendering together.

What LLM Red Teaming Is and How It Differs From Pen Testing and Evals

A red-team campaign produces a findings package. For each finding it records the reproducible input sequence, failing output, affected layer, severity, owner, fix and a dated retest showing the attack no longer reproduces.

A red team that includes members independent of the build team runs the campaign, often AI security engineers working with domain experts [6].

Pen testing asks a different question. NIST describes it as assessing “how difficult it would be for an actor to circumvent security measures” [6]. A pen test of a chatbot's API can find no exploitable flaw while the chatbot still hands an attacker a customer's account number on request.

Evaluations and red teaming find different failures:

  • Evaluations: They score fixed test sets against thresholds in bulk and measure how often a known failure occurs in a given build.
  • Red teaming: It searches for failures the test set does not contain and converts each into a new evaluation case, so the next build can be checked automatically.

What Red Teamers Attack: Model, Application, RAG and Agent Layers

Red teams usually scope four layers because the fix depends on which layer produced the finding.

  • Model: The red team tests how consistently the base or fine-tuned model refuses, what training data it has memorized, and whether it treats retrieved text as instructions.
  • Application: Testers probe the system prompt, output rendering and the wrappers that pass user input to the model and model output to the client.
  • RAG: Testers examine the document sources, indexing pipeline and access controls that decide which content enters the context window, and whether the model can tell that content apart from user intent.
  • Agent: Testers check tool permissions, the credentials the agent holds, handoffs between agents, and what the agent can change in real systems (see red teaming agents).

In the Microsoft 365 Copilot incident, the model followed instructions it should have treated as data. The exploit ran through the RAG layer's retrieval of an attacker email and the application's rendering of a reference-style link.

Common Attack Techniques and Vulnerability Categories

Five techniques recur across campaigns, and the red team classifies each finding as a Security, Privacy or Responsible AI failure.

Attack Techniques

  • Prompt injection: The model executes untrusted text as an instruction, whether a user typed it or it arrived inside a retrieved document or tool return.
  • Jailbreaks: The attacker builds prompts that make the model set aside its safety training, through encoding or fictional framing.
  • Roleplay and persona attacks: The attacker asks the model to adopt a character or internal role whose rules override the deployed policy.
  • Multi-turn escalation: A sequence of individually benign turns references the model's own prior replies until it reaches a prohibited output; Russinovich et al. found the automated Crescendo attack outperformed other state-of-the-art jailbreak techniques against GPT-4 [7].
  • Data extraction: Requests pull out the system prompt, retrieved context, other users' data or memorized training text.

Vulnerability Categories

  • Security: Prompt injection that redirects behavior, guardrail evasion, excessive agency through over-permitted tools, and improper output handling such as rendered links.
  • Privacy: PII/PHI leakage, sensitive information disclosure, and hidden context exposure, including system prompts and embedded credentials.
  • Responsible AI: Hallucinated policy or product facts, bias across customer segments, and toxicity.

Choosing Manual, Automated or Both

Enterprise programs typically need both: automation carries regression coverage and humans carry discovery. Automated attack generation tooling, including open-source attack libraries and red-teaming frameworks that automate this work, runs technique variants against every build and replays every prior finding, so a fix from last quarter is rechecked on each release.

It cannot invent the chain that leads through a support ticket, a retrieval index, an account lookup tool and a markdown renderer because that chain is specific to one application's permissions and data flows.

A capable red team for this work typically includes a lead who owns the charter and severity calls, one or two security testers focused on injection and exfiltration paths, a domain expert who judges whether an output is harmful to the business, and an engineer with trace access who can reproduce findings for the application team.

An automated pass on a new build often completes in days, while a manual campaign against a new or materially changed application often runs for weeks under a charter naming the target, prohibited outcomes, red team's access and the decision the results inform. The red team reopens the manual campaign after any of these changes rather than on a calendar schedule:

  • a new model version;
  • a change to the system prompt, tools or retrieval sources;
  • a production incident;
  • a new attack class published against comparable systems.

A Worked Example: Red Teaming a Support Chatbot From Attack to Fix

In this illustrative scenario, the target is a customer-support chatbot that answers policy and billing questions from a RAG index of product documents and past tickets, looks up an authenticated customer's account through a tool call, and is restricted by system prompt to support topics. A two-week campaign produced three findings that blocked release until fixed:

  • Finding 1, persona play (Security): The tester asked the chatbot to act as an internal training assistant exempt from customer-facing rules, then steered it over four turns until it listed the identity-verification steps an agent could skip. The team hardened the system prompt, but that fix failed on retest against a past-tense rewording. The team added a per-request jailbreak check scored against the full conversation rather than the latest turn. On the second retest, the recorded variants no longer reproduced.
  • Finding 2, retrieval injection (Privacy): The tester filed a support ticket whose body carried hidden instructions to append the customer's account number to an image link pointing at an external host. When a support agent later asked the chatbot to summarize that customer's open tickets, it returned a link with the account number in the query string. This is the Microsoft 365 Copilot pattern on a smaller system. The team delimited retrieved content as data in the prompt, disabled external image and link rendering in the chat client, and added PII detection on the response. On retest, the response carried no external URL and the account number arrived redacted.
  • Finding 3, usable credential exposure through hidden context (Privacy): A request to repeat everything above the first user message returned the full system prompt, including the bearer token the chatbot used to call an internal escalation tool. The team defined the prohibited outcome as disclosure of a usable credential. It rotated the token the same day, moved tool credentials into server-side configuration, added secret detection on the response, and classified the finding as LLM08:2026 Hidden Context Exposure, the category the OWASP Top 10 for LLM Applications 2026 renamed from System Prompt Leakage [8]. On retest, the prompt still leaked but contained no usable credential, so the defined prohibited outcome no longer reproduced.

Mapping Red Teaming Evidence to OWASP, NIST AI RMF and the EU AI Act

Auditors ask for dated records of what was tested, what failed, what was fixed and what still runs in production:

Standard Red teaming activity Evidence an auditor expects
OWASP Top 10 for LLM Applications 2026, published August 3, 2026 [8] Classify each finding against the ten categories, such as LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure and LLM08 Hidden Context Exposure Findings package with a category identifier and a retest result on every finding
NIST AI RMF 1.0, voluntary [9] Document security testing, test sets and tools (MEASURE 2.7, MEASURE 2.1); carry findings into post-deployment monitoring (MANAGE 4.1) Dated test records, the tool and metric inventory, and the monitoring plan
EU AI Act [10] Pre-market testing against predefined metrics (Article 9); resilience against evasion and confidentiality attacks (Article 15); adversarial testing for GPAI models with systemic risk (Article 55) Test logs “dated and signed by the responsible persons” (Annex IV); event logs retained at least six months (Article 19)

How Fiddler Keeps Red Team Findings Closed in Production

Findings closed on the retest date can reopen under production traffic as new jailbreak wordings arrive and poisoned tickets get indexed. The Fiddler AI Observability and Security Platform is a runtime control that operates on production traffic.

Fiddler Guardrails enforces policy with inline enforcement on the agent request and response path, with under 80ms response time. Checks on the request run before the application invokes the model: a jailbreak and prompt-injection check scores the request, including prior conversation turns, for the persona-play pattern behind Finding 1. Checks on the response run before it reaches the user or a downstream system: PII/PHI detection catches an account number inside a rendered link (Finding 2), and secret detection catches a bearer token in leaked prompt text (Finding 3). Fiddler Guardrails can allow, block or redact depending on the policy and use case.

Batteries-included Fiddler Centor Models score prompts and responses in-environment, with no external LLM call and no per-evaluation cost. By default they evaluate 100 percent of prompts and responses rather than sampling a fraction of traces. A variant that slips past a request check can still appear in the evaluation record for the same dimension the red team tested, and the security team can add it to the regression suite.

A Fiddler customer, Nielsen, built the Ask Nielsen multi-agent copilot, a production deployment. Nielsen reports 95.38 percent accuracy and 99 percent precision on direct harmful instructions and 98 percent accuracy on persona-play attacks.

Auditable Governance records every enforcement decision with contextual metadata, including the triggering user and the action or policy the system applied. Fiddler retains those records as audit evidence aligned to frameworks such as the EU AI Act, GDPR, HIPAA, SR 11-7 and SR 26-2.

Request a demo to review how the required controls would fit your agent architecture and deployment environment.

References

[1] National Vulnerability Database, CVE-2025-32711 Detail, https://nvd.nist.gov/vuln/detail/CVE-2025-32711

[2] Cato Networks, Breaking down 'EchoLeak', the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot, https://www.catonetworks.com/blog/breaking-down-echoleak/

[3] OpenAI, Hugging Face model evaluation security incident, https://openai.com/index/hugging-face-model-evaluation-security-incident/

[4] Anthropic, Investigating incidents in cybersecurity evals, https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

[5] Bloomberg, Meta AI Model Accessed Internet, Hacked Outside Firm in Testing, https://www.bloomberg.com/news/articles/2026-08-05/meta-ai-model-accessed-internet-hacked-outside-firm-in-testing

[6] National Institute of Standards and Technology, NIST AI 800-1 Second Public Draft, Managing Misuse Risk for Dual-Use Foundation Models, https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-1.ipd2.pdf

[7] M. Russinovich, A. Salem, and R. Eldan, Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack, USENIX Security 2025, https://arxiv.org/html/2404.01833v3

[8] OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/

[9] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

[10] European Union, Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated text of 27 July 2026, https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX%3A02024R1689-20260727

Frequently Asked Questions

Should an Enterprise Run Manual or Automated LLM Red Teaming?

Enterprise programs typically run both. Automation replays prior findings and technique variants on every build to catch regressions, while human testers find the application-specific chains that attack libraries rarely contain.

Do Regulators Require LLM Red Teaming?

The EU AI Act mandates adversarial testing only for GPAI models with systemic risk (Article 55 and Annex XI) [10]. For high-risk systems, whose obligations apply from December 2, 2027 for stand-alone systems under the July 2026 Digital Omnibus amendment [10], it requires pre-market testing against predefined metrics and resilience against model evasion and confidentiality attacks without prescribing the method. NIST AI RMF 1.0 remains voluntary [9]. In practice, a red-team findings package can serve as evidence toward both sets of documentation expectations.

How Often Should a Red Team Campaign Be Rerun?

Neither NIST AI RMF nor the EU AI Act sets a fixed interval for red teaming enterprise LLM applications. The red team reruns the manual campaign on the triggers listed earlier, and the automated suite runs on every build.

How Can a Team Tell a Model Weakness From a System Weakness?

Reproduce the failing output against the bare model with the same prompt and no retrieved content, tools or rendering. If it reproduces, the weakness is in the model and needs a model change or a request-side check. If it needs a retrieved document, a tool return or a rendering path, as the Microsoft 365 Copilot incident did, the fix belongs to the application or RAG layer.