LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
Installs into .claude/skills of the current project.
Are you the author of Llm Redteam?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/awarexone-llm-redteam)
---
name: llm-redteam
description: LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
---
# LLM Red-Team
> Reflection is not exploitation. A model that *echoes* your payload is noise; a model that *acts* on it — leaks its system prompt, emits your canary, calls a tool, renders raw HTML downstream — is a bug. Prove the action, then chain it to concrete impact.
This is the first-class LLM-attack skill. The scanner behind it is `tools/llm_redteam.py` (canary-based corpus runner, `/llm-redteam`). For auditing an MCP *server*, use `skills/mcp-server-audit`; for attacking a *deployed agent* black-box, use `skills/agentic-app-audit`.
## 0. QUICK KILL CHECKLIST
Kill the lead (classify Informational, move on) when:
- The model refuses and no canary / leak / tool-call ever lands. "It said something edgy" is not a finding.
- System-prompt "leak" is just generic assistant boilerplate ("I am a helpful assistant"), not the target's actual instructions.
- Prompt injection fires but reads/changes nothing the current user couldn't already access (no cross-tenant data, no tool with impact).
- Output looks like XSS but the host app escapes it before rendering (check the actual downstream sink, not the chat response).
- Exfil payload is a markdown image beacon but the client never fetches remote images.
## 1. ROUTING TABLE — signal on target -> what to run
| Signal | Run |
|---|---|
| Any chat/assistant endpoint | `tools/llm_redteam.py --url <endpoint> --field <json-field>` (full corpus) |
| OpenAI-style API | `--template '{"messages":[{"role":"user","content":"{{PAYLOAD}}"}]}' --response-path choices.0.message.content` |
| Only want one class | `--category jailbreak` (see `--list-categories`) |
| Uploads/RAG/"summarize this doc" | indirect-injection + token-smuggling (invisible unicode in the doc) |
| Agent has tools | excessive-agency category + `skills/agentic-app-audit` |
| Confirm blind exfil | plant a canary host, correlate with `tools/oob_listener.py` |
## 2. ATTACK CLASSES (corpus categories)
The runner fires a canary per run (`RT_PWNED_xxxx`); a category "lands" when the canary / a real leak / a tool-call shows up in the response.
- **prompt-injection** — direct instruction override, delimiter breaks.
- **jailbreak** — DAN / developer-mode persona escape.
- **system-prompt-leak** — repeat-above, keyword-anchor, scenario-escape.
- **data-exfil** — markdown-image beacon / canary emission to an attacker host.
- **indirect-injection** — payload framed as a retrieved document / tool result (the high-value one for RAG apps).
- **guardrail-bypass** — base64 / split-string instruction smuggling.
- **multi-turn-crescendo** — single-shot simulation of a gradual escalation jailbreak.
- **cross-lingual** — override instruction in another language to dodge English-only filters.
- **cipher-obfuscation** — leetspeak / ROT13-wrapped instruction.
- **token-smuggling** — invisible "Sneaky Bits" unicode (U+2062/U+2064) carrying a hidden instruction (shared encoder with `tools/hai_payload_builder.py`).
- **excessive-agency** — coerce an unsafe tool call / exfil (SSRF via "fetch URL", email send, payment).
- **insecure-output** — get the model to emit raw HTML/script that a downstream sink renders (XSS via AI).
## 3. CANONICAL OWASP ASI01-ASI10 (single source of truth)
This table is the ONE authoritative ASI mapping for the whole toolkit — `skills/bug-bounty` and `skills/web2-vuln-classes` both defer to it. If you edit the taxonomy, edit it here.
| ID | Class | What to test | Chain to |
|----|-------|--------------|----------|
| ASI01 | Prompt Injection / Goal Hijack | Override objectives via direct or indirect injection | IDOR / exfil / tool abuse |
| ASI02 | Tool Misuse | Attacker-controlled tool params ("fetch this URL", code tool) | SSRF / RCE |
| ASI03 | Privilege Compromise | Agent uses broader perms / admin tokens than the user | Priv-esc / cross-tenant |
| ASI04 | Supply Chain | Compromised plugin / MCP server / tool-output poisoning next agent | RCE / data theft |
| ASI05 | Code Execution | Unsafe code-interpreter / sandbox escape | RCE |
| ASI06 | Memory & Context Poisoning | Persistent RAG/memory corruption across sessions/users | Stored injection affecting all users |
| ASI07 | Agent Communication | Inter-agent spoofing / IDOR (agent A reads agent B's context) | Cross-tenant disclosure |
| ASI08 | Excessive Agency | Destructive action without confirmation; cascading failures | Funds/email/delete |
| ASI09 | Insecure Output Handling | AI output rendered as XSS / SQLi / command injection downstream | XSS / injection |
| ASI10 | Sensitive Information Disclosure | Leaks system prompt / keys / configs / user data | Secrets -> deeper access |
**Triage rule:** ASI alone = Informational. It is a bounty only when chained to IDOR / exfil / RCE / ATO with a working PoC.
## 4. CHAIN TO IMPACT (what makes it payable)
| Standalone | Chain | Result |
|---|---|---|
| Prompt injection | + reads another user's conversation/data (IDOR) | High |
| Indirect injection (RAG doc) | + persists in memory -> hits every user | High/Critical |
| Tool misuse "fetch URL" | + hits internal service / IMDS and returns data | SSRF (Medium/High) |
| Insecure output | + the host renders it -> stored XSS -> session theft | High |
| System-prompt leak | + the prompt contains an API key / internal URL | High |
| Excessive agency | + unconfirmed destructive action (send funds, delete) | High/Critical |
## 5. CONFIRMATION DISCIPLINE (no false positives)
- A category only counts when the **canary / real leak / tool-call** is observed in the response — never on "it sounded compliant."
- For blind exfil, require a correlated out-of-band callback (`tools/oob_listener.py`), not just a rendered beacon tag.
- For insecure-output XSS, confirm execution at the **downstream sink** with the XSS verifier (`tools/verifiers/xss.py` / `tools/dom_xss_harness.py`), not in the chat reply.
- Record the exact payload + response excerpt; the report must be reproducible verbatim.