Skip to content
Back to skills

Llm Redteam

ASecurity

LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.

  • 5,270 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added October 2, 2026
securitygosqlapi

Works with

  • cli
  • api
  • mcp

Security analysis

A100/100

Scanned October 2, 2026

npx -y skills add awarexone/Agentic-Bug-Hunter --skill llm-redteam --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Llm Redteam?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Llm Redteam
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/awarexone-llm-redteam/badge)](https://www.skillsdirectory.com/skills/awarexone-llm-redteam)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: llm-redteam
description: LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
---

# LLM Red-Team

> Reflection is not exploitation. A model that *echoes* your payload is noise; a model that *acts* on it — leaks its system prompt, emits your canary, calls a tool, renders raw HTML downstream — is a bug. Prove the action, then chain it to concrete impact.

This is the first-class LLM-attack skill. The scanner behind it is `tools/llm_redteam.py` (canary-based corpus runner, `/llm-redteam`). For auditing an MCP *server*, use `skills/mcp-server-audit`; for attacking a *deployed agent* black-box, use `skills/agentic-app-audit`.

## 0. QUICK KILL CHECKLIST

Kill the lead (classify Informational, move on) when:

- The model refuses and no canary / leak / tool-call ever lands. "It said something edgy" is not a finding.
- System-prompt "leak" is just generic assistant boilerplate ("I am a helpful assistant"), not the target's actual instructions.
- Prompt injection fires but reads/changes nothing the current user couldn't already access (no cross-tenant data, no tool with impact).
- Output looks like XSS but the host app escapes it before rendering (check the actual downstream sink, not the chat response).
- Exfil payload is a markdown image beacon but the client never fetches remote images.

## 1. ROUTING TABLE — signal on target -> what to run

| Signal | Run |
|---|---|
| Any chat/assistant endpoint | `tools/llm_redteam.py --url <endpoint> --field <json-field>` (full corpus) |
| OpenAI-style API | `--template '{"messages":[{"role":"user","content":"{{PAYLOAD}}"}]}' --response-path choices.0.message.content` |
| Only want one class | `--category jailbreak` (see `--list-categories`) |
| Uploads/RAG/"summarize this doc" | indirect-injection + token-smuggling (invisible unicode in the doc) |
| Agent has tools | excessive-agency category + `skills/agentic-app-audit` |
| Confirm blind exfil | plant a canary host, correlate with `tools/oob_listener.py` |

## 2. ATTACK CLASSES (corpus categories)

The runner fires a canary per run (`RT_PWNED_xxxx`); a category "lands" when the canary / a real leak / a tool-call shows up in the response.

- **prompt-injection** — direct instruction override, delimiter breaks.
- **jailbreak** — DAN / developer-mode persona escape.
- **system-prompt-leak** — repeat-above, keyword-anchor, scenario-escape.
- **data-exfil** — markdown-image beacon / canary emission to an attacker host.
- **indirect-injection** — payload framed as a retrieved document / tool result (the high-value one for RAG apps).
- **guardrail-bypass** — base64 / split-string instruction smuggling.
- **multi-turn-crescendo** — single-shot simulation of a gradual escalation jailbreak.
- **cross-lingual** — override instruction in another language to dodge English-only filters.
- **cipher-obfuscation** — leetspeak / ROT13-wrapped instruction.
- **token-smuggling** — invisible "Sneaky Bits" unicode (U+2062/U+2064) carrying a hidden instruction (shared encoder with `tools/hai_payload_builder.py`).
- **excessive-agency** — coerce an unsafe tool call / exfil (SSRF via "fetch URL", email send, payment).
- **insecure-output** — get the model to emit raw HTML/script that a downstream sink renders (XSS via AI).

## 3. CANONICAL OWASP ASI01-ASI10 (single source of truth)

This table is the ONE authoritative ASI mapping for the whole toolkit — `skills/bug-bounty` and `skills/web2-vuln-classes` both defer to it. If you edit the taxonomy, edit it here.

| ID | Class | What to test | Chain to |
|----|-------|--------------|----------|
| ASI01 | Prompt Injection / Goal Hijack | Override objectives via direct or indirect injection | IDOR / exfil / tool abuse |
| ASI02 | Tool Misuse | Attacker-controlled tool params ("fetch this URL", code tool) | SSRF / RCE |
| ASI03 | Privilege Compromise | Agent uses broader perms / admin tokens than the user | Priv-esc / cross-tenant |
| ASI04 | Supply Chain | Compromised plugin / MCP server / tool-output poisoning next agent | RCE / data theft |
| ASI05 | Code Execution | Unsafe code-interpreter / sandbox escape | RCE |
| ASI06 | Memory & Context Poisoning | Persistent RAG/memory corruption across sessions/users | Stored injection affecting all users |
| ASI07 | Agent Communication | Inter-agent spoofing / IDOR (agent A reads agent B's context) | Cross-tenant disclosure |
| ASI08 | Excessive Agency | Destructive action without confirmation; cascading failures | Funds/email/delete |
| ASI09 | Insecure Output Handling | AI output rendered as XSS / SQLi / command injection downstream | XSS / injection |
| ASI10 | Sensitive Information Disclosure | Leaks system prompt / keys / configs / user data | Secrets -> deeper access |

**Triage rule:** ASI alone = Informational. It is a bounty only when chained to IDOR / exfil / RCE / ATO with a working PoC.

## 4. CHAIN TO IMPACT (what makes it payable)

| Standalone | Chain | Result |
|---|---|---|
| Prompt injection | + reads another user's conversation/data (IDOR) | High |
| Indirect injection (RAG doc) | + persists in memory -> hits every user | High/Critical |
| Tool misuse "fetch URL" | + hits internal service / IMDS and returns data | SSRF (Medium/High) |
| Insecure output | + the host renders it -> stored XSS -> session theft | High |
| System-prompt leak | + the prompt contains an API key / internal URL | High |
| Excessive agency | + unconfirmed destructive action (send funds, delete) | High/Critical |

## 5. CONFIRMATION DISCIPLINE (no false positives)

- A category only counts when the **canary / real leak / tool-call** is observed in the response — never on "it sounded compliant."
- For blind exfil, require a correlated out-of-band callback (`tools/oob_listener.py`), not just a rendered beacon tag.
- For insecure-output XSS, confirm execution at the **downstream sink** with the XSS verifier (`tools/verifiers/xss.py` / `tools/dom_xss_harness.py`), not in the chat reply.
- Record the exact payload + response excerpt; the report must be reproducible verbatim.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…