Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Red Teaming Llms

ASecurity

Red-team LLM systems methodically — scoping, adversarial test design, vulnerability classification, and remediation tracking. Use when you need to find weaknesses before attackers do. Defensive security practice only.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsrustgotestingsecuritydocumentation

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill red-teaming-llms --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Red Teaming Llms?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Red Teaming Llms
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-red-teaming-llms/badge)](https://www.skillsdirectory.com/skills/aicodedecode-red-teaming-llms)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: red-teaming-llms
description: Red-team LLM systems methodically — scoping, adversarial test design, vulnerability classification, and remediation tracking. Use when you need to find weaknesses before attackers do. Defensive security practice only.
category: ai-research
---

# Red-Teaming LLMs

Red-teaming is structured adversarial testing: deliberately trying to make the system fail — 
produce harmful outputs, leak data, bypass controls, get hijacked — so you can fix it before 
deployment. It's a security discipline, not mischief: scoped, documented, and aimed at remediation.

## Overview

A red-team engagement: define scope (what system, what threat model, what's off-limits), design 
adversarial tests across vulnerability classes, execute methodically while documenting everything, 
classify findings by severity, and track remediation to closure. The output isn't a list of 
"gotchas" — it's a risk assessment with prioritized fixes and re-test results.

## When to use

- Before launching any user-facing LLM feature: the pre-deployment security review.
- After significant changes: new tools, new data sources, model updates.
- Periodic assessment of production systems: threats evolve.
- Building an internal AI security practice: establishing the methodology.

## Core concepts

- **Scoping**: the system under test, threat model (who attacks, with what access), and rules of 
engagement. Unauthorized testing is not red-teaming — get explicit permission.
- **Vulnerability classes**: harmful content generation, prompt injection (direct/indirect), data 
exfiltration (training data, system prompts, user data), tool abuse (injection + actions), access 
control bypass, denial of service (resource exhaustion). Cover each systematically.
- **Test design**: adversarial cases per class — from known techniques to creative variants. Keep 
tests private; they're dual-use.
- **Severity rating**: impact × exploitability. A data leak via trivial prompt injection outranks 
an exotic multi-turn bypass with no impact.
- **Documentation**: every test — input, output, classification, severity — recorded. Findings 
without evidence don't get fixed.
- **Remediation loop**: findings → fixes → re-test → closure. A red-team report that nobody 
acts on is theater.

## Practical workflow

1. Get authorization: written scope, threat model, rules of engagement, and handling rules for 
findings.
2. Map the attack surface: inputs, tools, data flows, trust boundaries, deployment context.
3. Design tests per vulnerability class; execute methodically; document every attempt and result.
4. Classify findings by severity with evidence; write the report for the people who'll fix things 
— concrete, prioritized.
5. Track remediation: each finding gets an owner, a fix, and a re-test. Nothing closes without 
verification.
6. Archive tests privately; share defensive learnings (patterns, mitigations) — never attack 
specifics — with the broader team.

```text
Engagement template:
SCOPE:     <system, version, threat model>
RULES:     <authorized by, off-limits, data handling>
SURFACE:   <inputs, tools, data flows mapped>
TESTS:     <per class: cases designed/executed>
FINDINGS:  <severity, evidence, affected component>
REMEDIATE: <owner, fix, re-test result per finding>
STATUS:    <open / remediated / accepted-risk>
```

## Common pitfalls

- **No authorization**: testing systems you weren't asked to test. Get it in writing first — 
always.
- **Gotcha hunting**: collecting clever bypasses without severity assessment or remediation. 
Findings must drive fixes.
- **Publishing attacks**: sharing working exploits publicly. Report to owners; publish defenses and 
patterns only.
- **One-and-done**: a single engagement before launch, never repeated. Threats evolve; so must 
testing.
- **Scope creep**: testing beyond authorization. Stay in scope; expand it formally if needed.
- **No re-test**: fixes assumed to work. Every remediation gets verified with the original test 
plus variants.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →