Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ai Safety

ASecurity

Practice AI safety across the lifecycle — risk assessment, alignment concepts, evaluation, deployment safeguards, and governance. Use when building or deploying AI systems responsibly.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgorailstesting

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill ai-safety --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Safety?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ai Safety
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-ai-safety/badge)](https://www.skillsdirectory.com/skills/aicodedecode-ai-safety)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: ai-safety
description: Practice AI safety across the lifecycle — risk assessment, alignment concepts, evaluation, deployment safeguards, and governance. Use when building or deploying AI systems responsibly.
category: ai-research
---

# AI Safety

AI safety is the practice of building systems that do what we intend, don't cause unintended harm, 
and remain under meaningful human control. It spans technical work (alignment, evaluation, 
robustness) and governance (policies, oversight, accountability) — this skill covers both at a 
practitioner level.

## Overview

Think in layers. Technical: train models to follow intent faithfully (alignment), test for 
dangerous capabilities and failure modes (evaluation), harden against adversaries (robustness), and 
keep humans meaningfully in control of consequential actions (oversight). Organizational: risk 
assessments before deployment, clear ownership, incident response plans, and honest communication 
about limitations. Safety isn't a feature added at the end — it's a property of how you build.

## When to use

- Planning any AI deployment: assessing risks before building.
- Designing safeguards for agents with real-world actions.
- Evaluating models for dangerous capabilities or problematic behaviors.
- Establishing team practices: review processes, incident response, accountability.

## Core concepts

- **Risk assessment**: identify hazards (what could go wrong?), assess likelihood and severity, and 
decide mitigations before deployment. Proportional to capability and autonomy — more powerful 
systems need more rigor.
- **Alignment**: the technical problem of models pursuing intended goals — instruction following, 
honesty, and not optimizing proxies in harmful ways. Practically: careful training objectives, 
evals for deceptive or sycophantic behavior, and humility about unsolved problems.
- **Capability evaluation**: testing what the system can do — including misuse-relevant 
capabilities — before deployment. Know your system's powers; don't discover them via incident.
- **Oversight and control**: humans approve consequential actions; systems are monitorable and 
interruptible; no autonomous operation beyond the assessed risk envelope.
- **Defense in depth**: no single safeguard suffices — combine training, evals, input/output 
controls, tool limitations, monitoring, and human gates.
- **Governance**: ownership of safety decisions, documented risk assessments, incident response 
plans, and channels for raising concerns. Safety needs a name attached.

## Practical workflow

1. Assess: for the planned system, list hazards, rate likelihood × severity, and identify the top 
risks.
2. Design mitigations per risk: technical (evals, guardrails, limits) and procedural (approvals, 
monitoring, rollback).
3. Evaluate before deployment: capability tests, adversarial tests, and failure-mode analysis — 
with acceptance criteria defined up front.
4. Deploy gradually: limited rollout, close monitoring, kill switches and rollback ready.
5. Monitor in production: anomaly detection, user reports, periodic re-evaluation as the system and 
environment change.
6. Learn: incident reviews feed back into assessments; update the risk picture as capabilities grow.

```text
Safety review template:
SYSTEM:    <what it does, what it can affect>
HAZARDS:   <what could go wrong — brainstormed broadly>
TOP RISKS: <likelihood × severity, top 5>
MITIGATIONS:<per risk: technical + procedural>
EVALS:     <acceptance criteria + results>
DEPLOY:    <gradual rollout plan + monitoring + kill switch>
OWNER:     <named person accountable>
```

## Common pitfalls

- **Safety as a checklist**: going through motions without engaging with actual risks. The 
assessment must be honest to be useful.
- **Capability blindness**: not knowing what your system can do until it does it publicly. Evaluate 
first.
- **No kill switch**: deployed systems without a way to stop them. Always have one, tested.
- **Autonomy creep**: gradually expanding what the system does without re-assessing. Re-assess on 
every expansion.
- **Diffused accountability**: "the team" owns safety, so nobody does. Name an owner.
- **Optimism about alignment**: assuming the model shares your goals because it usually behaves. 
Test adversarially; monitor continuously.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →