Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Evals Ci Gate

ASecurity

Operationalize a safety/prompt-injection eval suite into an enforced CI gate — not just a one-off report — using the ready-to-copy promptfoo/garak template, a regression baseline, and a burn-in rollout. Use when a safety eval already exists (or is being designed) and needs to actually block regressions on every release rather than being run manually once.

8 stars
0 votes
0 copies
1 views
Added 9/19/2026
ai-agentsgogitsecurity

Security Analysis

A100/100

Scanned 9/19/2026

$npx -y skills add jassics/awesome-claude-security --skill evals-ci-gate --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evals Ci Gate?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Evals Ci Gate
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/jassics-evals-ci-gate/badge)](https://www.skillsdirectory.com/skills/jassics-evals-ci-gate)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: evals-ci-gate
description: >-
  Operationalize a safety/prompt-injection eval suite into an enforced CI gate —
  not just a one-off report — using the ready-to-copy promptfoo/garak template,
  a regression baseline, and a burn-in rollout. Use when a safety eval already
  exists (or is being designed) and needs to actually block regressions on every
  release rather than being run manually once.
---

# Goal

A working, tuned CI gate that fails releases on a real safety/injection
regression — durable and enforced, not a stale one-time eval someone read once.

# Steps

1. **Pull in the test material.** Eval set, harm categories, and rubrics come
   from `ai-safety:safety-evaluation`; the prompt-injection payload taxonomy for
   the injection-resistance cases comes from `llm-security:prompt-injection-test`.
   This skill doesn't design the tests — it makes them an enforced gate.
2. **Install the gate.** Copy `templates/genai-eval-gates/` into the target repo
   (`promptfoo.config.yaml` + `.github/workflows/genai-eval-gate.yml`). Point the
   config's `providers` section at the real model/endpoint and replace the
   illustrative test cases with the real eval set from step 1.
3. **Establish the regression baseline.** Run the suite once, review results by
   hand, fix any ambiguous/poorly-worded rubrics, then commit the run as
   `eval-baseline.json` — every future run is compared against this, not against
   an absolute pass-rate target (LLM outputs have natural run-to-run variance).
4. **Roll out in report-only mode first**, then flip to blocking, per the
   burn-in stages in `templates/genai-eval-gates/EVAL-GATE-NOTES.md` — do not
   make this a required check on day one.
5. **Wire the decision, not just the number.** Point `ai-safety-engineer:safety-case`
   at this gate's *live* pass rate as ongoing evidence for its safety argument,
   instead of citing a stale one-time eval run.

# Output

A short description of the installed gate: what it blocks on (regression vs.
baseline, tolerance used), the current baseline's date/version, and which
rollout stage it's at (report-only / tightening / blocking).

# Notes

This is the operationalization step, not the test design — that lives in
`ai-safety:safety-evaluation` and `ai-safety:safety-red-team`. A gate that's
green because the eval set or baseline has gone stale is worse than an honest
red; keep the eval set versioned and revisit the baseline whenever the model,
prompt, or eval set changes meaningfully.

Attribution

jassicsjassics
View sourceSee grades on GitHubMore from jassics →
SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

696481 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes
View all in ai-agents →