Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Coval Compare Runs

ASecurity

Compare before/after Coval runs using matched cases, stable metric versions and inspected conversation evidence. Use to assess a regression or proposed improvement; does not automatically launch or tune anything.

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
businessgoapi

Works with

terminalcliapi

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add coval-ai/coval-external-skills --skill coval-compare-runs --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Coval Compare Runs?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Coval Compare Runs
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/coval-ai-coval-compare-runs/badge)](https://www.skillsdirectory.com/skills/coval-ai-coval-compare-runs)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: coval-compare-runs
description: Compare before/after Coval runs using matched cases, stable metric versions and inspected conversation evidence. Use to assess a regression or proposed improvement; does not automatically launch or tune anything.
---

# Compare runs without fooling yourself

Answer whether the observed change supports the customer's decision. Stay
read-only. Do not start a hill-climb, create variants or launch confirmation
calls unless explicitly requested within a budget.

## Freeze the comparison

Confirm organization/workspace, API environment, baseline and candidate run IDs,
the intended change, primary criterion and acceptable regressions. Read each
run and its individual conversations, selected metric outputs and relevant
recordings. Use the installed CLI's `runs get`, `simulated-conversations list
--run-id`, `get`, `metrics` and `metric-detail` commands. Check `ok` and paginate
through the public API if the CLI can't prove a complete list.

Metric lists default to the latest output per metric. Request
`include_superseded=true` when recovering historical scores, or retrieve exact
saved output IDs. Select the intended version and scoring occasion; do not count
re-scores as new conversations. Read full output details for composite criteria
missing from list responses.

Retain the exact launch requests and configuration/version evidence if available.
Do not substitute today's resource definition for the historical one. Record
unknown versions explicitly. A shared seed selects cases; it does not make
voice conversations deterministic.

Build a comparability table:

| Dimension | What must match or be explained |
|---|---|
| Cases | Same exact cases and expectations; same ID can have changed content |
| Persona | Same prompt, voice, language, audio condition and interruption settings |
| Agent | Only the intended agent change; preserve model/provider/config context |
| Metrics | Same criterion, polarity, units, scope and output metric version |
| Execution | Iteration counts, concurrency, timeout, time window and failure rates |
| Variants | Base and each mutation separated; mutations include a base run |

If agent and judge both changed, an improved score is confounded. Evaluate both
sets of existing recordings with one frozen judge only when that paid action
is authorized, or report the limitation. If cases differ, compare only the
matching subset and report both unmatched sets. Do not describe unrelated runs
as an A/B test merely because their means differ.

## Compute and inspect

Join on stable case identity/content, persona condition and intended variant
mapping. Treat iterations as repeated observations of a case, not new scenario
coverage. If iteration IDs cannot be paired, aggregate within each case and
compare cases; never zip API list order into fictional pairs.

For each metric and matched case report baseline/candidate valid counts and
values, pass rule if applicable, changed outcomes, and unavailable results.
Keep these counts separate: requested, observed, terminal, validly scored,
failed/skipped/null, and matched. Missing values never become zero or pass.
For composite scores, inspect mandatory criterion failures separately.

Use a useful measure for the data:
- Binary: pass/fail counts, false-pass concerns if the judge is unvalidated,
  and individual pass→fail/fail→pass cases.
- Numerical: unit-aware per-case deltas; distributions when there is enough
  data. Do not claim a meaningful p95 from three calls.
- Categorical: transitions and counts, not an invented category average.

Inspect every changed outcome in a small comparison. Read the transcripts and
metric explanations; listen to audio when the claimed change is acoustic or
turn-taking. A failed connection is an operational regression and also missing
quality evidence, not a successful conversation with a low task score.

For uncertainty with enough independent cases, use case/group-level resampling
or a suitable paired test; don't bootstrap individual turns or repeated calls
as independent. With a small smoke test, show exact cases and counts and say
the effect is inconclusive beyond those examples. No universal number of
repeats establishes statistical confidence.

## Decision and smallest next step

Return **regression observed**, **improvement observed on the matched sample**,
**no observed change**, or **inconclusive**, with evidence and qualifications.
An improvement on a development suite does not prove generalization to traffic.
Don't call a production release safe without the customer's release criteria
and sufficient evidence on the intended population.

Name any blocker, then propose one focused confirmation test with exact case
IDs, frozen metrics, one intended change and a total call/metric budget. Include
base + mutation multipliers. A recommendation is not execution authority.

Attribution

coval-aicoval-ai
View sourceMore from coval-ai →
SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Related Skills

Solution Architect

Designs system architecture, component specifications, and technical integration strategy. Use when: designing solutions, system architecture, technology stack, or integration approaches.

192 votes

Akorchak:Venture Assessment

Generate a comprehensive VC investment assessment report for a company

72 votes

Stock Analysis

Analyze stocks and cryptocurrencies using Yahoo Finance data. Supports portfolio management (create, add, remove assets), crypto analysis (Top 20 by market cap), and periodic performance reports (daily/weekly/monthly/quarterly/yearly). 8 analysis dimensions for stocks, 3 for crypto. Use for stock analysis, portfolio tracking, earnings reactions, or crypto monitoring.

6511 votes

Just Fucking Cancel

Find and cancel unwanted subscriptions by analyzing bank transactions. Detects recurring charges, calculates annual waste, and helps you cancel with direct URLs and browser automation. Use when: 'cancel subscriptions', 'audit subscriptions', 'find recurring charges', 'what am I paying for', 'save money', 'subscription cleanup', 'stop wasting money'. Supports CSV import (Apple Card, Chase, Amex, Citi, Bank of America, Capital One, Mint, Copilot) OR Plaid API for automatic transaction pull. Out...

6511 votes

Telegram Compose

Compose rich, readable Telegram messages using HTML formatting via direct Telegram API. Use when: (1) Sending any Telegram message beyond a simple one-line reply, (2) Creating structured messages with sections, lists, or status updates, (3) Need formatting unavailable via Clawdbot's Markdown conversion (underline, spoilers, expandable blockquotes, user mentions by ID), (4) Sending alerts, reports, summaries, or notifications to Telegram, (5) Want professional, scannable message formatting wit...

6511 votes
View all in business →