Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Review Llm Annotations And Improve Prompt

ASecurity

Analyze development-set disagreements between human annotations and one Coval text LLM judge, then propose a focused prompt revision. Use coval-calibrate-metric for independent trust measurements.

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
businessrustgoapi

Works with

cliapi

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add coval-ai/coval-external-skills --skill review-llm-annotations-and-improve-prompt --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Review Llm Annotations And Improve Prompt?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Review Llm Annotations And Improve Prompt
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/coval-ai-review-llm-annotations-and-improve-prompt/badge)](https://www.skillsdirectory.com/skills/coval-ai-review-llm-annotations-and-improve-prompt)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: review-llm-annotations-and-improve-prompt
description: Analyze development-set disagreements between human annotations and one Coval text LLM judge, then propose a focused prompt revision. Use coval-calibrate-metric for independent trust measurements.
argument-hint: "<project-id> <metric-id>"
---

# Improve a metric from reviewed disagreements

Use this for development, not to declare a judge calibrated. When the goal is a
trust or release claim, use `coval-calibrate-metric` if installed. This workflow
remains usable by itself with the boundaries below.

Confirm organization/workspace and one text judge metric. Read its definition,
version, project and completed human annotations using current CLI `context`
and `--help`. Binary, categorical and numerical text judges require different
error analysis; don't silently coerce one type into another.

Collect actual completed human labels and reviewer notes, exact machine output
IDs/versions and the relevant transcripts. Paginate the public reviews API when
the CLI can't prove completeness. Zero is a valid label; null or pending is
missing. AI-suggested labels are not human ground truth. Multiple reviewers on
one call need adjudication, not duplicate counting. Annotation `simulation_output_id` can identify an uploaded
conversation; retain the source collection and retrieve original metrics through
its matching public API/CLI resource.

Separate related calls into train (prompt examples), development and untouched
test groups before tuning. Inspect only the development disagreements and an
agreement sample. For binary judges include failure detection and pass recall,
not just raw agreement. For numerical judges fix tolerance before inspecting
scores; don't widen it to improve agreement. For categorical judges show the
actual confusion counts.

Read each disagreement against the rubric. The human or the judge may be wrong;
a domain expert must adjudicate a disputed human label. Never offer “make all
human labels match the machine” as a shortcut. Preserve original labels and notes.

Draft the smallest prompt change supported by observed evidence. Preserve the
criterion, output type and successful boundaries. Include only train examples;
never leak held-out examples into the prompt. Show the old/new prompt and affected
failure pattern before seeking any missing update authority. Prefer a separate
candidate metric so production defaults and historical comparison stay intact.

Use a bounded `coval --agent metrics test <candidate-id>
--simulation-output-ids <ids>` request to score existing outputs when authorized.
Inspect per-item responses and poll each returned `metric_output_ulid` using
`simulated-conversations metric-detail <simulation-id> <output-ulid>` or
`GET /v1/conversations/simulated/{simulation_id}/metrics/{metric_output_id}`.
The API exposes `explanation`; don't print credentials or use an invented nested
outputs route. Count re-scores against the agreed metric budget and stop on
ambiguity rather than enqueueing duplicates.

Report development changes and remaining errors. Freeze the candidate before
measuring it once on untouched human-labeled groups. Without that evidence, say
“prompt candidate improved on development examples,” not “validated judge.” No
automatic resimulation, human-label edits, default attachment or notifications.

Attribution

coval-aicoval-ai
View sourceMore from coval-ai →
SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Related Skills

Solution Architect

Designs system architecture, component specifications, and technical integration strategy. Use when: designing solutions, system architecture, technology stack, or integration approaches.

192 votes

Akorchak:Venture Assessment

Generate a comprehensive VC investment assessment report for a company

72 votes

Stock Analysis

Analyze stocks and cryptocurrencies using Yahoo Finance data. Supports portfolio management (create, add, remove assets), crypto analysis (Top 20 by market cap), and periodic performance reports (daily/weekly/monthly/quarterly/yearly). 8 analysis dimensions for stocks, 3 for crypto. Use for stock analysis, portfolio tracking, earnings reactions, or crypto monitoring.

6511 votes

Just Fucking Cancel

Find and cancel unwanted subscriptions by analyzing bank transactions. Detects recurring charges, calculates annual waste, and helps you cancel with direct URLs and browser automation. Use when: 'cancel subscriptions', 'audit subscriptions', 'find recurring charges', 'what am I paying for', 'save money', 'subscription cleanup', 'stop wasting money'. Supports CSV import (Apple Card, Chase, Amex, Citi, Bank of America, Capital One, Mint, Copilot) OR Plaid API for automatic transaction pull. Out...

6511 votes

Telegram Compose

Compose rich, readable Telegram messages using HTML formatting via direct Telegram API. Use when: (1) Sending any Telegram message beyond a simple one-line reply, (2) Creating structured messages with sections, lists, or status updates, (3) Need formatting unavailable via Clawdbot's Markdown conversion (underline, spoilers, expandable blockquotes, user mentions by ID), (4) Sending alerts, reports, summaries, or notifications to Telegram, (5) Want professional, scannable message formatting wit...

6511 votes
View all in business →