Skip to content
Back to skills

Judge Evaluation

ASecurity

Use when changing or evaluating judge prompts, scoring thresholds, usage floors, or judgment eval corpora.

  • 192 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 23, 2026
ai-agentsgit

Works with

  • cli

Security analysis

A100/100

Scanned September 23, 2026

npx -y skills add kunchenguid/compact-adviser --skill judge-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Judge Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Judge Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/kunchenguid-judge-evaluation/badge)](https://www.skillsdirectory.com/skills/kunchenguid-judge-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: judge-evaluation
description: Use when changing or evaluating judge prompts, scoring thresholds, usage floors, or judgment eval corpora.
user-invocable: false
metadata:
  internal: true
---

- Judge design rule, measured: two one-sentence atomic questions composed in code beat any single question that folds two judgments together (TypeSafe's own guidance). Hill-climb prompt changes with `eval/tools/earn.py` (paired bootstrap) and read the usage-floor ladder with `eval/tools/schedule.py`; a clause stays only if it earns its place.
- Judgment eval harness: `packages/pi-extension/eval/`. Session transcripts, labels, worksheets, and results stay in gitignored `eval/local/`. See that README for the corpus contract and label rubric.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…