Run a blinded, repeatable comparison of workflow or model variants using real artifacts and a frozen rubric.
Scanned 10/6/2026
npx -y skills add williamwue/oh-my-stack --skill eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/williamwue-eval)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: eval
description: "Run a blinded, repeatable comparison of workflow or model variants using real artifacts and a frozen rubric."
disable-model-invocation: true
---
# Eval
## Child session handoff
Read [the handoff contract](../poteto-mode/references/subagent-handoff.md).
New tasks, repair rounds, retries, and queue items use fresh child sessions
with the original brief, every later directive, prior findings and responses,
and unresolved objections. Reuse only for required costly live state, and
only when the host allows it. Stop and fence active writers before replacement.
A host-owned orchestrator's model catalog, workspace binding, child tools,
and review-round rules take precedence over the native binding above.
Keep its task handles and attribution receipts. Do not use a backing child
conversation as a new delegated review, or claim native-record verification
for a host-owned child. Report attribution evidence gaps explicitly.
Define the variant, observable success, baseline, and three to six rubric
criteria before running candidates. Freeze the task input, data snapshot,
runtime versions, and scoring procedure. Separate setup, candidate execution,
judgment, and root synthesis. A better score must survive an actual check of
the artifact, not merely a candidate's self-report.
Give each candidate an isolated equivalent workspace and the same organic
user request. Remove evaluation vocabulary from paths, prompts, and visible
files when blinding is part of the question. Do not hint that competitors or
grading exist. Keep judging criteria away from candidates. If a variant
necessarily changes visible instructions, identify that exposure as part of
the treatment rather than pretending it is blind.
Run candidates within an explicit budget, recording model/runtime identities,
starting state, actual outputs, and failures. [Arena](../arena/SKILL.md) can
coordinate independent attempts when available, but do not let it merge outputs
before scoring. Give a judge anonymous output labels and one shared rubric;
score all variants on the same scale in one pass. Verify tool use or Skill
loading from scoped, available transcripts and artifact shape, never from
self-description. Read every output and resolve disagreement against criteria.
Return the variant and baseline, prompt, rubric, each result and failure,
judge verdict, root checks, uncertainty, and promote/hold recommendation.
Report loss of blinding, unavailable transcripts, or unequal environments.
No promotion, Skill edit, or external publication is implicit in running eval.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!