Compute held-out prompted-task accuracy and prompt wording consistency for T0/P3-style recovery experiments.
Scanned 9/9/2026
Install to Claude Code
npx -y skills add VectorSpaceLab/AREX-Skill --skill zero_shot_prompt_robust_evaluation --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Zero Shot Prompt Robust Evaluation?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/vectorspacelab-zero-shot-prompt-robust-evaluation)More formats (shields.io, HTML) on the badges page.
---
name: zero_shot_prompt_robust_evaluation
description: Compute held-out prompted-task accuracy and prompt wording consistency for T0/P3-style recovery experiments.
---
# Zero-Shot and Prompt-Robust Evaluation
Use this skill after a model or proxy predictor has produced predictions for held-out prompted examples. It measures correctness and robustness across alternative templates for the same raw item.
## Inputs
- Prediction records with `example_id`, `template_id`, `prediction`, and `target`.
## Outputs
- Overall accuracy, per-template accuracy, and same-example prompt consistency.
## Workflow
1. Compare canonical prediction and target strings for exact-match accuracy.
2. Group records by template id for per-prompt diagnostics.
3. Group records by example id and measure whether all prompt variants produce the same prediction.
4. Return JSON-serializable metrics and counts.
## Validation
Run included deterministic metric tests.
## Limitations
This skill does not generate predictions; it only evaluates records already produced by an experiment.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!
Practical guide to testing web applications with screen readers for comprehensive accessibility validation.
在编写新功能、修复错误或重构代码时使用此技能。强制执行测试驱动开发,包含单元测试、集成测试和端到端测试,覆盖率超过80%。
Go测试模式包括表格驱动测试、子测试、基准测试、模糊测试和测试覆盖率。遵循TDD方法论,采用地道的Go实践。
Django测试策略,包括pytest-django、TDD方法论、factory_boy、模拟、覆盖率以及测试Django REST Framework API。
克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则