Evaluate whether benchmark measures its claimed capability
Scanned 6/2/2026
Install to Claude Code
npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill construct-validity-assessment --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Construct Validity Assessment?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/yogsoth-ai-construct-validity-assessment)More formats (shields.io, HTML) on the badges page.
---
name: construct-validity-assessment
description: Evaluate whether benchmark measures its claimed capability
execution: subagent
prompt: ./prompt.md
input: benchmark_name, claimed_capability, task_examples
used-by: benchmark-archaeology
---
# Construct Validity Assessment SOP
Evaluate whether a benchmark actually measures the capability it claims to measure, using psychometric validity frameworks adapted for AI evaluation.
## Input
- **benchmark_name**: Name of the benchmark
- **claimed_capability**: What the benchmark authors claim it measures
- **task_examples**: Representative examples from the benchmark
## Procedure
1. Define the construct (claimed capability) precisely
2. Analyze task requirements — what skills are actually needed to solve examples?
3. Assess content validity — do items representatively sample the construct?
4. Check convergent validity — correlation with other measures of same construct
5. Check discriminant validity — independence from unrelated constructs
6. Identify construct-irrelevant variance (confounds)
## Output
Validity verdict with evidence for each validity dimension.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!