Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Evaluation Protocol Comparison

ASecurity

Compare implementation differences of same benchmark across papers

417 stars
0 votes
0 copies
0 views
Added 6/1/2026
researchgo

Security Analysis

A100/100

Scanned 6/1/2026

Install to Claude Code

$npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill evaluation-protocol-comparison --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluation Protocol Comparison?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Evaluation Protocol Comparison
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yogsoth-ai-evaluation-protocol-comparison/badge)](https://www.skillsdirectory.com/skills/yogsoth-ai-evaluation-protocol-comparison)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: evaluation-protocol-comparison
description: Compare implementation differences of same benchmark across papers
execution: tactic
used-by: benchmark-archaeology
---

# Evaluation Protocol Comparison Tactic

Compare how different papers implement the same benchmark to expose hidden protocol variance that undermines cross-paper score comparability.

## Stages

### Stage 1: Paper Collection (Same Benchmark)

Collect 10-15 papers that report results on the target benchmark:
- Prioritize diversity: different labs, years, model families
- Include the original benchmark paper as reference protocol
- Include papers from different venues (top conferences, workshops, preprints)
- Search via dare-ss (ss_relevance_search) and dare-scholar (paper_searching)

**Search queries**: "[benchmark name] evaluation", "[benchmark name] results", "[benchmark name] state-of-the-art"

### Stage 2: Protocol Element Extraction

For each paper, run protocol-element-extraction SOP to extract:

| Element Category | Specific Parameters |
|-----------------|-------------------|
| **Data** | Split version, subset selection, preprocessing, filtering |
| **Prompting** | Template format, few-shot examples (count, selection), instruction wording |
| **Generation** | Decoding strategy, temperature, top-p/top-k, max tokens, stop criteria |
| **Evaluation** | Metric implementation, postprocessing, normalization, scoring script version |
| **Infrastructure** | Framework, precision (fp16/bf16/fp32), batch size, hardware |

### Stage 3: Difference Matrix Construction

Build a comparison matrix:
- Rows = protocol elements
- Columns = papers
- Cells = specific value used
- Highlight deviations from original protocol

Compute per-element variance:
- **None**: All papers use identical value
- **Low**: Minor variations (e.g., different random seeds)
- **Medium**: Substantive differences (e.g., different few-shot examples)
- **High**: Fundamental disagreements (e.g., different splits, different metrics)
- **Extreme**: Papers appear to evaluate different things under same name

### Stage 4: Impact Assessment

For each high-variance element:
- Search for ablation studies showing impact of that element
- Estimate score range attributable to protocol choice vs model quality
- Identify which protocol choices systematically favor certain model families
- Flag "protocol p-hacking" — suspicious correlation between protocol choice and reported improvement

## Output

```yaml
protocol_comparison:
  benchmark: string
  papers_compared: int
  reference_protocol: string  # original benchmark paper
  difference_matrix:
    - element: string
      category: data|prompting|generation|evaluation|infrastructure
      variance_level: none|low|medium|high|extreme
      values: list[{paper, value}]
      impact_estimate: string
  highest_variance_elements:
    - element: string
      score_impact: string
      favors: string  # which model family benefits
  protocol_p_hacking_flags:
    - paper: string
      suspicious_choice: string
      benefit: string
  cross_paper_comparability: high|moderate|low|unreliable
  standardization_recommendations:
    - element: string
      recommended_value: string
      rationale: string
```

## Yield Report

| Metric | Minimum |
|--------|---------|
| Papers compared | 8 |
| Protocol elements extracted per paper | 10 |
| High-variance elements identified | 2 |
| Impact estimates produced | 3 |

Attribution

yogsoth-aiyogsoth-ai
View sourceMore from yogsoth-ai →
SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Related Skills

Competitor Analysis

This skill provides comprehensive analysis of competitor SEO and GEO strategies, revealing what's working in your market and identifying opportunities to outperform the competition.

1823 votes

Deep Research

Universal deep research agent team. 13-agent pipeline for rigorous academic research on any topic. 7 modes: full research, quick brief, paper review, lit-review, fact-check, Socratic guided research dialogue, and systematic review with optional meta-analysis. Covers research question formulation, Socratic mentoring, methodology design, systematic literature search, source verification, cross-source synthesis, risk of bias assessment, meta-analysis, APA 7.0 report compilation, editorial review...

452202 votes

Paperclip Distill

Use when an operation issue is a Paperclip cursor-window, distill, or backfill — `operationType: "distill"` or `"backfill"` and the body references a Paperclip source bundle for a project or root issue. Turn raw Paperclip activity into a wiki-insightful project page, decisions log, and history note. This skill exists specifically to replace the stiff, datestamp-heavy templated output that the deterministic distiller produces.

805541 votes

Academic Pipeline

Orchestrator for the full academic research pipeline: research -> write -> integrity check -> review -> revise -> re-review -> re-revise -> final integrity check -> finalize. Coordinates deep-research, academic-paper, and academic-paper-reviewer into a seamless 10-stage workflow with mandatory integrity verification, two-stage peer review, and reproducible quality gates. Triggers on: academic pipeline, research to paper, full paper workflow, paper pipeline, end-to-end paper, research-to-publi...

452201 votes

Exa Search

Semantic search, similar content discovery, and structured research using Exa API

304951 votes
View all in research →