Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ml Benchmark Evaluation

ASecurity

Rigorous methodology for evaluating ML models on established benchmarks. Covers proper train/val/test splits, baseline verification from original papers, exact metric formula discrepancies, data-leak detection checklist, multi-seed robustness, and honest reporting templates. Use when claiming to beat published baselines, writing methods papers, or auditing existing results.

10 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythonrustgobashawsperformance

Security Analysis

A96/100
mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned 9/29/2026

$npx -y skills add FOURTEEN1416/academic-agent-toolkit --skill ml-benchmark-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ml Benchmark Evaluation?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ml Benchmark Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fourteen1416-ml-benchmark-evaluation/badge)](https://www.skillsdirectory.com/skills/fourteen1416-ml-benchmark-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: ml-benchmark-evaluation
description: Rigorous methodology for evaluating ML models on established benchmarks. Covers proper train/val/test splits, baseline verification from original papers, exact metric formula discrepancies, data-leak detection checklist, multi-seed robustness, and honest reporting templates. Use when claiming to beat published baselines, writing methods papers, or auditing existing results.
category: ml-training
version: 1.0.0
author: Synthetic Sciences
license: MIT
tags: [Evaluation, Benchmark, Validation, Train-Val-Test, Metrics, Data Leak, Reproducibility]
dependencies: ["torch", "numpy"]
---

# ML Benchmark Evaluation — Rigorous Methodology

## When to Use
- Claiming to beat published baselines on any ML benchmark
- Writing a methods paper that compares against prior work
- Any ML evaluation where the reported number matters for publication
- Auditing existing results for overfitting or data leakage

## The Golden Rule

**Your evaluation protocol must be AT LEAST as rigorous as the baseline you claim to beat. Ideally, stricter.**

## 1. Train/Val/Test Split

### The Wrong Way (common but weak):
```python
train = data[:9000]   # 90%
test = data[9000:]    # 10% — used for BOTH model selection AND final metric
# Problem: model selection on test set inflates reported performance
```

### The Right Way:
```python
N_TRAIN, N_VAL, N_TEST = 8000, 1000, 1000
train = data[:N_TRAIN]                          # Gradient updates
val = data[N_TRAIN:N_TRAIN+N_VAL]               # Model selection (checkpoints)
test = data[N_TRAIN+N_VAL:N_TRAIN+N_VAL+N_TEST] # Final metric (ONCE)

# During training:
if epoch % 50 == 0:
    val_metric = evaluate(model, val_loader)  # NOT test_loader
    if val_metric < best_val_metric:
        save_checkpoint(model)

# After training (ONCE):
model = load_best_checkpoint()
final_metric = evaluate(model, test_loader)  # This is the reported number
```

### When the Benchmark Doesn't Use Val Split:
Many benchmarks (PDEBench, some Kaggle, older CV benchmarks) don't use validation splits. When claiming to beat them:
1. Report results using **their protocol** for fair comparison
2. ALSO report results using **proper val split** for scientific rigor
3. Be transparent about both numbers in the paper

## 2. Verify Published Baselines (NEVER Trust Task Descriptions)

Published numbers can be wrong in secondary sources. Always verify from the original paper.

```bash
# Download and parse the original paper
wget -O paper.pdf "https://arxiv.org/pdf/XXXX.XXXXX"
npm i -g @llamaindex/liteparse
liteparse parse paper.pdf -o paper_parsed.md
grep "nRMSE\|accuracy\|F1" paper_parsed.md
```

**Common discrepancies found in practice:**
- Task specification says 5.9e-3 but paper says 9.7e-3 (wrong table, wrong metric)
- Baseline from a different parameter setting or model configuration
- RMSE vs nRMSE vs MSE confusion
- Different train/test split than claimed

## 3. Data Leak Detection Checklist

Run these 6 checks before reporting any result:

```python
# Check 1: IC window preserved exactly (for time-series / PDE problems)
assert np.allclose(preds[:,:,:INIT_STEP], targets[:,:,:INIT_STEP], atol=1e-10)

# Check 2: Predictions differ from targets in predicted window
assert not np.allclose(preds[:,:,INIT_STEP:], targets[:,:,INIT_STEP:], atol=1e-6)

# Check 3: No duplicate samples between train and test
train_hashes = set(hash(x.tobytes()) for x in train_data)
test_hashes = set(hash(x.tobytes()) for x in test_data)
assert len(train_hashes & test_hashes) == 0

# Check 4: Test set never used for gradients
# Verify by code inspection: test_loader only in torch.no_grad() blocks

# Check 5: No NaN/Inf in predictions
assert not np.isnan(preds).any() and not np.isinf(preds).any()

# Check 6: Error distribution is realistic
# If min error is near-zero for many samples, suspicious
assert (per_sample_error < 1e-6).mean() < 0.01  # <1% near-perfect
```

## 4. Metric Computation (Get It Exactly Right)

Different benchmarks use different metrics. Verify the EXACT formula from the benchmark code.

```python
# Example: PDEBench nRMSE
def calc_nrmse(preds, targets, init_step):
    """PDEBench nRMSE: per-timestep spatial RMSE, normalized, averaged."""
    p = preds[:,:,init_step:,:].permute(0,3,1,2)   # [N, C, X, T]
    tg = targets[:,:,init_step:,:].permute(0,3,1,2)
    err = torch.sqrt(torch.mean((p-tg)**2, dim=2))  # spatial RMSE: [N, C, T]
    nrm = torch.sqrt(torch.mean(tg**2, dim=2)) + 1e-20
    return torch.mean(err / nrm).item()

# ALWAYS cross-check against benchmark's own metrics code
# e.g., pdebench/models/metrics.py
```

### CRITICAL: Metric Formula Discrepancies

The SAME metric name (e.g., "nRMSE") can have multiple valid definitions that give different numbers (up to 5-10% difference):

```python
# Formula A: Per-timestep, then average (common in code implementations)
err_per_t = sqrt(mean_spatial((pred-target)^2))  # [N, C, T]
nrm_per_t = sqrt(mean_spatial(target^2))
nrmse_A = mean(err_per_t / nrm_per_t)  # average over N, C, T

# Formula B: Frobenius norm ratio per sample (canonical PDEBench definition)
nrmse_B = mean_over_N(||pred_i - target_i||_F / ||target_i||_F)

# Formula C: Global RMSE / global RMS
nrmse_C = sqrt(mean_all((pred-target)^2)) / sqrt(mean_all(target^2))
```

**These are NOT equivalent.** For a single benchmark comparison, the difference can be 1-10%. Always:
1. Read the benchmark's actual metrics.py code (not just the paper)
2. Compute your metric using the EXACT same formula
3. If in doubt, report BOTH formulas and show you beat under both
4. Document which formula you used in results.json

## 5. Multi-Seed Robustness

Single-seed results can be lucky. For strong claims:
```python
seeds = [42, 123, 7, 2024, 31415]
results = []
for seed in seeds:
    torch.manual_seed(seed)
    model = train(seed)
    results.append(evaluate(model))

mean = np.mean(results)
std = np.std(results)
print(f"nRMSE = {mean:.4e} ± {std:.4e} (n={len(seeds)} seeds)")
```

**Minimum for publication:** 3 seeds for key results, report mean ± std.

## 6. Honest Reporting Template

```markdown
## Results

| Test | Our nRMSE | Published | Improvement | Seeds | Val Split |
|------|-----------|-----------|-------------|-------|-----------|
| A    | X.Xe-Y ± Z.Ze-Y | P.Pe-Q | N.N× | 3 | Yes (8K/1K/1K) |

### Comparison Fairness Notes:
- Our model uses [more modes / higher resolution / ...] than the baseline
- These confounding factors are documented in Table X
- Ablation study (Table Y) isolates the contribution of each change

### Limitations:
- [List every weakness honestly]
```

## 7. Physics-Informed Validation (for PDE/Scientific ML)

Beyond standard ML metrics, verify physical consistency:

| Check | What to Compute | Pass Criterion |
|---|---|---|
| Conservation laws | Mass/momentum/energy integral over time | Drift < 5% of baseline |
| Physical bounds | Density ≥ 0, Temperature ≥ 0, etc. | Zero violations |
| Symmetry | If PDE has symmetry, solution must respect it | Error < 1% |
| Known limits | Analytical solution exists for special case | Match to <1% |
| Error vs time | nRMSE at each rollout step | No exponential blowup |
| Spectral content | FFT of prediction vs truth | No spurious high-freq |

## Common Pitfalls

| Pitfall | Why It's Wrong | Fix |
|---|---|---|
| Model selection on test set | Optimistic bias (1-10%) | Use validation split |
| Single seed | Could be lucky | Report 3+ seeds with std |
| Trusting secondary baseline numbers | Often wrong | Parse original paper |
| Comparing against wrong metric | RMSE ≠ nRMSE ≠ MSE | Read benchmark code |
| Not reporting confounding factors | Unfair comparison | Table of ALL differences |
| Cherry-picking best epoch | Not reproducible | Report final epoch OR val-selected |
| Hiding failure cases | Dishonest | Show worst-case sample explicitly |

Attribution

FOURTEEN1416FOURTEEN1416
View sourceSee grades on GitHubMore from FOURTEEN1416 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →