Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Convert To Benchmark Script

ASecurity

Turn the script a user really runs into a benchmark script an optimization loop can run over and over. Same shape as the real job, a few minutes long, repeatable, and leaving nothing behind. It produces the script and a table of what changed, and it does not run it.

3 stars
0 votes
0 copies
0 views
Added 9/29/2026
datago

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add OpenPerfAgent/who-ate-my-flops --skill convert-to-benchmark-script --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Convert To Benchmark Script?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Convert To Benchmark Script
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/openperfagent-convert-to-benchmark-script/badge)](https://www.skillsdirectory.com/skills/openperfagent-convert-to-benchmark-script)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: convert-to-benchmark-script
description: Turn the script a user really runs into a benchmark script an optimization loop can run over and over. Same shape as the real job, a few minutes long, repeatable, and leaving nothing behind. It produces the script and a table of what changed, and it does not run it.
---

# convert-to-benchmark-script

In: the script the user actually runs. 

Out: a new script that does the same job within a few minutes (mostly by reducing total steps), gives the same answer twice, and leaves nothing behind.

## Four rules

1. **It has to be the same job.** Same model size, batch size, sequence length,
   precision, parallelism, and world size etc. Shrink any of them and the bottleneck
   moves somewhere else, and then the loop spends a day optimizing the wrong
   thing.
2. **It has to be short.** Aim for one run within 8 minutes. Every candidate in
   the loop pays this cost, so a 20 minute benchmark is a 20 minute tax on every
   experiment. Short is a budget, not an instruction to cut: if the job already
   fits, leave it whole.
3. **It has to leave nothing behind.** No runs on someone's dashboard, no
   checkpoints eating a quota, no writes into the checkout.
4. **It has to give the same answer twice.** Pin every seed and every source of
   ordering. Use the script's own flag where it has one, and otherwise seed from
   your wrapper before the workload starts. `optimize`'s setup smoke-runs it 3
   times and reads the tolerance off those runs, so drift you leave in here comes
   back there as a tolerance too wide to gate anything on.

## What to turn off

Reach for the script's own flag or an environment variable every time. A deleted
line is a change to their code that nobody agreed to.

**Estimate the cost before you switch anything off.** You cannot measure it here,
so do the arithmetic on paper:

- a checkpoint: `bytes on disk / write bandwidth x saves per run`
- an eval pass: `eval batches x 1/3 of a train step / steps between evals`

Divide either by the run's total wall clock. Under a few percent, turn it off.
Above that, either measure it as a phase of its own rather than buried in
steady-state step time, or turn it off and record the exclusion, with your
estimate, in the contract's **Anything else** row. An exclusion nobody wrote down
becomes a speedup that never shows up in the real job.

| | how | why it matters |
|---|---|---|
| trackers: wandb, tensorboard, mlflow | the script's own flag, or `WANDB_MODE=disabled` | the loop makes hundreds of runs, and every one of them lands on a real dashboard |
| checkpoint writes | the script's own flag | the loop runs hundreds of times, so price the disk as well as the clock before you leave it on |
| eval or validation loops inside training | flag them off, or measure them as their own phase | otherwise every Nth steady-state step has an eval buried in it |
| uploads and notifications: hub pushes, slack, mail on finish | flag | same as trackers, and harder to take back |
| the output directories the script picks for itself: `./outputs`, `./wandb`, `./lightning_logs`, `./checkpoints` | route them all through one flag or environment variable | the workspace does not exist yet during `init`, so `optimize` points that flag at its `runs/` and nothing lands in a candidate commit |

## How much to run

**If the whole job fits the budget, run it whole.** Judge that on the real job: a
5 minute smoke config belonging to a 10 hour run is still the 10 hour run's
benchmark.

If you have to cut, keep both phases: 
- the warmup, paid once. 
- the steady step, paid many times. 

Run enough warmup steps that step time stops changing, then at least 5 to 10
steady steps after it. With fewer, the profiler has nowhere to place its capture
window, and a real speedup cannot be told apart from run-to-run noise.

### Cutting the dataset

You are allowed to use a smaller dataset where it cuts overhead. 
It moves the numbers as well as the clock, starting with dataset initialization time, so record what you cut.

## Two modes

The mode sets the step count. It does not relax anything under "the same job".

| | speed | accuracy |
|---|---|---|
| what it has to show | warmup and steady state, told apart | the quantities the contract's **What must not change** row names |
| steps | the floor above. More buys nothing | what the check needs: bit-exact settles in a few steps, a convergence check cannot be short |
| determinism | seeds fixed, `torch.use_deterministic_algorithms` left off because it costs speed | seeds fixed, data order fixed, deterministic algorithms on when the check is bit-exact |

## Hand it back

Show the user a table first: what you changed, the flag or environment variable
it maps to, and why. Anything you had to leave alone because fixing it would mean
editing their code goes in the same table, as a limitation. Then have them
confirm it.

Keep it short. Reading it costs the user attention, and attention is the scarce
thing in `init`. 

Attribution

OpenPerfAgentOpenPerfAgent
View sourceSee grades on GitHubMore from OpenPerfAgent →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Rank Tracker

This skill helps you track, analyze, and report on keyword ranking positions over time. It monitors both traditional SERP rankings and AI/GEO visibility to provide comprehensive search performance insights.

1821 votes

Youtube Competitor Analyzer

Find and analyze YouTube competitor channels using YouTube Data API v3. Discover competitors through keyword search, category matching, content similarity, and related channel discovery. Compare metrics, content strategies, and market positioning. Use when users want to (1) Find competitors for their YouTube channel, (2) Analyze competitor performance metrics, (3) Compare their channel against competitors, (4) Identify content gaps and opportunities, (5) Benchmark against similar creators, (6...

31 votes

Xlsx

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the...

1798860 votes

Weather Fetcher

Instructions for fetching current weather temperature data for Karachi, Pakistan from wttr.in API

672240 votes

Weather

Get current weather and forecasts (no API key required).

486960 votes
View all in data →