Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Proxy First Eval

ASecurity

Judge whether a change made things better with demonstrable proxy measures tied to the mechanism it targets, not wall-clock time. Load before choosing acceptance measures for any before/after or A/B claim — CI and fleet speed-ups, build speed, DSP/audio quality, render fidelity. Each proxy names its data source, detection floor and sample size, and every zero is paired with a control on the same instrument.

22 stars
0 votes
0 copies
0 views
Added 9/30/2026
devopsgoc++nodeapi

Works with

api

Security Analysis

A100/100

Scanned 9/30/2026

$npx -y skills add danielraffel/pulp --skill proxy-first-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Proxy First Eval?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Proxy First Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/danielraffel-proxy-first-eval/badge)](https://www.skillsdirectory.com/skills/danielraffel-proxy-first-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: proxy-first-eval
description: Judge whether a change made things better with demonstrable proxy measures tied to the mechanism it targets, not wall-clock time. Load before choosing acceptance measures for any before/after or A/B claim — CI and fleet speed-ups, build speed, DSP/audio quality, render fidelity. Each proxy names its data source, detection floor and sample size, and every zero is paired with a control on the same instrument.
---

# Proxy-first evaluation

Load this **before** you pick how to measure a change, not after the numbers
are in. Any claim of the shape "X is now better / faster / closer" is in scope:
a CI routing fix, a fleet capacity change, a build-speed tweak, a DSP rewrite,
an importer or renderer fix.

## The rule

**Pick the proxy from the mechanism the change targets.** Ask: what does this
change physically alter? Measure that, at the point where it happens, in a
form someone else can re-count from a log line, a label or an annotation.

Wall-clock time is almost never that. It mostly tracks load, queue depth,
neighbouring jobs and noise. A change can make things better while wall time
rises (the fleet got busier) or leave them unchanged while wall time falls (a
quiet evening). **Wall time is context, never the verdict** — report it, label
it load-dependent, and do not let it decide.

## Proxies by mechanism

### CI / fleet

| Change targets | Proxy | Source |
|---|---|---|
| Placement (routing a job to the right runner class) | share of jobs that landed on a runner of the intended class | completed job `runner_name` / `runner_group_name` / labels, not the workflow's `runs-on:` |
| Starvation | share of jobs cancelled before any runner was assigned | jobs with `conclusion=cancelled` and empty `runner_name` |
| Queue wait | wait **per job ahead in the queue**, not raw wait | `created_at` → `started_at`, divided by jobs queued ahead at `created_at` |
| Gate cost | required-gate runs (or minutes) per merged PR | `shipyard metrics gate-cost` |
| Merge-queue churn | merge-queue attempts per merged PR | `merge_group` runs / PRs merged in the window |
| Queue ejections | ejections **by cause** (red check, timeout, wedge, neighbour failure) | merge-queue events + the failing check-run's `output.title` |
| Build speed | compile units rebuilt, cache hit rate, blast radius of a header | `tools/scripts/build_speed_scorecard.py report --split <ISO-time>` |

Normalise every count for volume: a rate per job, per PR or per merge, never a
raw total across windows of different traffic.

### DSP / audio

| Change targets | Proxy | Tool |
|---|---|---|
| Correctness vs reference | null residual **with alignment** (dB) | `assert_null_near`, quality-lab `compare` |
| Aliasing / distortion | tone residual by least-squares projection, THD/THD+N | `tone_residual_db()` prior art, Audio Doctor |
| Perceptual artifacts | detector counts with timestamps (transient smear, dulling, metallic HF, graininess) | `pulp tool run audio-quality-lab -- compare` (`/audio-compare`) |
| Filter shape | magnitude response at named frequencies | `signal::frequency_response`, Audio Doctor |

A listening impression is a pointer to where to measure, not a verdict. See
the `audio-harness` skill for the lanes and the window-floor traps (Hann cannot
see −100 dBc; the default `OversamplerT` kind has ~7 dB alias rejection).

### Render / import fidelity

| Change targets | Proxy | Tool |
|---|---|---|
| Layout | per-node box deltas in px | `layout_parity.py` |
| Material survival | properties present in the envelope | `material_audit.mjs` |
| A named region | per-region score | `diff_against_reference_regions.py` |
| Controls work | driven-control assertions | the `prove-before-showing` skill |

A whole-image similarity score is position-blind triage, not a fidelity
verdict. Read each tool's **Cannot see** line in the CLAUDE.md tool registry
before quoting its number.

## Demonstrable means three numbers per proxy

1. **Data source** — the exact log line, label, annotation or API field, so a
   reviewer can re-count it.
2. **Detection floor** — the smallest effect this instrument can see. Prove it
   with a negative control (run it on a case with the defect removed and show
   the reading collapses), don't derive it.
3. **Sample size** — n per side. Small n (a handful of runs, one merge window,
   one render) is **"insufficient sample"**, not a verdict in either direction.

## Controls and instrument traps

- **Pair every zero with a control** on the same instrument and target that
  must return non-zero. If the control is also zero, the instrument is broken;
  report nothing. Compare the control's *count* to what you expect, not just
  "non-zero".
- **Identical results across different filters means the filter is ignored.**
  Example: `actions/runs?workflow_id=...` silently ignores the parameter and
  returns every workflow's runs; use `actions/workflows/<file>/runs`.
- **Do not grep whole job logs** for a marker. Logs echo the step's own script,
  so the pattern matches its own source. Count annotations, `##[notice]` /
  `##[error]` lines, or check-run `output` instead.
- **`actions/jobs/<id>` handed a check-run id returns a coherent, wrong job.**
  Use `check-runs/<id>` for the merge gate's own record.
- **Watch the failure shape, not only success.** A job queued with no runner
  ever assigned, a merge-queue entry with no `merge_group` run, a test that
  SKIPs — all read as "nothing bad happened" to a success-only query.
- **Confirm the before and after measured the same thing**: same workflow, same
  job name, same stimulus, same canvas size, same build type (Release).

## Checklist

- [ ] Named the mechanism the change targets, in one sentence.
- [ ] Chose a proxy measured at that mechanism, normalised per job / PR / merge.
- [ ] Wrote down the data source a reviewer can re-count.
- [ ] Stated the detection floor, proven by a negative control.
- [ ] Ran a positive control for every zero.
- [ ] Recorded n per side; declared "insufficient sample" if it is small.
- [ ] Checked the failure shape (no runner, no run, SKIP), not only success.
- [ ] Reported wall time as load-dependent context only.

## Report template

```
Mechanism:   <what the change physically alters>
Proxy:       <measure, normalised>            before -> after
Source:      <log line / label / API field / tool invocation>
Floor:       <smallest detectable effect; negative control used>
n:           <before n> / <after n>   (insufficient sample if < ...)
Controls:    <positive control for each zero, with its count>
Verdict:     better | worse | no change | insufficient sample
Context:     wall time <before -> after>, load-dependent, not the verdict
```

## Tools

- `shipyard metrics gate-cost` — gate minutes per merged PR, batch fullness,
  receipt reuse. `shipyard metrics compare` for before/after windows; treat
  its timing columns as context.
- `tools/scripts/build_speed_scorecard.py report --split <ISO-time>` — build
  proxies split at the change.
- `audio-harness` skill (C++ lane, gating) and quality-lab / `/audio-compare`
  (advisory A/B with timestamped detectors).
- Visual-compare tools in the CLAUDE.md tool registry, each with its
  **Cannot see** caveat.
- `trace-analysis` skill when the question is "why is this slow" inside one
  process; a trace gives wall and CPU time per slice, which is a mechanism
  measure, unlike end-to-end wall time.

Attribution

Generous-CorpGenerous-Corp
View sourceSee grades on GitHubMore from danielraffel →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Terraform Module Library

Build reusable Terraform modules for AWS, Azure, and GCP infrastructure following infrastructure-as-code best practices. Use when creating infrastructure modules, standardizing cloud provisioning, or implementing reusable IaC components.

401991 votes

sematext-otel

Wire a service's OpenTelemetry output to Sematext Cloud. Walks through region, App-type, instrumentation flow (managed OTLP endpoint vs Sematext Agent), and signal selection (traces/metrics/logs), then produces the exact env-var block and points at a runnable reference example in this repo. Invoke when instrumenting a new app for Sematext.

01 votes

Deployment Patterns

Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up deployment infrastructure or planning releases.

2699140 votes

Babysit

Watch a pull request or review cycle until it is ready to merge. Use when asked to babysit, monitor, or keep checking PR comments, reviews, and CI until all actionable issues are resolved.

971540 votes

V7 Roster

Interact with the Paperclip control plane API for task coordination and governance. Use when checking assignments, updating issue status, posting comments, delegating work, managing routines, or calling Paperclip API endpoints.

953190 votes
View all in devops →