Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Statistical Qa Of Metrics

ASecurity

Decide whether a dashboard metric movement, comparison, or trend is signal or noise — and annotate it honestly (significance, confidence interval, "not enough data yet"). The interop seam with data-platform — invoked by `data-platform/dashboard-builder` when a widget shows a comparison/trend that needs a statistical-validity annotation. data-platform answers "is this number correct?"; this skill answers "is it real?". Used by `applied-statistician` (primary) + `data-platform/dashboard-builder`.

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsrustrails

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill statistical-qa-of-metrics --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Statistical Qa Of Metrics?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Statistical Qa Of Metrics
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-statistical-qa-of-metrics/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-statistical-qa-of-metrics)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: statistical-qa-of-metrics
description: Decide whether a dashboard metric movement, comparison, or trend is signal or noise — and annotate it honestly (significance, confidence interval, "not enough data yet"). The interop seam with data-platform — invoked by `data-platform/dashboard-builder` when a widget shows a comparison/trend that needs a statistical-validity annotation. data-platform answers "is this number correct?"; this skill answers "is it real?". Used by `applied-statistician` (primary) + `data-platform/dashboard-builder`.
---

# Skill: statistical-qa-of-metrics

> **Invoked by:** `applied-statistician` (primary) **and** `data-platform/dashboard-builder` (the seam — when a widget shows a comparison/trend that needs a significance/CI annotation).
>
> **The seam:** `data-platform` owns *"is this number correct?"* — present, in-range, reconciled, fresh (its [`data-quality-tests`](../../../data-platform/skills/data-quality-tests/SKILL.md) skill). **This skill owns *"is this number real?"*** — is the movement signal or sampling noise; does the comparison have enough data; is the trend honest. Non-overlapping by design.
>
> **When to invoke:** "revenue is up 18% — is that real or noise?"; "this KPI tile shows a WoW change — should we annotate it?"; "the dashboard shows variant B winning — can we trust it?"
>
> **Output:** a signal-vs-noise verdict + the annotation to put on the widget (CI, significance flag, or "insufficient data") + the caveat.

## Procedure

1. **Classify the widget's claim:** a *point comparison* (this period vs last), a *trend* (slope over time), or a *group comparison* (segment A vs B).
2. **Get the denominators.** A "+18%" on n=40 is very different from n=40,000. No sample size → the only honest annotation is "insufficient data to call."
3. **Attach uncertainty, not just a point estimate:**
   - **Rate/proportion KPI** → Wilson confidence interval on the rate; flag whether the period-over-period change's CI excludes zero.
   - **Mean/continuous KPI** → CI on the difference (t-based or bootstrap if skewed).
   - **Count/rare events** → Poisson CI; warn that small counts swing wildly.
4. **Distinguish noise from signal explicitly.** "Up 18%, 95% CI [−4%, +40%]" → the honest annotation is *"within normal variation — not yet distinguishable from noise,"* not "up 18%."
5. **For trends,** prefer a fitted slope + CI (or a control chart band) over eyeballing two points. Two-point "trends" are the most common dashboard lie.
6. **Return the annotation** the dashboard widget should display, plus the one-line caveat.

## Annotation patterns for widgets

| Widget shows | Honest annotation |
|---|---|
| "Revenue +18% WoW" (small n) | "+18% (95% CI −4% to +40%) — within normal weekly variation" |
| "Conversion 6.5% vs 5.0%" | "+1.5pp (95% CI +0.6 to +2.4pp), significant at α=0.05" |
| "Signups trending up" (2 points) | replace with a slope + band, or "1 week of data — trend not yet established" |
| A/B winner on the dashboard | route to [`../experiment-analysis/SKILL.md`](../experiment-analysis/SKILL.md) for the full verdict |

## Guardrails
- **Never let a widget assert a movement without an uncertainty band** when the sample is small enough for noise to dominate.
- A KPI provenance note (source query, date range, baseline) is `data-platform`'s job (its house opinion #7); the **statistical** annotation is this skill's job. Don't duplicate; complement.
- Anomaly **detection method** lives here; anomaly **operational alerting** lives in `data-platform`. Keep the seam at "method vs ops."

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →