Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Data Validation

ASecurity

Load when data or a computed result must be checked against rules that should already hold. E.g. validate rows against a schema; check uniqueness, ranges, and referential integrity across files; QA analysis output before delivery.

60 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentsgoperformance

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add ufo-ai/ufo-core --skill data-validation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Validation?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Data Validation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ufo-ai-data-validation/badge)](https://www.skillsdirectory.com/skills/ufo-ai-data-validation)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: data-validation
description: Load when data or a computed result must be checked against rules that should already hold. E.g. validate rows against a schema; check uniqueness, ranges, and referential integrity across files; QA analysis output before delivery.
---
# Data Validation

Pre-delivery QA checks, common pitfalls, and result sanity checking.

## Pre-Delivery Checklist

**Data quality:**

- Source tables are correct for this question; freshness noted ("as of" date)
- No unexpected gaps in time series or missing segments
- Null rates checked; nulls handled explicitly (excluded, imputed, or flagged)
- No double-counting from bad joins or duplicate source records

**Calculations:**

- GROUP BY includes all non-aggregated columns; aggregation grain matches the question
- Rate/percentage denominators are correct and non-zero
- Time period comparisons use equal-length complete periods
- JOIN types are appropriate; row counts verified before and after joins
- Metric definitions match stakeholder expectations

**Reasonableness:**

- Numbers in plausible range; cross-referenced against dashboards or prior reports
- No unexplained jumps/drops in time series
- Edge cases checked (empty segments, zero-activity periods, new entities)

## Common Pitfalls

**Join explosion:** A many-to-many join silently multiplies rows. Always check `COUNT(*)` before and after. Use `COUNT(DISTINCT entity_id)` when counting through joins.

**Incomplete period comparison:** Comparing a partial month to a full month. Filter to complete periods or compare same number of days.

**Denominator shifting:** Rate improves because the definition of "eligible" changed, not because performance improved. Use consistent definitions across compared periods.

**Average of averages:** Averaging pre-aggregated averages ignores group size differences. Aggregate from raw data.

**Timezone mismatch:** Event timestamps in UTC vs. user-facing dates in local time cause off-by-one-day errors. Standardize to a single timezone before analysis.

**Selection bias in segmentation:** "Power users have higher revenue" is circular if you defined power users BY revenue. Define segments on pre-treatment characteristics.

## Sanity Checks

| Metric type | Check                                   |
| ----------- | --------------------------------------- |
| User counts | Match known MAU/DAU figures?            |
| Revenue     | Right order of magnitude vs. known ARR? |
| Rates       | Between 0-100%? Match dashboard?        |
| Growth      | Is 50%+ MoM realistic or a data issue?  |
| Percentages | Segment percentages sum to ~100%?       |

**Red flags:** >50% period-over-period change without cause, exact round numbers, rates at exactly 0% or 100%, results that perfectly confirm the hypothesis, identical values across segments.

**Cross-validation:** Calculate the same metric two ways. Spot-check individual records. Reverse-engineer totals (per-user \* users ~= total).

## Gotchas

- **LEFT JOIN hides missing data** — A LEFT JOIN silently fills unmatched rows with NULLs. If you then filter on a right-table column in WHERE (not ON), it becomes an INNER JOIN. Put right-table filters in the ON clause.
- **Rounding before aggregation compounds error** — Round display values at the end, not intermediate calculations. Rounding percentages per-row then summing can produce totals that don't add to 100%.
- **"No rows returned" is not "zero"** — A query that returns no rows for a segment is different from a segment with zero activity. Missing rows won't show up in aggregations; use COALESCE or a complete date/segment spine.
- **Dashboard numbers include filters you forgot** — When cross-referencing against a dashboard, check its hidden filters (date range, excluded segments, bot filtering). Mismatch is usually a filter difference, not a bug.

Attribution

ufo-aiufo-ai
View sourceMore from ufo-ai →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →