Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Data Scientist

ASecurity

Run the data science workflow end to end — framing, exploration, feature engineering, modeling, validation, and communicating results. Use when turning raw data into decisions or predictions.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgo

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill data-scientist --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Scientist?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Data Scientist
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-data-scientist/badge)](https://www.skillsdirectory.com/skills/aicodedecode-data-scientist)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: data-scientist
description: Run the data science workflow end to end — framing, exploration, feature engineering, modeling, validation, and communicating results. Use when turning raw data into decisions or predictions.
category: ai-research
---

# Data Scientist

Data science is decision-making under uncertainty with data. The workflow is: frame the question, 
explore the data, engineer features, model carefully, validate honestly, and communicate so the 
result changes something.

## Overview

Most of data science is not modeling — it's framing the right question and understanding the 
data. A model built on a misunderstood problem or leaky data is worse than no model: it produces 
confident wrong answers. The discipline is in the checks: train/test hygiene, baselines before 
complexity, and communicating uncertainty alongside every number.

## When to use

- A business question that data could answer: churn, demand, pricing, risk, targeting.
- Building a predictive model from tabular, text, or time-series data.
- Exploring an unfamiliar dataset to find what's in it and what it's good for.
- Evaluating whether an existing model is still working.

## Core concepts

- **Problem framing**: translate the business question into a modeling task — what is the target, 
what's the decision, what's the cost of being wrong? The most important step.
- **Exploratory analysis**: distributions, missingness, correlations, outliers, and data quality. 
Plot first; model later.
- **Feature engineering**: creating informative inputs from raw data — often more impactful than 
model choice. Domain knowledge lives here.
- **Baselines**: a simple model (mean prediction, logistic regression, heuristics) before anything 
fancy. Complexity must earn its place.
- **Validation**: honest train/test splits, cross-validation, and — for time data — time-based 
splits. Leakage (future information in training) is the classic silent killer.
- **Uncertainty**: point estimates plus intervals and calibration. "70% ± 15" is more useful than 
"70%."

## Practical workflow

1. Write the decision the analysis will inform and the metric that matters — before touching the 
data.
2. Explore: profile every column, check missingness and outliers, verify joins and grain (one row = 
what?).
3. Build the simplest baseline and a proper validation scheme; record the baseline score.
4. Engineer features iteratively, measuring each addition against the baseline — keep what moves 
the needle.
5. Tune modestly; prefer robust simple models over brittle complex ones unless the gain is real and 
validated.
6. Communicate: the finding, the uncertainty, the limitations, and the recommended action — in 
that order.

```text
Analysis checklist:
[ ] Question and decision defined
[ ] Data grain and quality understood
[ ] Baseline model + score recorded
[ ] Validation scheme resists leakage
[ ] Uncertainty quantified
[ ] Limitations stated plainly
```

## Common pitfalls

- **Leakage**: future or target-derived information in features. Always ask: "would I know this at 
prediction time?"
- **Skipping the baseline**: jumping to complex models. If logistic regression is within noise of 
the fancy model, ship the simple one.
- **Metric myopia**: optimizing accuracy on imbalanced data, or a metric nobody's decision actually 
uses. Match the metric to the cost of errors.
- **Overfitting the validation set**: tuning until the test score looks good. Hold out a final set 
you touch once.
- **Ignoring the data-generating process**: models assume the future resembles the past. Regime 
changes, selection bias, and feedback loops break this.
- **Analysis without a decision**: beautiful notebooks that change nothing. Start from the decision 
and work backward.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →