Skip to content
Back to skills

Senior Data Scientist

ASecurity

Use when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model performance", "set up MLOps pipeline", "analyze time series", "calculate sample size", or "deploy a model to production". Expert data science covering statistical modeling, experimentation, causal inference, feature engineering, ML deployment, and advanced analytics with Python, R, and SQL.

  • 847 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added May 29, 2026
ai-agentspythongosqldockerkubernetestestingapibackenddevopsci/cd

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 7 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add borghei/Claude-Skills --skill senior-data-scientist --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Senior Data Scientist?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Senior Data Scientist
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/borghei-senior-data-scientist/badge)](https://www.skillsdirectory.com/skills/borghei-senior-data-scientist)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: senior-data-scientist
description: >
  Use when the user asks to "design an experiment", "build a predictive model",
  "run A/B test analysis", "perform causal inference", "engineer features",
  "evaluate model performance", "set up MLOps pipeline", "analyze time series",
  "calculate sample size", or "deploy a model to production".
  Expert data science covering statistical modeling, experimentation, causal inference,
  feature engineering, ML deployment, and advanced analytics with Python, R, and SQL.
license: MIT + Commons Clause
metadata:
  version: 1.1.0
  author: borghei
  category: engineering
  domain: data-science
  updated: 2026-06-17
  tags: [data-science, ml, statistics, experimentation, python, mlops]
---
# Senior Data Scientist

Expert data science for statistical modeling, experimentation, ML deployment, and data-driven decision making — A/B test design and analysis, feature engineering, model training/evaluation, production deployment, and causal inference.

## Keywords

data-science, machine-learning, statistics, a-b-testing, causal-inference,
feature-engineering, mlops, experiment-design, model-deployment, python,
scikit-learn, pytorch, tensorflow, spark, airflow

## Core Capabilities

- **Experiment design & analysis** — hypothesis framing, power analysis and sample sizing, randomization, SRM monitoring, and post-hoc significance testing.
- **Feature engineering** — profiling, candidate generation (temporal/aggregation/interaction/text), selection (variance, correlation, SHAP/RFE), and leakage validation.
- **Model training & evaluation** — stratified/temporal splits, baselines, hyperparameter tuning, cross-validation, calibration, and fairness checks.
- **Production deployment** — containerized serving, input/output drift monitoring (KS/PSI), canary rollouts, and latency/error SLAs.
- **Causal inference** — propensity score matching, difference-in-differences, regression discontinuity, instrumental variables, and assumption/placebo testing.

## When to Use

- Designing or analyzing an A/B test.
- Building a feature engineering pipeline.
- Training, evaluating, or deploying an ML model.
- Estimating treatment effects from observational data.

## Clarify First

Before running an analysis or pipeline, confirm these inputs. If any is unknown or vague, ASK — do not assume:

- [ ] **Task** — A/B test design / feature engineering / model evaluation / causal inference (selects the script and workflow)
- [ ] **Dataset & target variable** — what you are modeling or measuring (drives feature generation and leakage validation)
- [ ] **Decision metric & minimum effect** — the metric and the smallest effect worth detecting (drives power analysis and sample size)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

## Tools

| Script | Purpose |
|--------|---------|
| `scripts/experiment_designer.py` | A/B test design, power analysis, sample size calculation |
| `scripts/feature_engineering_pipeline.py` | Automated feature generation, correlation analysis, feature selection |
| `scripts/model_evaluation_suite.py` | Model comparison, cross-validation, deployment readiness checks |

> For hypothesis testing, causal inference, and regression analysis, use inline scipy/statsmodels — see [references/statistical_methods_advanced.md](references/statistical_methods_advanced.md).

## References

Load the reference that matches the task — keep this file lean and pull detail on demand:

- **[references/ds-workflows.md](references/ds-workflows.md)** — quick-start commands, tech stack, the five end-to-end workflows (A/B testing, feature pipeline, train/evaluate, deploy, causal inference) with Python snippets, performance targets, and common commands. Read when executing any data-science task.
- **[references/ds-operations.md](references/ds-operations.md)** — troubleshooting table, success criteria, and the full CLI flag reference for each script. Read when diagnosing issues or running the tools.
- **[references/statistical_methods_advanced.md](references/statistical_methods_advanced.md)** — advanced statistical methods reference (hypothesis testing, causal inference, regression). Read for statistical depth.
- **[references/experiment_design_frameworks.md](references/experiment_design_frameworks.md)** — experiment design frameworks and power-analysis foundations. Read when designing rigorous experiments.
- **[references/feature_engineering_patterns.md](references/feature_engineering_patterns.md)** — feature engineering patterns and selection techniques. Read when building features.

## Scope & Limitations

**This skill covers:**
- End-to-end experiment design including power analysis, randomization, and post-hoc analysis
- Feature engineering pipelines with profiling, generation, selection, and validation
- Model training evaluation including cross-validation, calibration, and fairness checks
- Production model deployment with monitoring, drift detection, and canary rollouts

**This skill does NOT cover:**
- Data engineering infrastructure (ETL orchestration, pipeline scheduling, data lake management) -- see `senior-data-engineer`
- Deep learning model architecture design and training at scale (distributed GPU training, custom layers) -- see `senior-ml-engineer`
- Prompt engineering, RAG systems, and LLM fine-tuning workflows -- see `senior-prompt-engineer`
- Computer vision pipelines (object detection, segmentation, video processing) -- see `senior-computer-vision`

## Integration Points

| Skill | Integration | Data Flow |
|-------|-------------|-----------|
| `senior-data-engineer` | Feature pipeline ingests data from ETL outputs; shares data quality validation patterns | Raw data stores --> feature engineering pipeline --> feature store |
| `senior-ml-engineer` | Trained models handed off for MLOps deployment; shares model registry and serving configs | Evaluated model artifacts --> deployment pipeline --> production serving |
| `senior-prompt-engineer` | Embedding features from LLMs feed into ML pipelines; experiment frameworks apply to prompt A/B tests | LLM embeddings --> feature vectors; experiment designs --> prompt evaluation |
| `senior-architect` | Model serving architecture reviewed for scalability; data platform design aligned with training infrastructure | Architecture specs --> deployment topology --> monitoring dashboards |
| `senior-backend` | Model inference endpoints integrated into backend services; API contracts defined for prediction requests | REST/gRPC model API --> backend service layer --> client applications |
| `senior-devops` | CI/CD pipelines extended for model retraining triggers; containerized model images deployed via infrastructure-as-code | Docker images --> Kubernetes manifests --> production clusters |

Files in this skill

  • SKILL.md15.8 KB
  • references/experiment_design_frameworks.md1.4 KB
  • references/feature_engineering_patterns.md1.4 KB
  • references/statistical_methods_advanced.md1.4 KB
  • scripts/experiment_designer.py2.7 KB
  • scripts/feature_engineering_pipeline.py2.8 KB
  • scripts/model_evaluation_suite.py2.7 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…