Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 7,777–7,800 of 13,285 skills
Statistical test selection decision tree, per-test assumptions/formulas/interpretation guide, effect size, and power analysis. Use this skill for statistical analysis method selection involving 'statistical test', 't-test', 'ANOVA', 'chi-squared', 'correlation analysis', 'p-value', 'hypothesis testing', 'normality test', 'nonparametric test', 'effect size', etc. Enhances the analyst's statistical analysis capabilities. Note: data cleaning and visualization are outside this skill's scope.
A full analysis pipeline where an agent team collaborates to perform exploratory data analysis (EDA), data cleaning, statistical analysis, visualization, and report writing. Use this skill for 'analyze this data', 'do EDA', 'exploratory analysis', 'statistical analysis', 'data visualization', 'write an analysis report', 'analyze CSV', 'extract data insights', 'data cleaning', 'outlier analysis', and other data analysis tasks. Note: real-time data streaming, ML model training/deployment, and B...
A full ML pipeline where an agent team collaborates to perform data preparation, model design, training, evaluation, and deployment readiness. Use this skill for 'design an ML experiment', 'train a model', 'machine learning project', 'build a deep learning model', 'classification model', 'regression model', 'data preprocessing', 'model evaluation', 'hyperparameter tuning', 'MLOps setup', 'XGBoost model', 'PyTorch model', and other ML experiment tasks. Supports data-preprocessing-only or evalu...
Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology. Use this skill for ML experiment management involving 'experiment tracking', 'MLflow', 'W&B', 'Weights and Biases', 'reproducibility', 'model registry', 'experiment comparison', 'hyperparameter logging', etc. Enhances the training-manager's experiment management capabilities. Note: model architecture design and feature engineering are out...
Airflow DAG pattern, of , retry strategy, etc. , strategy etc. data pipeline orchestration guide. 'Airflow DAG', 'DAG ', 'of', 'retry strategy', 'etc.', '', 'pipeline orchestration', 'Dagster', 'Prefect' etc. pipeline scheduling this for. scheduler-engineerof DAG -ize. , data rule of monitoring dashboard this of scope .
Database normalization/denormalization pattern library. An extension skill for data-modeler that provides 1NF-BCNF criteria, functional dependency analysis, step-by-step normalization procedures, strategic denormalization patterns, and common domain ERD templates. Use when data modeling involves 'normalization', 'denormalization', 'ERD patterns', 'functional dependencies', 'table splitting', 'relationship design', etc. Note: DDL generation and query optimization are outside the scope of this ...
Visualization Chart Selection Guide, Information Hierarchy Design , number-Chart matrix info-architect Extended Skill. 'Chart Selection', ' Visualization', ' ', 'Information Design', 'Chart ', ' Storytelling' etc. Visualto expression . , Chart BI of .
Clinical-grade deep analysis of Apple Health export data. Parses the XML from iPhone's Health app export and produces a comprehensive clinical-quality health assessment report using 20+ peer-reviewed statistical methods including Granger Causality, Transfer Entropy, Convergent Cross Mapping, Sample Entropy, DFA, Cosinor analysis, Kovatchev glucose risk indices, Bayesian change-point detection, and biological age estimation. Use this skill whenever the user mentions Apple Health data, health e...
Use this skill for processing and analyzing large tabular datasets (billions of rows) that exceed available RAM. Vaex excels at out-of-core DataFrame operations, lazy evaluation, fast aggregations, efficient visualization of big data, and machine learning on large datasets. Apply when users need to work with large CSV/HDF5/Arrow/Parquet files, perform fast statistics on massive datasets, create visualizations of big data, or build ML pipelines that don't fit in memory.
Model interpretability and explainability using SHAP (SHapley Additive exPlanations). Use this skill when explaining machine learning model predictions, computing feature importance, generating SHAP plots (waterfall, beeswarm, bar, scatter, force, heatmap), debugging models, analyzing model bias or fairness, comparing models, or implementing explainable AI. Works with tree-based models (XGBoost, LightGBM, Random Forest), deep learning (TensorFlow, PyTorch), linear models, and any black-box mo...
Molecular machine learning toolkit. Property prediction (ADMET, toxicity), GNNs (GCN, MPNN), MoleculeNet benchmarks, pretrained models, featurization, for drug discovery ML.
Parallel/distributed computing. Scale pandas/NumPy beyond memory, parallel DataFrames/Arrays, multi-file processing, task graphs, for larger-than-RAM datasets and parallel workflows.
Access AlphaFold's 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
Output format specifications for healthcare data standards. Use when generating or converting data to FHIR R4, HL7v2 (ADT/ORM/ORU), C-CDA, X12 (837/835/834/270-271), NCPDP D.0, CDISC SDTM/ADaM, CSV, SQL, or dimensional analytics formats.
Run the Phase 0 research workflow to scaffold research artifacts before task planning.
Run the Phase 0 research workflow to scaffold research artifacts before task planning.
Analyze word counts, character statistics, keyword frequency, and estimated reading time. 支持字数统计、词频分析及阅读时长预测。Use when auditing content length, optimizing SEO keywords, or checking document stats.
Provides a development tool that gives detailed information about the execution of any request web profiler bundle, twig, component, dev, php, symfony.
Sort and organize files into folders by type, date, or rules. Use when decluttering dirs, checking structure, running cleanup, generating reports.
Open source annotation tool for machine learning practitioners. text-annotator, python, annotation-tool, data-labeling, dataset, datasets.
Tool for shell commands execution, visualization and alerting. Configured with a simple YAML file. terminal-dashboard, go, alerting, charts, cmd.
Check system health with CPU, memory, disk, and process stats. Use when scanning resources, monitoring load, reporting disk usage, alerting thresholds.
Sorting algorithm and system reference — comparison sorts, distribution sorts, parallel sorting, and industrial sortation. Use when choosing sorting strategies, understanding algorithmic complexity, or designing physical sortation systems.
睡眠改善工具。睡眠分析、改善建议、作息规划、睡眠环境优化、小睡指南、睡眠日记。Sleep tracker with analysis, improvement tips, schedule planning, environment optimization, nap guide.