
Claude Skills by thedixitjain
github.com/thedixitjainBayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
'Transform pyspark transformer operations. Auto-activating skill for Data Pipelines. Triggers on: pyspark transformer, pyspark transformer Part of the Data Pipelines skill category. Use when working with pyspark transformer functionality. Trigger with phrases like \"pyspark transformer\", \"pyspark transformer\", \"pyspark\". '
Use Therapeutics Data Commons through the PyTDC Python package for registry discovery, approved dataset access, task-aware splits, evaluator metrics, benchmark groups, and bounded molecular-oracle workflows.
Fully sharded data-parallel training for large models.
Deep learning framework (PyTorch Lightning / lightning package). Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard, MLflow), distributed training (DDP, FSDP, DeepSpeed), for scalable neural network training.
Clean training loops with built-in distributed support.
'Build pytorch model trainer operations. Auto-activating skill for ML Training. Triggers on: pytorch model trainer, pytorch model trainer Part of the ML Training skill category. Use when working with pytorch model trainer functionality. Trigger with phrases like \"pytorch model trainer\", \"pytorch trainer\", \"pytorch\". '
PyTorch深度学习模式与最佳实践,用于构建稳健、高效且可复现的训练流程、模型架构和数据加载。
日本語翻訳:このファイルは pytorch-patterns 用の日本語翻訳が必要です
PyTorch deep learning patterns and best practices for building robust, efficient, and reproducible training pipelines, model architectures, and data loading.
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.
Automatically invoke this skill whenever the user asks to refresh a semantic model or a dataset. Can also be used to manage, optimize, troubleshoot, or configure a refresh or a refresh schedule.
Use when the causal-identification or measurement strategy is the bottleneck for a The Review of Economics and Statistics (REStat) manuscript — a DID / RD / IV / shift-share design, or a measurement / measurement-error problem. Stress-tests the design to REStat's applied-econometrics-and-measurement bar before exhibits are finalized.
Use when a The Review of Economics and Statistics (REStat) decision letter (R&R or reject-and-resubmit) has arrived and the manuscript needs a response-to-referees letter and a revision plan. Drafts the response strategy; it does not run new analysis (route that to the analysis skills).
Use when anticipating the objections a The Review of Economics and Statistics (REStat) referee will raise for a given design, and pre-empting them in the manuscript before submission. Maps threats to defenses; it does not draft the post-decision response letter (that is restat-rebuttal).
Use when the headline estimate of a The Review of Economics and Statistics (REStat) manuscript needs to survive specification, sample, measurement, and inference choices before submission. Builds the robustness suite that REStat referees expect; it does not establish the primary identification.
Use when running the final pre-submission preflight for The Review of Economics and Statistics (REStat) via Editorial Express — format rules, abstract limit, online-appendix cap, article categories, submission fee, the proprietary-data fee hold, and house style. Final checks; it does not draft content.
Use when building or revising the exhibits of a The Review of Economics and Statistics (REStat) manuscript so the main result is legible in one table or figure in AEA/REStat house style. Designs publication-grade exhibits; it does not run the underlying estimation.
Use when deciding how much theory or structure a The Review of Economics and Statistics (REStat) manuscript should carry — right-sizing a model so it interprets or disciplines the empirical estimate without becoming the contribution. Calibrates the theory's role; it does not develop new theory for its own sake.
Use when deciding whether an applied-economics project fits The Review of Economics and Statistics (REStat) rather than AER / AEJ:Applied / J. Econometrics / a field journal, and sharpening the question to REStat's empirical-and-measurement bar. Decides venue fit and frames the question; it does not run the analysis.
Use when deciding which restat-* sub-skill to invoke next, or when sequencing manuscript work from topic selection through rebuttal for a The Review of Economics and Statistics (REStat) submission. Routes — it does not replace — the specialized skills.
Use when targeting Review of Economics and Statistics (REStat) or deciding whether an applied econometrics manuscript fits this venue. Encodes the journal's fit, framing, method-and-evidence bar, house style, official-submission re-check, and desk-reject heuristics.
Classifies agent tasks into 4 risk tiers (GREEN/YELLOW/RED/CRITICAL). Use when assessing action reversibility before committing to an approach.
Medical device risk management specialist implementing ISO 14971 throughout product lifecycle. Provides risk analysis, risk evaluation, risk control, and post-production information analysis. Use when user mentions risk management, ISO 14971, risk analysis, FMEA, fault tree analysis, hazard identification, risk control, risk matrix, benefit-risk analysis, residual risk, risk acceptability, or post-market risk.
'Manage roc curve plotter operations. Auto-activating skill for ML Training. Triggers on: roc curve plotter, roc curve plotter Part of the ML Training skill category. Use when working with roc curve plotter functionality. Trigger with phrases like \"roc curve plotter\", \"roc plotter\", \"roc\". '
Use when packaging the artifacts behind an RSS (Robotics: Science and Systems) paper — code, trained policies, trial ledgers, hardware documentation, and footage — first as anonymous review-time evidence and then as the public release the venue's free open-access proceedings culture expects, without any badge program to structure it.
'Analyze datasets by running clustering algorithms (K-means, DBSCAN, hierarchical) to identify data groups. Use when requesting \"run clustering\", \"cluster analysis\", or \"group data points\". Trigger with relevant phrases based on skill purpose. '
Train RuView models — camera-free WiFlow pose (10 sensor signals, no labels), camera-supervised pose (MediaPipe + ESP32 CSI → 92.9% PCK@20, ADR-079), RuVector contrastive embeddings (AETHER, ADR-024), domain generalization (MERIDIAN, ADR-027), local SNN environment adaptation, plus GPU training on GCloud and Hugging Face publishing. Use when building, fine-tuning, evaluating, or shipping a model.
Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for exploratory scRNA-seq analysis with established workflows. For deep learning models use scvi-tools; for data format questions use anndata.
Scanpy is a scalable Python toolkit for analyzing single-cell RNA-seq data, built on AnnData. Apply this skill for complete single-cell workflows including quality control, normalization, dimensionality reduction, clustering, marker gene identification, visualization, and trajectory analysis.
Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.
論文、提案書、文献レビュー、方法論セクション、証拠の質、引用サポート、研究論文フィードバックのための構造化された学術的作業評価。
Use to enforce Science's statistics and reproducibility reporting — n and replication, test choice and assumptions, effect sizes with uncertainty, multiple-comparison control, randomization/blinding, and pre-registration where relevant.
Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
Machine learning in Python with scikit-learn. Use for classification, regression, clustering, model evaluation, and ML pipelines.
RNA velocity analysis with scVelo. Estimate cell state transitions from unspliced/spliced mRNA dynamics, infer trajectory directions, compute latent time, and identify driver genes in single-cell RNA-seq data. Complements Scanpy/scVI-tools for trajectory inference.
Statistical visualization with pandas integration. Use for quick exploration of distributions, relationships, and categorical comparisons with attractive defaults. Best for box plots, violin plots, pair plots, heatmaps. Built on matplotlib. For interactive plots use plotly; for publication styling use scientific-visualization.
Seaborn is a Python visualization library for creating publication-quality statistical graphics. Use this skill for dataset-oriented plotting, multivariate analysis, automatic statistical estimation, and complex multi-panel figures with minimal code.
Computer vision engineering skill for object detection, image segmentation, and visual AI systems. Covers CNN and Vision Transformer architectures, YOLO/Faster R-CNN/DETR detection, Mask R-CNN/SAM segmentation, and production deployment with ONNX/TensorRT. Includes PyTorch, torchvision, Ultralytics, Detectron2, and MMDetection frameworks. Use when building detection pipelines, training custom models, optimizing inference, or deploying vision systems.
World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking — using Python (NumPy, Pandas, Scikit-learn), R, and SQL. Use when designing or analysing controlled e...
../../../engineering-team/skills/senior-ml-engineer/SKILL.md
Use when packaging a SenSys artifact for the Artifact Evaluation Committee — choosing which of the three ACM badges (Available, Functional, Reproduced) to pursue, building a hardware-optional evaluation path for reviewers without your testbed, documenting energy and hardware provenance, and passing the smoke run that proves functionality.
Use when designing or auditing a SenSys evaluation — energy and low-power measurement with a named instrument, real-testbed and deployment realism, honest sensor ground truth, on-device latency and memory, and same-hardware baselines, so the evidence meets SenSys's built-and-measured bar rather than a simulation or offline-benchmark one.
vLLM: high-throughput LLM serving, OpenAI API, quantization.
'Implement machine learning experiment tracking using MLflow or Weights & Biases. Configures environment and provides code for logging parameters, metrics, and artifacts. Use when asked to \"setup experiment tracking\" or \"initialize MLflow\". Trigger with relevant phrases based on skill purpose. '
Configure the project's game engine and version. Pins the engine in CLAUDE.md, detects knowledge gaps, and populates engine reference docs via WebSearch when the version is beyond the LLM's training data.
Provisions the oracle ML inference daemon with onnxruntime via uv. Use when setting up local ONNX model inference for skill quality evaluation.
Use when packaging an ACM SIGCOMM paper's code, traces, topologies, and configuration for the artifact-evaluation committee — choosing ACM badges (Artifacts Available, Evaluated, Results Reproduced) as claim calibration, building downscaled topologies and trace substitutes, and making a networking testbed result rebuildable by a reviewer.
Use when pursuing the Graphics Replicability Stamp (GRSI) or Code Replicability in Computer Graphics (CRCG) recognition for an accepted SIGGRAPH / TOG paper, covering how graphics replicability differs from ACM artifact badging, what volunteers actually run, deterministic result reproduction, Software Heritage archiving, and the separate post-acceptance timing.
Use when packaging code, run files, test collections, or judgments for a SIGIR submission — deciding between an artifact inside a full/short paper and a standalone Resources track paper, building reviewer-runnable IR repositories, run-file and qrels hygiene, licensing and datasheets, and single- vs double-anonymous handling.