
Claude Skills by thedixitjain
github.com/thedixitjainllama.cpp local GGUF inference + HF Hub model discovery.
Configure RuVLLM local inference with model selection, MicroLoRA fine-tuning, and SONA adaptation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.
Choose methods for measuring long-term product impact after or beyond an A/B test. Use when comparing long-term holdbacks, post-period analysis, continuous monitoring, CLV models, delayed effects, short-term versus long-term metric tradeoffs, or lower-cost alternatives to long-term holdbacks.
Configure LoRA and QLoRA supervised fine-tuning with current best-practice hyperparameters. Use when writing or reviewing a LoRA/QLoRA training configuration, choosing rank/alpha/target modules, or deciding between LoRA, QLoRA, and full fine-tuning.
A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule. Also determines when NOT to A/B test and what to do instead. Use when asked to \"design an A/B test\", \"should we test this\", \"experiment design\", \"how do we know if this works\", \"what's the sample size\", or \"set up an experiment\".
Product analyst — metrics architecture, funnel analysis, A/B test design, retention, and growth measurement.
Use when targeting Machine Learning for Health (ML4H) or deciding whether a computer-science manuscript fits this venue. Encodes conference fit, framing, evidence bar, submission-cycle checks, rebuttal posture, and desk-reject risks for AI for health.
This skill must be used whenever a response needs mathematical notation — equations, filters, set-builder notation, statistics, calculus, linear algebra, logic, ratios, drops, or counts. Load it before composing, including when the user explicitly mentions math-unicode. Emit terminal-native Unicode inline, never raw LaTeX delimiters or commands.
EU MDR 2017/745 compliance specialist for medical device classification, technical documentation, clinical evidence, and post-market surveillance. Covers Annex VIII classification rules, Annex II/III technical files, Annex XIV clinical evaluation, Art. 86 PSUR schedules, and EUDAMED integration. Use when classifying a medical device under MDR, building or gap-checking a technical file, planning clinical evaluation or PMS/PSUR cadence, or preparing for notified body review (e.g., 'what class i...
Use when preparing a MICRO artifact for post-acceptance evaluation — packaging simulators, configs, traces, and scripts so evaluators can regenerate the paper's figures, targeting the ACM Available/Functional/Reproducible badges, handling licensed workloads and long simulations, and earning the optional artifact appendix.
Use when designing or auditing the evaluation of a MICRO paper — choosing the right instrument on the ladder from analytical model to cycle-level simulator to RTL to silicon, tuning baselines the PC will respect, selecting workload suites, running ablations and sensitivity sweeps, and reporting geomeans with full overhead accounting.
'Execute MindTickle primary workflow: Training Content Management. Trigger: \"mindtickle training content management\", \"primary mindtickle workflow\". '
Use when choosing and defending the research design for a MIS Quarterly manuscript — a behavioral survey/experiment, an economics-of-IS identification strategy, a design-science build-and-evaluate cycle, or an organizational/qualitative design. Matches the method to the IS claim and the manuscript category; it designs the study and hands estimation/evaluation to misq-data-analysis.
Use after a MIS Quarterly revision decision to plan the revision and draft the point-by-point response — prioritizing the Senior Editor's letter, addressing tradition-specific rigor concerns (identification, validity/CMB, artifact evaluation, trustworthiness), updating the transparency materials, and keeping the revised manuscript within its page limit. Drafts the response after the manuscript is actually revised; interpret the decision first with misq-review-process.
'Manage mixed precision trainer operations. Auto-activating skill for ML Training. Triggers on: mixed precision trainer, mixed precision trainer Part of the ML Training skill category. Use when working with mixed precision trainer functionality. Trigger with phrases like \"mixed precision trainer\", \"mixed trainer\", \"mixed\". '
End-to-end methodology for AI agents and software engineers to add machine learning algorithms to existing non-ML codebases. Covers problem framing, data readiness, architectural decoupling, and baseline model integration.
Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring.
Plan evaluation strategies for machine-learning product changes. Use when deciding between offline evaluation, interleaving, online A/B tests, multi-armed bandits, or model filtering for ranking, recommendation, search, personalization, or other ML-powered user experiences.
'Configure mlflow tracking setup operations. Auto-activating skill for ML Training. Triggers on: mlflow tracking setup, mlflow tracking setup Part of the ML Training skill category. Use when working with mlflow tracking setup functionality. Trigger with phrases like \"mlflow tracking setup\", \"mlflow setup\", \"mlflow\". '
Use when packaging an accepted MLSys paper's code, configs, and measurement scripts for the venue's post-acceptance artifact evaluation, targeting the Availability, Functional, and Reproducible badges, writing the Artifact Appendix, handling hardware that AE reviewers cannot access, and answering anonymous evaluator questions.
Use when packaging a MobiCom artifact for the evaluation committee — choosing among the three ACM badges as a calibration of what you can prove, building a hardware-optional reproduction path for evaluators without radios, writing the smoke run that establishes functionality, and making code, data, and traces available.
Use when packaging a MobiSys artifact for the Artifact Evaluation Committee — choosing among the three independent ACM badges (Available, Evaluated–Functional, Results Reproduced), building the workflow scripts the single-blind AEC runs on its own machines, and providing a hardware-optional path for evaluators who lack the exact phone, wearable, or board.
Use when planning a MobiSys project timeline backward from the single early-December paper deadline through registration, two-round review with an early-reject cut, the rebuttal, notification, artifact evaluation, camera-ready, and the June conference — with device-experiment lead time built in and the one-deadline-per-year risk made explicit.
Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. Use when deploying or serving AI/ML models, running GPU-accelerated workloads (training, fine-tuning, inference), serving web endpoints, scheduling batch jobs, or scaling Python code to cloud containers with the Modal SDK.
'Manage model checkpoint manager operations. Auto-activating skill for ML Training. Triggers on: model checkpoint manager, model checkpoint manager Part of the ML Training skill category. Use when working with model checkpoint manager functionality. Trigger with phrases like \"model checkpoint manager\", \"model manager\", \"model\". '
'Build model evaluation metrics operations. Auto-activating skill for ML Training. Triggers on: model evaluation metrics, model evaluation metrics Part of the ML Training skill category. Use when working with model evaluation metrics functionality. Trigger with phrases like \"model evaluation metrics\", \"model metrics\", \"model\". '
'Build model explainability tool operations. Auto-activating skill for ML Training. Triggers on: model explainability tool, model explainability tool Part of the ML Training skill category. Use when working with model explainability tool functionality. Trigger with phrases like \"model explainability tool\", \"model tool\", \"model\". '
'Build model quantization tool operations. Auto-activating skill for ML Deployment. Triggers on: model quantization tool, model quantization tool Part of the ML Deployment skill category. Use when working with model quantization tool functionality. Trigger with phrases like \"model quantization tool\", \"model tool\", \"model\". '
Use to build Molecular Cell's mandatory STAR Methods — the Key Resources Table, the Resource Availability subsections, Experimental Model and Subject Details, Method Details, and Quantification and Statistical Analysis, in the required order and with the reagent transparency Molecular Cell reviewers demand.
Generates SQL validation notebooks for dbt PR changes with before/after comparison queries.
Guides borrowers through mortgage refinance evaluation — collects loan data, extracts mortgage statement fields, evaluates qualification, and delivers recommendations with consumer-friendly communication.
Use when packaging datasets, models, prompts, or annotation materials for a NAACL-bound submission — building the artifact around the Responsible NLP checklist's artifact questions, documenting provenance and licensing for language data, and handling community-owned or Indigenous-language resources correctly.
Use when preparing an NDSS artifact for evaluation after conditional acceptance — targeting the Available, Functional, and Reproduced badges, packaging attacks and measurements responsibly, writing the 2-page artifact appendix, and handling dangerous or embargoed material.
Use after NEJM reviews arrive — often including a dedicated statistical reviewer and an editor letter — to triage the decision, answer statistical comments rigorously, and draft a point-by-point response that quotes each comment, the response, and the revised manuscript text. Do not run before the main text is actually revised.
Use to enforce NEJM's clinical statistical reporting — confidence intervals over bare P values, the pre-specified intention-to-treat primary analysis, multiplicity control for secondary endpoints, pre-specified subgroups with interaction tests, missing-data handling, and absolute risk with NNT alongside relative measures.
Use as the final preflight before submitting to NEJM — a complete clinical checklist across significance, registration, reporting guidelines, abstract, statistics, display items, ethics, references, and required files. Bundles a checklist and a clinical cover-letter template.
Curate LLM training data: dedupe, filter, PII redaction.
Create, analyze, and visualize complex networks and graphs in Python with NetworkX. Use when working with network/graph data structures, computing graph algorithms (shortest paths, centrality, clustering), detecting communities, generating synthetic networks (random, scale-free, small-world), reading/writing graph file formats, or drawing network topologies. Common applications include social, biological, transportation, and citation networks.
> Neural pattern training with SONA (Self-Optimizing Neural Architecture), MoE (Mixture of Experts), and EWC++ for knowledge consolidation. Use when: pattern learning, model optimization, knowledge transfer, adaptive routing. Skip when: simple tasks, no learning required, one-off operations.
Use when packaging NeurIPS code, data, models, demos, benchmarks, or other research artifacts for anonymous review, reproducibility, public release, or MLRC-style artifact scrutiny.
Use when preparing an accepted NSDI paper's artifact for badge evaluation — packaging code, traces, and testbed recipes for the AEC, choosing which badges to pursue, meeting Zenodo-style permanence expectations, and timing public release to stay eligible for the Community Award.
Use when planning an NSDI campaign across the spring and fall deadlines — sequencing abstract and paper gates, notification waits, one-shot revision windows, artifact evaluation, and the May symposium — so a networked-systems project always knows which of the two yearly gates it is really building toward.
Strategic DDD — bounded context discovery, context mapping patterns, subdomain classification, ubiquitous language, and organizational alignment
Evaluation criteria and scoring for data engineering artifact reviews
Generates 3-5 divergent design directions through JTBD analysis, competitive research, structured brainstorming, and taste evaluation before convergence. Use when the team has a validated problem but hasn't chosen a solution approach.
JTBD workflow classification and routing - ODI two-phase framework, five job types with workflow sequences, baseline type selection, workflow anti-patterns, and common recipes
Design taste evaluation framework — DVF primary filter, Apple/Google/Jobs design principles as explicit scoring criteria, weighted decision matrix, and option ranking for the DIVERGE wave