Track machine learning experiments with reproducible parameters and metrics
Scanned 5/31/2026
Install via CLI
openskills install diegosouzapw/awesome-omni-skill---
name: ML Experiment Tracking
description: Track machine learning experiments with reproducible parameters and metrics
category: ml
version: 1.0.0
triggers:
- model-training
- experiment-run
- hyperparameter-search
globs: "**/models/**,**/experiments/**,**/ml/**"
---
# ML Experiment Tracking Skill
Track machine learning experiments with reproducible parameters and metrics.
## Trigger Conditions
- Model configuration changes or hyperparameter updates
- New experiment run initiated
- User invokes with "track experiment" or "compare models"
## Input Contract
- **Required:** Experiment parameters (model, hyperparameters, data)
- **Required:** Evaluation metrics
- **Optional:** Baseline comparison, hypothesis
## Output Contract
- Experiment log entry with full reproducibility info
- Comparison table against baseline/prior runs
- Recommendation on whether to promote or iterate
## Tool Permissions
- **Read:** Model configs, training data metadata, metric logs
- **Write:** Experiment logs, comparison reports
- **Execute:** Metric collection commands
## Execution Steps
1. Record experiment hypothesis and parameters
2. Capture environment (dependencies, data version, code commit)
3. Execute or observe training run
4. Collect metrics and artifacts
5. Compare against baseline and prior experiments
6. Recommend: promote, iterate, or abandon
## Success Criteria
- Experiment is fully reproducible from logged parameters
- Metrics compared against baseline
- Clear recommendation with rationale
## Escalation Rules
- Escalate if model performance degrades vs. baseline
- Escalate if data drift detected in training set
- Escalate if experiment requires new infrastructure
## Example Invocations
**Input:** "Compare the BERT-base and DistilBERT models for our classification task"
**Output:** Experiment log: BERT-base (F1: 0.92, latency: 45ms, size: 440MB) vs DistilBERT (F1: 0.89, latency: 12ms, size: 260MB). Recommendation: DistilBERT for production (3% F1 trade-off for 73% latency improvement). Promote to staging for A/B test.
No comments yet. Be the first to comment!
This skill provides comprehensive analysis of competitor SEO and GEO strategies, revealing what's working in your market and identifying opportunities to outperform the competition.
Universal deep research agent team. 13-agent pipeline for rigorous academic research on any topic. 7 modes: full research, quick brief, paper review, lit-review, fact-check, Socratic guided research dialogue, and systematic review with optional meta-analysis. Covers research question formulation, Socratic mentoring, methodology design, systematic literature search, source verification, cross-source synthesis, risk of bias assessment, meta-analysis, APA 7.0 report compilation, editorial review...
Use when an operation issue is a Paperclip cursor-window, distill, or backfill — `operationType: "distill"` or `"backfill"` and the body references a Paperclip source bundle for a project or root issue. Turn raw Paperclip activity into a wiki-insightful project page, decisions log, and history note. This skill exists specifically to replace the stiff, datestamp-heavy templated output that the deterministic distiller produces.
Orchestrator for the full academic research pipeline: research -> write -> integrity check -> review -> revise -> re-review -> re-revise -> final integrity check -> finalize. Coordinates deep-research, academic-paper, and academic-paper-reviewer into a seamless 10-stage workflow with mandatory integrity verification, two-stage peer review, and reproducible quality gates. Triggers on: academic pipeline, research to paper, full paper workflow, paper pipeline, end-to-end paper, research-to-publi...
Semantic search, similar content discovery, and structured research using Exa API