Skip to content
Back to skills

Monitor Experiment

ASecurity

Monitor running experiments, check progress, collect results. Use when user says ”check results”, ”is it done”, or ”check experiment progress”.

  • 10 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 24, 2026
ai-agentspythonbashapi

Works with

  • api

Security analysis

A100/100

Pro scans all 14 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add FOURTEEN1416/academic-agent-toolkit --skill monitor-experiment --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Monitor Experiment?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Monitor Experiment
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fourteen1416-monitor-experiment/badge)](https://www.skillsdirectory.com/skills/fourteen1416-monitor-experiment)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: monitor-experiment
description: "Monitor running experiments, check progress, collect results. Use when user says ”check results”, ”is it done”, or ”check experiment progress”."
argument-hint: [server-alias or screen-name]
allowed-tools: Bash(ssh *), Bash(echo *), Read, Write, Edit
---

# Monitor Experiment Results

Monitor: $ARGUMENTS

## Workflow

### Step 1: Check What's Running
```bash
ssh <server> "screen -ls"
```

### Step 2: Collect Output from Each Screen
For each screen session, capture the last N lines:
```bash
ssh <server> "screen -S <name> -X hardcopy /tmp/screen_<name>.txt && tail -50 /tmp/screen_<name>.txt"
```

If hardcopy fails, check for log files or tee output.

### Step 3: Check for JSON Result Files
```bash
ssh <server> "ls -lt <results_dir>/*.json 2>/dev/null | head -20"
```

If JSON results exist, fetch and parse them:
```bash
ssh <server> "cat <results_dir>/<latest>.json"
```

### Step 3.5: Pull W&B Metrics (when `wandb: true` in AGENTS.md)

**Skip this step entirely if `wandb` is not set or is `false` in AGENTS.md.**

Pull training curves and metrics from Weights & Biases via Python API:

```bash
# List recent runs in the project
ssh <server> "python3 -c \"
import wandb
api = wandb.Api()
runs = api.runs('<entity>/<project>', per_page=10)
for r in runs:
    print(f'{r.id}  {r.state}  {r.name}  {r.summary.get(\"eval/loss\", \"N/A\")}')
\""

# Pull specific metrics from a run (last 50 steps)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
history = list(run.scan_history(keys=['train/loss', 'eval/loss', 'eval/ppl', 'train/lr'], page_size=50))
print(json.dumps(history[-10:], indent=2))
\""

# Pull run summary (final metrics)
ssh <server> "python3 -c \"
import wandb, json
api = wandb.Api()
run = api.run('<entity>/<project>/<run_id>')
print(json.dumps(dict(run.summary), indent=2, default=str))
\""
```

**What to extract:**
- **Training loss curve** — is it converging? diverging? plateauing?
- **Eval metrics** — loss, PPL, accuracy at latest checkpoint
- **Learning rate** — is the schedule behaving as expected?
- **GPU memory** — any OOM risk?
- **Run status** — running / finished / crashed?

**W&B dashboard link** (include in summary for user):
```
https://wandb.ai/<entity>/<project>/runs/<run_id>
```

> This gives the auto-review-loop richer signal than just screen output — training dynamics, loss curves, and metric trends over time.

### Step 4: Summarize Results

Present results in a comparison table:
```
| Experiment | Metric | Delta vs Baseline | Status |
|-----------|--------|-------------------|--------|
| Baseline  | X.XX   | —                 | done   |
| Method A  | X.XX   | +Y.Y              | done   |
```

### Step 5: Interpret
- Compare against known baselines
- Flag unexpected results (negative delta, NaN, divergence)
- Suggest next steps based on findings

### Step 6: Feishu Notification (if configured)

After results are collected, 检查 `~/.acat/feishu.json`:
- Send `experiment_done` notification: results summary table, delta vs baseline
- If config absent or mode `"off"`: skip entirely (no-op)

## Key Rules
- Always show raw numbers before interpretation
- Compare against the correct baseline (same config)
- Note if experiments are still running (check progress bars, iteration counts)
- If results look wrong, check training logs for errors before concluding

## 补充参考(v2.0 全部吸收批)

- `references/nature-experiment-log/`:实验知识日志(图片/语音/文字标准化记录 → YAML frontmatter Obsidian 日志 + 材料归档)——实验记录沉淀时读(Apache-2.0,来源 nature-skills)。溯源见该目录 UPSTREAM.md。

Files in this skill

  • SKILL.md3.6 KB
  • references/UPSTREAM.md525 B
  • references/nature-experiment-log/README.md2.2 KB
  • references/nature-experiment-log/README_EN.md2.3 KB
  • references/nature-experiment-log/SKILL.md6.3 KB
  • references/nature-experiment-log/UPSTREAM.md556 B
  • references/nature-experiment-log/agents/openai.yaml238 B
  • references/nature-experiment-log/manifest.yaml1.5 KB
  • references/nature-experiment-log/references/example-electrochemical.md1.8 KB
  • references/nature-experiment-log/references/example-log.md2 KB
  • references/nature-experiment-log/references/example-thermal-stability.md1.7 KB
  • references/nature-experiment-log/templates/anomaly-log.md577 B
  • references/nature-experiment-log/templates/equipment-tracking.md870 B
  • references/nature-experiment-log/templates/experiment-index.md653 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…