Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Wandb Experiment Analyzer

ASecurity

When the user asks to analyze a Weights & Biases (wandb) project or experiment results, particularly when they need to identify the best-performing run based on validation metrics and find the optimal training step within that run. This skill handles querying wandb projects via GraphQL API, extracting run summaries and history data, comparing validation scores across experiments, analyzing step-level performance within the best run, and exporting results to structured formats like CSV files. ...

8 stars
0 votes
0 copies
3 views
Added 9/6/2026
toolsapiperformance

Works with

terminalapi

Security Analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned 9/6/2026

$npx -y skills add zjunlp/Skills --skill wandb-experiment-analyzer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Wandb Experiment Analyzer?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Wandb Experiment Analyzer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/zjunlp-wandb-experiment-analyzer/badge)](https://www.skillsdirectory.com/skills/zjunlp-wandb-experiment-analyzer)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: wandb-experiment-analyzer
description: When the user asks to analyze a Weights & Biases (wandb) project or experiment results, particularly when they need to identify the best-performing run based on validation metrics and find the optimal training step within that run. This skill handles querying wandb projects via GraphQL API, extracting run summaries and history data, comparing validation scores across experiments, analyzing step-level performance within the best run, and exporting results to structured formats like CSV files. It's triggered by requests involving wandb URLs, experiment comparison, validation performance analysis, or finding optimal training checkpoints.
---
# Instructions

## Overview
This skill analyzes wandb projects to identify the best-performing experiment based on validation metrics and finds the optimal training step within that experiment. It outputs results to a CSV file.

## Core Workflow

### 1. Parse User Request
- Extract the wandb project URL from the user's request (format: `https://wandb.ai/<entity>/<project>`)
- Identify the target validation metric (default: `val/test_score/` or similar validation score metrics)
- Determine output requirements (CSV file location, format)

### 2. Query Project Information
- Use the `wandb-query_wandb_tool` with GraphQL to fetch:
  - Project metadata (entity, project name, run count)
  - List of all runs with their summary metrics
  - Look for validation score metrics in summary data

### 3. Identify Best Experiment
- Parse summary metrics from all runs
- Extract validation scores (look for metrics like `val/test_score/`, `val_score`, `validation_score`)
- Compare scores across all runs to identify the best-performing experiment
- Note: Some runs may have null or missing validation scores

### 4. Fetch Detailed History for Best Run
- Query the history data for the best run using GraphQL
- Request sufficient samples to capture all validation score entries
- Handle large response data that may be truncated or saved to files

### 5. Analyze Step-Level Performance
- Parse the history data to extract validation scores and corresponding steps
- Filter out null/missing validation scores
- Sort by validation score in descending order
- Identify the step with the highest validation score

### 6. Generate Output
- Create a CSV file with columns: `best_experiment_name`, `best_step`, `best_val_score`
- Save to the workspace directory as specified by the user
- Provide a summary of findings to the user

## Key Considerations

### Validation Metric Identification
- The validation metric name may vary across projects
- Common patterns: `val/test_score/`, `val_score`, `validation/score`, `eval/score`
- Check summary metrics first to identify the correct metric name

### Data Handling
- Large history responses may be truncated; use the `local-search_overlong_tooloutput` tool to search within saved files
- Some steps may have null validation scores (only recorded at evaluation intervals)
- Ensure proper JSON parsing of nested data structures

### Error Handling
- Handle missing or inaccessible projects/runs gracefully
- Provide clear error messages if validation metrics cannot be found
- Handle authentication/permission issues with wandb API

## Tools Required
- `wandb-query_wandb_tool`: For GraphQL queries to wandb API
- `local-search_overlong_tooloutput`: For searching within large saved tool outputs
- `filesystem-write_file`: For creating CSV output files
- `filesystem-read_file`: For reading saved data files
- `terminal-run_command`: For running analysis scripts (when needed)

## Output Format
The skill produces a CSV file with the following structure:

Attribution

zjunlpzjunlp
View sourceSee grades on GitHubMore from zjunlp →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

ucoz-landing-skill

Create and edit uCoz homepage landing pages via MCP: custom templates, hero sections, lead forms, navigation menus, SEO, and responsive layout. Includes a visual design system (style selection, layout/grid, section recipes, typography/spacing, color tokens, component states, icons, modern CSS/JS, motion, imagery, social proof, copy/voice, accessibility). Uses ucoz-mcp tools for templates, site file uploads, and site modules.

107 votes

Paperclip

Interact with the Paperclip control plane API for task coordination and governance. Use when checking assignments, updating issue status, posting comments, delegating work, managing routines, or calling Paperclip API endpoints.

953191 votes

Pptx

Presentation toolkit (.pptx). Create/edit slides, layouts, content, speaker notes, comments, for programmatic presentation creation and modification.

471861 votes

Daw Music

Digital Audio Workstation usage, music composition, interactive music systems, and game audio implementation for immersive soundscapes.

761 votes

Instantly Rdsthomas Mission Control

Instantly.ai cold email outreach API - manage campaigns, leads, accounts, and analytics. Use for cold email automation, lead management, campaign creation/monitoring, and email account warmup.

761 votes
View all in tools →