All authors
tmuskal avatar

Claude Skills by tmuskal

github.com/tmuskal
9 skillsA× 90 installs0 views
Benchmark AdderA

Given a benchmark repository URL (agentic envs like arc-agi, memory/eval benchmarks like longmemeval, QA/code/tool-use benchmarks, etc.), orchestrate the creation of a full Claude Code plugin that benchmarks the current harness setup against it. Wraps the babysitter:babysit skill with the benchmark-plugin-creator process.

ai-agentsbashgit
0
3
Browse TestsA

Explore available ARC-AGI environments - lists games, shows details with ASCII grid visualization, and displays historical scores

ai-agentspythongo
0
3
Compare RunsA

Compare two or more ARC-AGI benchmark runs - shows score deltas, config changes, and trends to track improvement or regression

ai-agentspythonbash
0
3
Cross HarnessA

Cross-harness benchmarking - generate instructions for Codex/Gemini/OpenCode, import results, and compare across harnesses

ai-agentspythonrust
0
3
ReportA

Generate and display comprehensive reports from completed ARC-AGI benchmark runs - shows scores, per-game breakdowns, and performance analysis

ai-agentspythonbash
0
3
Run BenchmarkA

Execute benchmark runs against ARC-AGI games - plays games with Claude Code as the agent and records scores

ai-agentspythongo
0
3
SetupA

Set up the ARC-AGI benchmarking environment - installs dependencies, configures API access, and verifies the setup works

ai-agentspythonbash
0
3
JudgeA

LLM-as-judge shim for LongMemEval - wraps upstream get_anscheck_prompt and calls Anthropic (default) or OpenAI (fallback) with exponential backoff

ai-agentspythonbash
0
3
ResumeA

Detect an incomplete LongMemEval run and continue it from the last checkpoint

ai-agentspythonbash
0
3