Skip to content
Skills Directory
Skills
Learn
Security
Categories
Docs
Blog
Pro
Submit a skill
All authors
Claude Skills by runta-dev
github.com/runta-dev
1 skill
A
× 1
0
installs
0
views
Frontierharness Eval
A
Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. Provisions a clean runtime bound to the GitHub repo under evaluation, installs the Harbor and Pier stacks needed for Terminal-Bench and DeepSWE tasks, freezes a golden checkpoint, runs tasks from identical fresh restores while saving trajectories as evidence, generates a comparison diagram, and builds a shareable report. Use when evaluating, benchmarking, scoring, or comparing a coding agent harnes...
ai-agents
go
bash
Votes:
0
GitHub stars:
250
Claude Skills by runta-dev — 1 skill, security-graded | Skills Directory