Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Agentic Benchmark Top5

ASecurity

Mount your personal Top-5 agentic benchmarks aligned to your work instead of a global index. Use when choosing models, building a model stack, or comparing performance/cost/speed. Use quando montar top-5 de benchmarks, escolher modelos, comparar custo x performance, fugir de índice global genérico.

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
ai-agentsgorailsperformance

Works with

terminalcli

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add majinmagros/magros.ai-skills --skill agentic-benchmark-top5 --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agentic Benchmark Top5?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Agentic Benchmark Top5
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/majinmagros-agentic-benchmark-top5/badge)](https://www.skillsdirectory.com/skills/majinmagros-agentic-benchmark-top5)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: agentic-benchmark-top5
description: Mount your personal Top-5 agentic benchmarks aligned to your work instead of a global index. Use when choosing models, building a model stack, or comparing performance/cost/speed. Use quando montar top-5 de benchmarks, escolher modelos, comparar custo x performance, fugir de índice global genérico.
---

# Agentic Benchmark Top-5

Monte seu Top-5 pessoal de benchmarks agênticos e escolha uma stack de modelos, não um modelo único.

## Quando usar

Quando o índice global fica borrado, quando dois modelos parecem empatados no score mas diferem em custo/velocidade, ou quando você precisa justificar SOTA vs. workhorse vs. leve para trabalho agêntico real.

## Passos

1. **Declare o trabalho-alvo** — escreva em 1 frase o que seus agentes farão sem supervisão (ex. "codar features longas + operar rotinas de backoffice"). O alvo define quais benchmarks valem.
2. **Monte a ficha dos 5** — use estes como base, troque no máximo 1-2 pelo seu domínio:
   - **Terminal-Bench:** coding agêntico puro. Agente em container preparado, loop comando→resultado, verificador valida estado final. Sinal: capacidade bruta de engenharia.
   - **Apex Agents:** proxy de knowledge work. Banking, consulting, legal, com tasks criadas por experts. Prompt vago + workspace de documentos. Sinal: transferência para domínios duros fora de SWE.
   - **Automation Bench:** tasks em domínios reais (finance, HR, marketing, operations, sales, support) com apps reais. Score = objetivo cumprido SEM violar guardrails. Sinal: alinhamento/instruction-following.
   - **Omniscience:** honestidade. Respostas graded como correct / incorrect / partial / not-attempted, sem penalidade para "não sei". Sinal: taxa de alucinação e custo da honestidade.
   - **Deep SWE:** SWE de horizonte longo a partir de issues/PRs, com prompts curtos realistas. Sinal: autonomia com spec mínima.
3. **Aplique o triângulo performance/custo/velocidade** — para cada benchmark levante: score, tokens por task, tempo por task, custo por task. Métrica final: output útil de agente por hora por dólar. Performance sozinha não decide.
4. **Aplique a regra variância-vs-saturação** — descarte benchmark "flat line" (todos empatados, saturação ~85-90%+). Sem variância não há alfa. Procure a curva com queda: ali está a informação de qual tier usar.
5. **Monte o índice pessoal** — fixe 1 modelo controle, separe em 3 tiers: SOTA / workhorse / leve. SOTA para horizonte longo e crítico, workhorse 10-20x mais barato para volume, leve para triagem/delegação com handoff bem desenhado.
6. **Reporte a decisão** — tabela: benchmark | o que proxya | quem vence | custo/velocidade caveat | decisão de tier. Declare o que ficou de fora e por quê.

## Regras

- NEVER decidir por 1 benchmark ou pelo índice global sozinho.
- NEVER comparar só score: sempre anexe tokens, tempo e custo por task.
- Descarte benchmark saturado, salvo se seu trabalho vive exatamente do 1-3% residual.
- Prefira proxies alinhados ao seu trabalho, não os mais famosos.
- Modelo que você não pode pagar é irrelevante: preço entra na definição de "melhor".

## Related skills

- `agent-eval` — avaliação first-party de agentes via CLI.
- `llm-leaderboard-tracker` — acompanhamento contínuo de leaderboards.
- `benchmark-methodology` — desenho e scoring de benchmarks próprios.

Attribution

majinmagrosmajinmagros
View sourceMore from majinmagros →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.

1023331 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

686011 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3331 votes

catchup

Recovers prior coding-agent session context by running `catchup <agent> --since-compact`, which extracts a clean summary of a previous Codex, Claude Code, Antigravity, OpenCode, or Pi Agent session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", or asks to recover/summarize a previous session before continuing. Do NOT use for the current conversation, git history, or any non-agent log.

611 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →