LLM evaluation — automated metrics, human feedback, benchmarking. Use when testing performance, measuring AI quality, or establishing evaluation frameworks.
Scanned 5/28/2026
Install to Claude Code
npx -y skills add martineserios/thebrana --skill llm-evaluation --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Llm Evaluation?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/martineserios-llm-evaluation)More formats (shields.io, HTML) on the badges page.
---
name: llm-evaluation
description: "LLM evaluation — automated metrics, human feedback, benchmarking. Use when testing performance, measuring AI quality, or establishing evaluation frameworks."
group: brana
keywords: [llm, evaluation, eval, testing, llm-judge, a-b-testing, benchmarking, bleu, rouge, bertscore, langsmith]
allowed-tools: [Read, Glob, Grep, AskUserQuestion]
status: experimental
source: "https://github.com/wshobson/agents @llm-evaluation"
acquired: "2026-04-30"
quarantine: true
---
<!-- PROCEDURE_FILE: procedures/llm-evaluation.md -->
Read and execute the full procedure from `system/procedures/llm-evaluation.md`.
> QUARANTINE: Community-tier skill. Read-only tools only. Verify patterns against official docs before applying.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!