Evaluate held-out direct and CoT instruction examples, compute accuracy delta, and emit FLAN recovery mechanism checks.
Scanned 9/9/2026
Install to Claude Code
npx -y skills add VectorSpaceLab/AREX-Skill --skill flan_heldout_instruction_evaluator --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Flan Heldout Instruction Evaluator?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/vectorspacelab-flan-heldout-instruction-evaluator)More formats (shields.io, HTML) on the badges page.
---
name: flan_heldout_instruction_evaluator
description: Evaluate held-out direct and CoT instruction examples, compute accuracy delta, and emit FLAN recovery mechanism checks.
---
# FLAN Held-Out Instruction Evaluator
Use this skill after an instruction-finetuning run to measure held-out generalization.
## Inputs
- Held-out examples and predictions before/after finetuning.
- Mixture audit, training trace, and target metadata.
## Outputs
- Accuracy before/after/delta and mechanism checks.
## Workflow
1. Check held-out ids are absent from training ids.
2. Compute exact-match accuracy before and after.
3. Emit mechanism checks for CoT, optimizer, and target matching.
## Validation
Run `python tests/test_evaluator.py`.
## Limitations
Supports deterministic proxy metrics, not full benchmark evaluation.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!
Practical guide to testing web applications with screen readers for comprehensive accessibility validation.
在编写新功能、修复错误或重构代码时使用此技能。强制执行测试驱动开发,包含单元测试、集成测试和端到端测试,覆盖率超过80%。
Go测试模式包括表格驱动测试、子测试、基准测试、模糊测试和测试覆盖率。遵循TDD方法论,采用地道的Go实践。
Django测试策略,包括pytest-django、TDD方法论、factory_boy、模拟、覆盖率以及测试Django REST Framework API。
克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则