Skip to content
Back to skills

Copilot Agent Eval Harness

ASecurity

Build the golden-prompt regression set + pre-publish evaluation gate for a Microsoft 365 Copilot agent — author representative + adversarial prompts, check grounding accuracy and citation correctness, stress the declarative-agent hard limits (50/25/4096/45s, no-loop), and gate publish on the results. Use before declaring any declarative or custom-engine agent done, and on every manifest/grounding change.

  • 7 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 23, 2026
ai-agentsgoapisecurity

Works with

  • cli
  • api

Security analysis

A100/100

Scanned September 23, 2026

npx -y skills add mcorbett51090/RavenClaude --skill copilot-agent-eval-harness --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Copilot Agent Eval Harness?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Copilot Agent Eval Harness
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-copilot-agent-eval-harness/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-copilot-agent-eval-harness)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: copilot-agent-eval-harness
description: "Build the golden-prompt regression set + pre-publish evaluation gate for a Microsoft 365 Copilot agent — author representative + adversarial prompts, check grounding accuracy and citation correctness, stress the declarative-agent hard limits (50/25/4096/45s, no-loop), and gate publish on the results. Use before declaring any declarative or custom-engine agent done, and on every manifest/grounding change."
---

# Copilot agent eval harness

Cross-agent playbook (used by `declarative-agent-engineer`, `api-plugin-engineer`, `graph-connector-engineer`, `agents-sdk-engineer`). House rule: **no agent ships without a golden-prompt regression set** — schema-valid ≠ behaviorally correct.

## 1. Author the golden-prompt set
- **Representative** prompts covering each in-scope task + each grounding source.
- **Boundary** prompts (just inside / just outside scope).
- **Adversarial** prompts: prompt-injection over ingested content (→ flag to `ravenclaude-core/security-reviewer`), out-of-scope coercion, requests that should be refused, ACL-bypass attempts (a low-privilege identity probing for content it shouldn't see).

## 2. Score each run
| Dimension | Check |
|---|---|
| Grounding accuracy | did it use the right source + return correct facts? |
| Citation correctness | are citations present, correct, and clickable (labels working)? |
| Refusal | did it decline what it should? |
| ACL trimming | does a low-privilege test identity get correctly trimmed results? |
| Tone/scope | on-brand, in-scope? |

## 3. Stress the hard limits (declarative agents)
Prompts that push toward 50 grounding items / 25 response items / ~4,096 tokens / 45 s — confirm graceful behavior at the ceiling, and that nothing needs a **loop** (a loop = wrong platform → `agents-sdk-engineer`).

## 4. Gate publish
Run the set on every manifest/grounding/auth change. A regression = no publish. Pair with manifest schema + **RAI** validation. Keep the set in source control next to the agent project.

## Anti-patterns
- "It worked once" instead of a versioned set; no adversarial/ACL prompts; no citation check; gating only on schema validity; not re-running on grounding changes.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…