Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation"
slug: "benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation"
description: "Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases."
github_stars: 358
verification: "listed"
source: "https://github.com/awslabs/agent-evaluation"
author: "AWS Labs"
publisher_type: "open_source_project"
category: "Runbooks & Diagnostics"
framework: "Custom Agents"
tool_ecosystem:
github_repo: "awslabs/agent-evaluation"
github_stars: 358
---
# Benchmark virtual agents with scripted multi-turn conversations using Agent Evaluation
Run concurrent scripted conversations against a target agent to measure whether it stays on task, responds correctly, and holds up in repeatable test cases.
## Prerequisites
Python environment, target agent endpoint or integration, optional AWS services such as Bedrock or SageMaker
## Installation
No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.
- Source: https://github.com/awslabs/agent-evaluation
## Documentation
- https://awslabs.github.io/agent-evaluation/
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-virtual-agents-with-scripted-multi-turn-conversations-using-agent-evaluation/)
No comments yet. Be the first to comment!