
Claude Skills by coval-ai
github.com/coval-aiSet up and connect a new Coval agent for evaluation. Guides through agent type selection, endpoint configuration, system prompt, and default resource attachment. Use when user says "set up an agent", "connect my agent", "create an agent", "add an agent", or "configure agent endpoint".
Build or improve a Coval dashboard with metric visualizations backed by real data. Creates new dashboards from scratch or rebuilds existing ones by analyzing usage patterns, metric frequency, and data density. Use when user says "create a dashboard", "build a dashboard", "improve my dashboard", "add widgets", "visualize my metrics", "make a performance dashboard", or "dashboard for my runs".
Validate one Coval judge against independent human labels, with grouped development/test separation, class-specific errors and uncertainty. Use for trust claims or calibration; without real labels, prepare review rather than inventing ground truth.
Compare before/after Coval runs using matched cases, stable metric versions and inspected conversation evidence. Use to assess a regression or proposed improvement; does not automatically launch or tune anything.
Discover failure modes in Coval runs, uploaded conversations or local conversation exports using scoped sampling and transcript, audio and trace evidence. Prepare human review without inventing ground truth or launching new runs.
Audit an existing Coval evaluation setup for coverage, misleading scores, missing evidence, judge validation and execution risk. Read-only; use when inheriting an eval or deciding whether its results support a release decision.
Choose the next useful step for evaluating a voice or chat agent with Coval. Use for getting started or deciding what to do next; use a specific skill directly when the task is already clear.
Analyze development-set disagreements between human annotations and one Coval text LLM judge, then propose a focused prompt revision. Use coval-calibrate-metric for independent trust measurements.
Select and configure a Coval metric for a concrete product criterion, choosing transcript, audio, trace or deterministic evidence and testing the result. Use for metric creation or scoring design, not for final human calibration.
Migrate configuration from Bluejay voice AI testing platform to Coval. Use when customer says "migrate from bluejay", "bluejay migration", "import bluejay config", or needs to transfer agents, simulations, metrics, and schedules from Bluejay to Coval.
Connect a customer agent to Coval and complete a small first evaluation with inspected voice or chat results. Use when no usable evaluation exists yet; reuse existing recordings and resources when available.
Comprehensive overview of ALL Coval platform resources, their hierarchy, relationships, API endpoints, and ID formats. Use when user asks about Coval resources, data model, how things relate, what endpoints exist, or needs context about the platform structure before making API calls.
Design and create a simulation persona for testing an AI agent. Guides through use case selection, voice and language configuration, behavior prompt crafting, and interruption calibration. Use when user says "create a persona", "design a persona", "set up a test persona", "configure simulation persona", or "build a caller profile".
Derive a SET of simulation personas for an agent from product artifacts — backend payloads, UI screenshots, journey/product docs, and sample real user messages — instead of designing one persona by hand. Identifies who actually interacts with the agent and how they behave, then creates the personas via the CLI. Best for text/chat agents and for new agents with no interaction history. Use when the user says "make personas from these screenshots/payloads", "who are my users", "create a set of p...
Analyze a Coval accent testing report from runs across different speaker accents. Use when a user provides a Coval report URL, report export, run IDs, screenshots, or metric summary and wants evidence-backed next steps such as prompt changes, STT/confirmation adjustments, accent-robust routing, or expanded accent coverage.
Analyze a Coval adversarial / red-team testing report and turn it into an agent-hardening plan. Use when a user provides a Coval report URL, report export, run IDs, screenshots, or a per-scenario scorecard from an adversarial sweep and wants evidence-backed next steps such as prompt/guardrail changes, refusal hardening, verification fixes, escalation routing, or expanded attack coverage.
Analyze a Coval audio-quality testing report from runs across different voice, speaking-style, volume, interruption, and background-noise scenarios. Use when a user provides a Coval report URL, report export, run IDs, screenshots, or metric summary and wants evidence-backed next steps such as prompt changes, tool handling fixes, STT/TTS adjustments, trace setup, or expanded audio-scenario coverage.
Launch a specific Coval evaluation with explicit cases and execution bounds. Use when resources are already selected; for result analysis use quick-eval.
Execute a bounded Coval evaluation on explicit test cases, then audit conversation, recording and metric evidence. Use when the user requests a run; does not automatically rerun, tune an agent or schedule evaluations.
End-to-end Coval accent testing workflow. Creates one persona per accent (each using a distinct accent voice and mirroring your Standard Customer behavior), launches one run per accent against the same voice agent + test set + metrics, polls for completion, builds a per-persona comparison table from the results, and creates the saved multi-run report (grouped by Persona) via the public API. Use when a user wants to follow the Testing Across Accents cookbook (https://docs.coval.dev/guides/test...
End-to-end Coval adversarial / red-team testing workflow. Builds one adversarial test set (12 core attack vectors plus legitimate controls, each with an expected-behavior checklist), creates a persistent "Adversarial User" persona and a Composite Evaluation metric that scores each scenario against its own expected behaviors, launches a multi-iteration run against the agent (voice or chat), polls for completion, builds a per-scenario pass/fail scorecard, and creates a saved report grouped by T...
End-to-end Coval audio-quality testing workflow. Launches one run per audio-robustness scenario against the same voice agent + test set + metrics, polls for completion, and produces the multi-run report URL grouped by Persona. Use when a user wants to follow the Testing Across Audio Qualities cookbook (https://docs.coval.dev/guides/testing-across-audio-qualities) without doing each step by hand.
Monitor a Coval run's progress with live updates. Use when user wants to check run status or wait for completion.
Download audio recordings from Coval voice simulations. Use when user wants to listen to or analyze call recordings.
Retrieve and analyze simulation results from a Coval run. Use when user wants to review evaluation outcomes or debug agent behavior.
Delegate a read-only Coval evaluation, simulation, monitoring, product-how-to, or troubleshooting subtask to Sofia through the official Coval MCP consult_sofia tool. Use when the user needs Coval-specific FDE judgment, analysis of organization data, metric or test-set recommendations, or a diagnosis grounded in Coval runs and conversations.
Build a compact Coval test set from product requirements, observed conversation failures or a coverage gap. Design realistic scenarios with separate expected behaviors, provenance and controlled voice/persona variation.
Turn a large dataset (an existing oversized Coval test set, an export of past conversations, or a CSV/JSON of cases) into a small, high-signal Coval test set by removing duplicates, identifying unique scenarios, and selecting a representative, failure-weighted subset — then bulk-loading it with no row cap. Use when the user says "I have thousands of cases", "dedupe my test set", "my test set is too big", "turn this dataset into a test set", "pick representative scenarios", or "my CSV import o...
Import datasets from HuggingFace and convert them to Coval test sets. Use when the user wants to create test cases from HuggingFace dataset or repository.
Recommend, create, preview, and attach Coval trace-based metrics from OpenTelemetry spans. Use when a user has Coval traces and wants custom trace metrics, trace-aware LLM judge metrics with include_traces, latency, token usage, provider failures, tool behavior, vertical-specific workflow signals, or production monitoring based on trace attributes.
Troubleshoot Coval OpenTelemetry trace ingestion, missing trace UI, sparse traces, bad simulation or conversation correlation, auth/org errors, oversized payloads, duplicate spans, and production debugging with Trace Search.
Improve Coval trace quality after basic ingestion works. Use when traces are sparse, missing useful STT/LLM/TTS/tool spans, missing attributes needed for Coval built-in metrics, or when a customer wants maximum debugging and observability value from agent traces.
Configure an AI agent to send OpenTelemetry traces to Coval. Use when a user wants to add Coval tracing, instrument an agent for simulations or conversation monitoring, make traces show up in Coval, handle SIP/PSTN/WebSocket trace correlation, or replace the one-command wizard with a security-reviewable manual setup.