Running EvalsA
This skill should be used when the user asks to "set up an eval", "compare prompt variants", "test two prompts", "check if this prompt is better", "score model outputs", "add an eval to CI", or otherwise wants to measure whether a prompt or model change is an improvement. Covers authoring eval.yaml, choosing scorers, running offline without API keys, and gating on regressions with evaling.
ai-agentspythongo