Turn successful traces into reusable skills, then benchmark those skills across models before you trust them in production.
Scanned 6/8/2026
Install via CLI
openskills install agentskillexchange/skills---
name: "Generate and evaluate agent skills from traces before shipping them into repeatable production workflows with UPskill"
slug: "generate-and-evaluate-agent-skills-from-traces-before-shipping-them-into-repeatable-production-workflows-with-upskill"
description: "Turn successful traces into reusable skills, then benchmark those skills across models before you trust them in production."
github_stars: 477
verification: "security_reviewed"
source: "https://github.com/huggingface/upskill"
author: "Hugging Face"
publisher_type: "open_source"
category: "Code Quality & Review"
framework: "Multi-Framework"
tool_ecosystem:
github_repo: "huggingface/upskill"
github_stars: 477
---
# Generate and evaluate agent skills from traces before shipping them into repeatable production workflows with UPskill
Turn successful traces into reusable skills, then benchmark those skills across models before you trust them in production.
## Prerequisites
uv or Python environment, upskill CLI
## Installation
Use the upstream install or setup path that matches your environment:
- uv pip install upskill
- # or just use uv
- uv sync --extra dev
- uv run scripts/format.py
Requirements and caveats from upstream:
- ## Python API
- python
Basic usage or getting-started notes:
- bash
- uvx upskill
- Create a new skill
- Source: https://github.com/huggingface/upskill
- Extracted from upstream docs: https://raw.githubusercontent.com/huggingface/upskill/HEAD/README.md
## Documentation
- https://github.com/huggingface/upskill
## Source
- [Agent Skill Exchange](https://agentskillexchange.com/skills/generate-and-evaluate-agent-skills-from-traces-before-shipping-them-into-repeatable-production-workflows-with-upskill/)
No comments yet. Be the first to comment!
Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.