Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Uxr Usability Benchmark

ASecurity

Builds a Usability Benchmark Study with the same tasks each round, task success, time on task, SEQ and SUS defined before data, the baseline and comparison point, a sample note and a re-run date, with numbers only from sessions you ran. Use for "run uxr-usability-benchmark", "UX benchmark", "SUS score", "did the redesign make it better", "before and after usability metrics", "task success rate", "Single Ease Question", "how to score SUS", part of the UX Research with Claude Pack by Polar Bear.

2 stars
0 votes
0 copies
0 views
Added 10/5/2026
ai-agents

Works with

cli

Security Analysis

A100/100

Scanned 10/5/2026

$npx -y skills add polar-bear-org/claude-skills --skill uxr-usability-benchmark --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Uxr Usability Benchmark?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Uxr Usability Benchmark
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/polar-bear-org-uxr-usability-benchmark/badge)](https://www.skillsdirectory.com/skills/polar-bear-org-uxr-usability-benchmark)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: uxr-usability-benchmark
description: Builds a Usability Benchmark Study with the same tasks each round, task success, time on task, SEQ and SUS defined before data, the baseline and comparison point, a sample note and a re-run date, with numbers only from sessions you ran. Use for "run uxr-usability-benchmark", "UX benchmark", "SUS score", "did the redesign make it better", "before and after usability metrics", "task success rate", "Single Ease Question", "how to score SUS", part of the UX Research with Claude Pack by Polar Bear.
---

# Usability Benchmark Study

## When To Use
Leadership asks whether the redesign made things better and you have opinions, not a baseline. Run it before the first measured round, so the tasks and metrics are frozen before any data exists. It answers: on the same tasks, measured the same way, did the numbers move between rounds?

## When Not To Use
To find and rate problems from one round, use Usability Test Findings after a Usability Test Plan. For questionnaire data outside test sessions, use Survey Results Analysis.

## Inputs
- The flows to track, the comparison point (previous version, before and after a redesign, a set date) and when each round runs
- Results you collected, if any rounds have run: per task and per session, anonymised (P1, P2)
If you have none of this, I start from the flows and build the protocol with every result cell reading [not yet collected], marked as a first draft.

## Approach
UX benchmarking as NN/g sets it out in Benchmarking UX: Tracking Metrics (3 May 2020): fixed tasks, metrics chosen in advance, a comparison point, repeated on a cadence. Attitude is measured with the Single Ease Question (Jeff Sauro, MeasuringU) and the System Usability Scale (John Brooke, 1986, via MeasuringU). SUS is not a percentage and does not say why; a person can rate a failed task as easy. The failure it prevents: a "usability score" in the leadership deck that nobody can trace to a session.

## Workflow
1. Ask up to three questions: what the comparison point is, which flows matter to the decision, and how many sessions you can run per round.
2. Freeze the tasks: identical wording, test data and start point every round. Any change breaks the comparison, so log it.
3. Define each metric before data. Task success: binary, with criteria per task. Time on task: successful attempts only, and no think-aloud in benchmark sessions. SEQ after each task: one 7-point question, "Overall, how difficult or easy was the task to complete?". SUS at the end: ten items on a 5-point scale.
4. Write the SUS scoring steps into the protocol: odd items minus one, five minus even items, sum times 2.5, range 0 to 100. No average or norm is shipped as a target.
5. Write the sample note: benchmarks need larger samples than qualitative rounds. You set the number; I state what that sample can and cannot show.
6. Fill the results table only from sessions you collected and pasted: per task, per round, with n. Empty cells read [not yet collected]. I never estimate, round up or extrapolate a score.
7. Set the re-run date and owner. What counts as a change worth acting on is set by a named person before the next round, not after seeing it.

## Output Format
```markdown
# Usability Benchmark Study
**Comparison point:** [baseline vs comparison] | **Cadence:** [when] | **Owner:** [name]
## Frozen tasks
| # | Task wording | Start point and test data | Success criteria |
|---|---|---|---|
| 1 | "[task]" | [screen, data] | [criteria] |
## Metric definitions
| Metric | Definition | When collected |
|---|---|---|
| Task success | [binary, per criteria] | Each task |
| Time on task | [successful attempts only] | Each task |
| SEQ | 7-point, "Overall, how difficult or easy was the task to complete?" | After each task |
| SUS | 10 items, 5-point; (odd minus 1) + (5 minus even), sum x 2.5 | End of session |
## Results (from collected sessions only)
| Task | Round | n | Success | Time on task | SEQ | Source |
|---|---|---|---|---|---|---|
| [1] | [baseline] | [n] | [not yet collected] | [not yet collected] | [not yet collected] | [session files] |
**SUS per round:** [baseline: not yet collected, n = ] [comparison: not yet collected, n = ]
## Sample note
[Sample set by you; what it can show; what it cannot]
## Decision
[Named person] sets the change worth acting on by [date, before the next round] and owns the re-run on [date].
```

## Done When
- Tasks, data and start points are frozen, and every metric is defined before the first session
- Every filled cell traces to a pasted session and carries its n; every other cell reads [not yet collected]
- The re-run date, owner and the person who sets the change threshold are named

## Quality Bar
- No benchmark averages, industry norms or SUS grades are quoted; only your own rounds are compared
- SUS is reported as a score from 0 to 100, never as a percentage or a diagnosis
- Results are per task and round, never per participant in the shared output
- Every number comes from sessions you ran; Claude never reports a metric it did not see collected

## Next
Run uxr-readout-deck (Research Readout Deck) to take the before and after to the people who asked.

## About the makers

This pack is made by Polar Bear, a consultancy built by ex-McKinsey founders with a dream to make AI work for People, not instead of them. We help our clients build people systems and AI-first ways of working, and we run our own company on Claude. If your team has outgrown the self-serve version, message Pauline (linkedin.com/in/paulinebertry).

Attribution

polar-bear-orgpolar-bear-org
View sourceSee grades on GitHubMore from polar-bear-org →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →