Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ln 72 Product Outcome Evaluator

ASecurity

Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment.

566 stars
0 votes
0 copies
0 views
Added 9/21/2026
developmentrustgorails

Security Analysis

A100/100

Scanned 9/21/2026

Install to Claude Code

$npx -y skills add levnikolaevich/claude-code-skills --skill ln-72-product-outcome-evaluator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ln 72 Product Outcome Evaluator?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ln 72 Product Outcome Evaluator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/levnikolaevich-ln-72-product-outcome-evaluator/badge)](https://www.skillsdirectory.com/skills/levnikolaevich-ln-72-product-outcome-evaluator)

More formats (shields.io, HTML) on the badges page.

Download with Pro
Files
SKILL.md
---
name: ln-72-product-outcome-evaluator
description: "Evaluates observed product outcomes against a prior hypothesis; does not run experiments or change user treatment."
---

# Product Outcome Evaluator

**Goal:** Determine what available evidence supports about a delivered product outcome and recommend continuation, adjustment or stopping. Remain read-only: do not change instrumentation, experiments, user treatment, campaigns or product files.

**Execution contract:** The checklist defines completion. Track each item internally as `PENDING`, `PROVEN` with evidence, `CLEARED` with evidence its condition is absent, or `UNPROVEN` with a gap; reading, delegation, or tool failure is not proof. Reconcile after each section. Before returning, resolve all `PENDING`, count only `PROVEN` and `CLEARED`, and apply verdict and approval rules to every gap.
Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method.
Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims.
On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority.
Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.


## Tool Routing

| Need | Preferred capability | Fallback |
|---|---|---|
| Original hypothesis | Product intent, baseline, experiment/measurement plan and accepted targets | Reconstruct from attributable sources; keep missing targets unknown |
| Outcome evidence | Authorized analytics, experiment results, customer behavior and cost/support evidence | Sanitized exports with explicit measurement limits |
| Analysis | Reproducible queries/statistics appropriate to the study design | Transparent arithmetic and qualitative inference; no fabricated causal confidence |

## Domain Rules

- Distinguish delivered behavior, observed metric movement and causal product impact. A release or acceptance test proves neither adoption nor business value.
- Do not choose success thresholds after seeing the result. Separate predeclared criteria from exploratory findings and owner preferences.
- Use only authorized data with necessary minimization. A recommendation is not permission to run an experiment or contact users.

## Checklist

### 1. Frame the Outcome Decision

- [ ] Resolve the delivered capability, intended audience, original hypothesis, decision horizon and outcome decision requested.
- [ ] Identify the released/deployed version, rollout/exposure window and relevant baseline or comparison group.
- [ ] Recover predeclared primary metrics, guardrails, targets and stop rules; mark absent criteria rather than inventing them.
- [ ] Separate product intent and owner preference from measured behavior and external assumptions.

### 2. Assess Measurement Fitness

- [ ] Inspect metric definitions, units, denominators, event coverage, deduplication, identity joins and missing data.
- [ ] Check whether users were actually exposed and whether observation duration supports the intended outcome.
- [ ] Assess cohort composition, selection bias, seasonality, concurrent changes and other confounders.
- [ ] For experiments, inspect assignment, contamination, sample imbalance and uncertainty using the actual study design.
- [ ] Distinguish trustworthy measurements, reported results, estimates, qualitative signals and unavailable evidence.

### 3. Evaluate Value and Harm

- [ ] Compare outcomes with valid baselines or controls using reproducible calculations and appropriate uncertainty.
- [ ] Check guardrails and material regressions in user experience, reliability, support burden, cost or data quality.
- [ ] Separate aggregate effects from relevant segments and expose tradeoffs without fishing for favorable subgroups.
- [ ] Distinguish causal conclusions supported by the design from correlations and exploratory interpretations.
- [ ] Identify whether failure lies in adoption, interaction, correctness, measurement or the original value hypothesis.

### 4. Recommend the Next Decision

- [ ] Recommend continue, adjust or stop only to the degree supported by the evidence; explain what could reverse the recommendation.
- [ ] For uncertainty, define the cheapest next measurement or experiment with audience, signal, boundary and decision criterion without executing it.
- [ ] Return results linked to the original requirement/hypothesis and observed deployment state.
- [ ] Report data and causal limitations explicitly; do not transform lack of proof into proof of no effect.

## Verdict

- `SUPPORTED`: evidence supports the intended outcome within the stated population, window and causal limits.
- `NOT_SUPPORTED`: valid evidence contradicts the declared outcome or violates a required guardrail.
- `INCONCLUSIVE`: evidence cannot establish the outcome or causal interpretation.
- `BLOCKED`: essential hypothesis, exposure identity or authorized data is unavailable.

## Self-Check

- [ ] **Reconcile before returning.** Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.

## Output Contract

Report in the user's language, in this order; retain all five fields and state each fact once. Small results may use one line per field; omit empty tables and do not copy linked artifacts:

1. **Result:** Skill-specific verdict and supported outcome.
2. **Scope:** Reviewed/changed scope, exclusions, baseline, and material assumptions.
3. **Evidence:** Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
4. **Verification:** Checks/results, unavailable evidence, and applicable cleanup/external state.
5. **Completion:** `Checklist: X/Y complete`; `Incomplete: None` or each `UNPROVEN` item's reason, outcome impact, and exact next action; residual risks and required decisions.

**Skill-specific evidence:** Hypothesis, deployed exposure, baseline/control, metric definitions and quality, reproducible results and uncertainty, guardrails, causal limits, recommendation and next evidence action.

Attribution

levnikolaevichlevnikolaevich
View sourceMore from levnikolaevich →
SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

284072 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2192 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

9881 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →