Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Harbor Harness Improvement Loop

ASecurity

Analyze Harbor benchmark jobs and ATIF trajectories, identify evidence-backed weaknesses across LibrAgent builtin tools, prompts, agent guidance, execution policy, and benchmark instrumentation, then design controlled improvements and rerun comparisons. Use when analyzing Harbor or Terminal-Bench results, optimizing the harness, investigating tool-call failures or retries, auditing prompt/context efficiency, or running repeated BM → analysis → improvement cycles.

10 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentspythongobashgitapibackend

Works with

terminalapi

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add fritzprix/libr-agent --skill harbor-harness-improvement-loop --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Harbor Harness Improvement Loop?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Harbor Harness Improvement Loop
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fritzprix-harbor-harness-improvement-loop/badge)](https://www.skillsdirectory.com/skills/fritzprix-harbor-harness-improvement-loop)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: harbor-harness-improvement-loop
description: Analyze Harbor benchmark jobs and ATIF trajectories, identify evidence-backed weaknesses across LibrAgent builtin tools, prompts, agent guidance, execution policy, and benchmark instrumentation, then design controlled improvements and rerun comparisons. Use when analyzing Harbor or Terminal-Bench results, optimizing the harness, investigating tool-call failures or retries, auditing prompt/context efficiency, or running repeated BM → analysis → improvement cycles.
---

# Harbor Harness Improvement Loop

Build a repeatable **benchmark → diagnosis → smallest intervention → rerun**
cycle. Optimize general agent capability, not benchmark-specific shortcuts.

## Analysis scope (non-negotiable)

A job in this repo is a **harness** test. The only question is: which tool,
prompt, handler, adapter, or telemetry contract failed?

**Don't care (never diagnose, compare, or narrate):** model, serving engine.

Do not put those in findings, owning layers, hypotheses, outcome commentary,
contract tables, or next-cycle.

Allowed output only:

- tool sequences (name, args, order, repeats)
- observations (size, hint, error contract)
- schema / handler / response / prompt / adapter / telemetry source

Reconstruct `intent → tool call → observation → next tool`. First divergence
is a tool-pattern class. No harness contract broken → write the pattern and
**stop**.

## Hard rules

- Treat reward as an outcome, not a cause. Inspect trajectories and source before
  recommending changes.
- Separate **measured fact**, **trace interpretation**, and **hypothesis**.
- Do not claim causality from one trial unless a deterministic contract violation
  is visible. Label single-trial findings as hypotheses.
- Compare runs only when dataset/task selection, assistant, attempts,
  concurrency, execution mode, workspace mode, timeout/resource policy, and
  verifier configuration are equivalent.
- Never modify benchmark tasks, verifiers, official timeout/resource limits, or
  add task-specific prompt clues to improve scores.
- Do not recommend a prompt change when a schema, handler, response, or
  instrumentation fix is the narrower source-of-truth solution.
- Do not duplicate the same instruction across system prompt, assistant prompt,
  tool description, and result hint.
- Ask before editing protected or git-tracked product files. Reports and experiment
  ledgers go under `.libragent/work/`.

## Cycle

### 1. Freeze the experiment contract

Record:

- job paths and git revision
- dataset, included tasks, attempts, and concurrency
- assistant ID/version, execution mode, workspace mode
- timeout/resource and verifier settings

If these cannot be established, analyze the run but mark cross-run conclusions
as unverified.

### 2. Inventory and validate artifacts

Collect recursively:

- job `result.json`
- trial `verifier/reward.txt`
- agent `trajectory.json` (ATIF)
- relevant adapter/session logs when present

Run the bundled analyzer:

```bash
python .agents/skills/harbor-harness-improvement-loop/scripts/analyze_harbor_results.py \
  jobs/<baseline> [jobs/<candidate>] \
  --output .libragent/work/harbor-harness-cycle/summary.json
```

The script produces descriptive metrics and explicitly labels heuristic error
signals. It does not establish causality.

Before interpretation, flag:

- missing or invalid trajectories
- reward without completed trajectory
- absent token data
- workspace-mode mismatch
- errors, cancelled/incomplete runs, and verifier failures
- unequal task or attempt coverage between runs

### 3. Analyze outcomes and traces

Analyze successful and failed trials separately. For each representative trace:

1. Reconstruct the sequence: intent → tool call → observation → next tool.
2. Find the **first divergence** from an efficient successful path (tool pattern).
3. Count tool selection, invalid arguments, retries, repeated calls, oversized
   outputs, missing verification, premature completion, and unrecovered errors.
4. Compare tool sequences with successful traces for the same task family.
5. Check token/turn/tool-call distributions; do not rely on means alone.

Read [references/evidence-model.md](references/evidence-model.md) before assigning
a root cause.

### 4. Map each symptom to the owning layer

Use the narrowest layer that can fix the observed contract:

| Symptom                                                  | Inspect first                                                      |
| -------------------------------------------------------- | ------------------------------------------------------------------ |
| Correct tool is unavailable or consistently not selected | tool exposure, name, description, input schema                     |
| Tool selected with invalid/missing arguments             | schema required fields, enums, examples, validation error          |
| Same failed action repeats                               | error recovery hint, state feedback, loop/escalation prompt        |
| Tool succeeds but result misleads or bloats context      | handler output, truncation/pagination, structured content, hints   |
| Agent uses many micro-tools for one outcome              | tool boundaries, consolidation, backend automation                 |
| Agent skips planning/verification across unrelated tools | assistant/system/workspace prompt                                  |
| Prompt tokens grow or cache ratio degrades               | stable/volatile prompt split, tool schema volume, repeated context |
| Correct actions still produce wrong workspace state      | handler semantics, isolation/sync, process lifecycle               |
| Metrics are missing or contradictory                     | Harbor adapter, Session API telemetry, aggregation script          |

Inspect the actual schema, dispatcher, handler, response builder, prompt assembly,
and tests for the suspected layer. Use `critique-builtin-tool`,
`lean-builtin-tool-auditor`, or `refactor-builtin-tool` when the evidence points
to builtin implementation details.

### 5. Form and rank hypotheses

For every proposal record:

- observed evidence and affected trial count
- successful-trace counterexample check
- owning layer and inspected code paths
- mechanism: why the change should alter behavior
- expected measurable effect
- regression and benchmark-overfitting risk
- confidence: high / medium / low

Prefer broad, repeated, high-confidence defects. Keep low-confidence ideas in a
backlog. Do not convert all correlations into work items.

### 6. Choose the smallest controlled intervention

Change one causal variable per experiment when practical. Examples:

- clarify one ambiguous schema field
- make one error response actionable and state-aware
- trim or paginate one oversized result
- hide one internal tool
- move one repeated instruction to its owning layer
- add missing telemetry without changing agent behavior

Define acceptance criteria before editing. Preserve a held-out task set to detect
overfitting.

### 7. Implement and verify

After user approval:

1. Apply the minimal change.
2. Run focused unit/contract tests.
3. Run focused lint or formatting checks if necessary. Never run the full `pnpm refactor:validate` pipeline automatically (run only if explicitly requested by the user to avoid system resource starvation).
4. Rerun a small matched task slice.
5. If the signal is positive and no regression appears, rerun the broader suite
   with multiple attempts.

Do not call a change successful from reward alone. Compare:

- task reward/pass rate and errors
- turns and total/per-task tool calls
- input/output/cache tokens
- repeated/failed calls and recovery rate
- completion integrity and workspace correctness

### 8. Record and continue

Write each cycle under:

```text
.libragent/work/harbor-harness-cycle/<cycle-id>/
├── experiment.md
├── baseline-summary.json
├── candidate-summary.json
└── analysis.md
```

Use [references/report-template.md](references/report-template.md). End each cycle
with one decision: **adopt**, **revert**, **iterate**, or **inconclusive**.

## Stop conditions

Stop and report instead of changing code when:

- baseline and candidate are not comparable
- telemetry is insufficient to localize the issue
- the proposed fix depends on benchmark-specific knowledge
- the only evidence is hidden reasoning text
- variance exceeds the observed difference
- the tool pattern is clear but no schema/handler/response/adapter contract is broken
  (report the pattern; do not invent an owner outside the harness)

Attribution

fritzprixfritzprix
View sourceMore from fritzprix →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →