Best practices for leading Ask compete and bakeoff workflows. Use when a user asks for competing models, isolated candidate implementations, winner selection, feature harvesting, approach comparison, model bakeoffs, or a creator competition where $ask should route browser and API handlers through Tau and the project agent must judge results against local evidence.
Scanned 9/11/2026
Install to Claude Code
npx -y skills add grahama1970/agent-skills --skill best-practices-competition --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Best Practices Competition?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/grahama1970-best-practices-competition)More formats (shields.io, HTML) on the badges page.
---
name: best-practices-competition
description: >
Best practices for leading Ask compete and bakeoff workflows. Use when a
user asks for competing models, isolated candidate implementations,
winner selection, feature harvesting, approach comparison, model bakeoffs,
or a creator competition where $ask should route browser and API handlers
through Tau and the project agent must judge results against local evidence.
triggers:
- competition best practices
- lead a competition
- Ask compete
- model competition
- implementation bakeoff
- competing models
- isolated candidates
- winner selection
- feature harvesting
- compare candidate implementations
provides:
- competition-leadership-protocol
- isolated-candidate-contract
- winner-selection-gate
- feature-harvesting-contract
- winner-revision-request-pattern
composes:
- ask
- best-practices-tau-dag
- brave-search
- github-search
- dogpile
- agentic-evals
complies:
- best-practices-skills
- best-practices-tau-dag
taxonomy:
- orchestration
- competition
- review
- validation
disciplines:
- engineering-standards
- agentic-orchestration
---
# Best Practices: Competition
Use this skill to run a competition that produces usable engineering signal:
isolated candidates, identical task packets, explicit judging criteria, local
verification, and a bounded revision request for the selected winner.
A competition is not a roundtable. Competitors do not deliberate with one
another, and the project agent does not share participant information between
candidate lanes or rounds. The project agent judges the work against the
codebase, skill contracts, and deterministic evidence.
## Core Rule
Compete candidates are claims until checked.
The project agent owns the competition contract, candidate isolation, artifact
collection, feature verification, scorecard, winner decision, and final
winner-only continuation. `$ask` owns the runtime entrypoint and should compile
the competition into Tau-owned candidate nodes and a join node. For iterative
competitions, the control plane is a dynamically expanding Tau DAG: each round
adds the next candidate or winner-continuation nodes under the same immutable
goal, or launches an explicitly linked next-round Tau DAG when the installed
runtime cannot append nodes in place. `$tau`, `$surf`, `$browser-oracle`, and
`$scillm` own transport and provider execution.
Browser-backed handlers and API-backed handlers are peers. Do not privilege or
discard a candidate merely because it arrived through a different transport.
The project agent may coach each participant between its own iterations with
the project agent's review of that participant's output, local deterministic
findings, and external research from `$brave-search`, `$github-search`, or
`$dogpile`. That help must stay lane-local unless it is shared public evidence
added to a fresh common task packet for every candidate. Never leak another
candidate's approach, code, score, feature ideas, or failure analysis.
## Use This When
- The user asks for a competition, compete run, bakeoff, isolated candidates, a
winner, or "best of several implementations."
- The task has multiple plausible implementation paths and comparison will
improve the final patch.
- The candidates may mix browser seats such as `webgpt`, `webclaude`,
`webkimi`, or `webgemini` with `$scillm` model seats.
- The project agent can locally inspect candidate output and verify reusable
features before promoting them.
## Do Not Use This When
| Request shape | Better route |
| --- | --- |
| Shared deliberation or synthesis | `$ask` roundtable with `$best-practices-roundtable` |
| N independent answers the human reads, no winner | `$ask one-shot` with `$best-practices-one-shot` |
| One answer from one handler | `$ask` single handler |
| Creator then pass/fail reviewer | `$ask tau-dag --topology sequential` |
| A deterministic repair is already obvious | Apply and test the repair directly |
| The judge cannot inspect the target repo or artifact | Stop with `NEEDS_ATTENTION` |
| The user needs a human policy decision | Ask the human directly |
## Source-Derived Step Model
1. **Define the competition contract.**
State the objective, immutable goal or acceptance bar, target repo/path,
allowed files, candidate count, judging criteria, output schema, and proof
boundary.
2. **Freeze the shared task packet.**
Every candidate receives the same task, context, constraints, and expected
output. Do not tailor hidden context to favor one handler.
3. **Choose isolated candidates.**
Use at least two handlers. Mix browser and API handlers when useful. Treat
`webclaude` as a strong default candidate for hard work; when available in
`$ask`, prefer `Opus 5 High` for that browser seat.
4. **Compile through `$ask compete`.**
Use `$ask` as the front door. The competition should become a Tau DAG with
concurrent candidate nodes and a join node. If the competition spans more
than one round, preserve the same immutable goal hash and link each
next-round DAG to the previous round's receipts.
5. **Preserve all receipts.**
Read `request.json`, `dag.json`, command specs, candidate receipts,
candidate responses, scorecard, and winner revision request. Missing,
blocked, stale-tab, or rate-limited candidates must be visible in the
scorecard.
6. **Normalize candidate outputs.**
Extract concrete changes, proposed files, test commands, risk notes, and
`VERIFIED_FEATURE:` claims. Do not let prose style or confidence decide the
winner.
7. **Review each candidate iteration.**
For each participant, inspect only that participant's artifacts plus the
original task packet, local repo state, and tool-backed research. Use
`$brave-search`, `$github-search`, or `$dogpile` to unblock the participant
when research can answer a concrete implementation question. Send feedback
back only to that same lane.
8. **Verify reusable features locally.**
A candidate feature is promotable only after the project agent checks it
against repository state, skill contracts, and the narrow deterministic proof
command. Unchecked features stay out of the winner request.
9. **Harvest useful features after N rounds.**
After the configured round count, or earlier if the evidence clearly
converges, decide feature-by-feature what is useful. Losing participants may
provide no useful ideas, one useful feature, or several useful features; the
project agent decides from local evidence and records accepted, rejected, and
unchecked feature claims. Do not share the harvested list with candidates
until the isolated competition phase has closed.
10. **Score with evidence.**
Score criteria such as correctness, minimality, maintainability, contract
fit, proof quality, and failure handling. Tie every score to artifacts or
local checks.
11. **Pick a winner or fail closed.**
Pick a winner only when one candidate has a clear evidence-backed advantage.
If candidates are tied, incomplete, blocked, unverifiable, or all wrong,
report `NEEDS_ATTENTION` instead of fabricating a winner.
12. **Continue with the winning participant.**
Once the winner is chosen, stop the broad competition and continue
iterating with the winning participant until the immutable goal is met or a
real `NEEDS_ATTENTION` blocker is recorded. Ask the winner to keep its own
implementation as the base and add only locally verified features harvested
from other candidates. The winner-continuation request is a next step, not
proof that the revision happened.
13. **Verify the final implementation locally.**
Competition output can guide the patch, but closure requires local
deterministic evidence appropriate to the task.
## Competition Packet
Every substantial competition prompt should include:
```text
Objective:
Immutable goal or acceptance bar:
Target repo/path:
Allowed files or boundaries:
Shared context:
Candidate handlers:
Judging criteria:
Expected candidate output:
Forbidden claims:
Proof boundary:
```
Ask every candidate for:
- `APPROACH`: concise strategy.
- `CHANGES`: specific files, functions, commands, or artifacts.
- `VERIFIED_FEATURE`: only features that the project agent can check locally.
- `RISKS`: failure modes, omitted cases, and assumptions.
- `PROOF_COMMANDS`: exact commands the project agent should run.
- `BLOCKERS`: only missing input, credentials, or external state.
## Iteration And Research Rules
The project agent may run iterative competitions, but isolation still applies.
Treat those iterations as a dynamically expanding Tau DAG: round 1 creates
isolated candidate nodes and a join node; later rounds add lane-local repair
nodes, fresh common-task candidate nodes, or a winner-continuation node with
explicit dependencies on the prior receipts. If the current `$ask` or `$tau`
runtime cannot mutate an existing DAG, the project agent must launch a linked
next-round DAG with the same immutable goal hash and cite the previous run
directory as input evidence.
Allowed lane-local help:
- ask a participant to repair its own failed proof gate;
- give a participant the project agent's review of that participant's output;
- provide deterministic local errors from that participant's attempted patch;
- run `$brave-search web`, `$github-search`, or `$dogpile` for a concrete
blocker and provide the retrieved public evidence to that participant;
- update all participants with a fresh common task packet when the same public
evidence should apply to everyone.
Forbidden cross-lane leakage:
- another participant's code, approach, prompt, score, or review;
- "candidate B solved this by..." hints;
- merged feature lists before winner selection;
- using one participant as an uncredited reviewer of another participant;
- sharing failure analysis from one lane with another lane unless the human
explicitly converts the competition into a roundtable.
If cross-lane leakage happens, mark the competition contaminated and restart
from a fresh shared packet or convert it to a roundtable with human approval.
## Winner Continuation
After N rounds, do not keep all candidates alive by default. The project agent
must choose a clear winner when evidence supports one, then continue iterating
with the winning participant until the immutable goal has been met.
Winner continuation packet:
```text
Winning participant:
Immutable goal:
Winner base to keep:
Verified features to add:
Rejected or unchecked features to exclude:
Required proof commands:
Stop condition:
```
Only locally verified harvested features may be included. A losing participant
may or may not contribute useful ideas; the project agent decides feature by
feature and records the reason. If the winner cannot make progress after a
focused continuation attempt, either run one explicit fallback round with the
remaining candidates or report `NEEDS_ATTENTION` with the failing evidence.
## Count candidates that ANSWERED, not candidates dispatched
"Fewer than two candidates" must be measured on answers. Dispatching two seats
and receiving one is a single opinion wearing a competition's artifacts.
Observed 2026-08-16: a compete run dispatched `webgpt` and `webclaude`, one
answered, and the scorecard still read `candidates: 2`. That run stayed honest
only because it also reported `NEEDS_ATTENTION`; a scorecard naming a winner
off one answer would have been indistinguishable from a real competition.
```bash
skills/ask/run.sh panel-audit <run-dir> --mode compete
```
The audit fails the run when fewer than two candidates produced a non-empty
response and the status is not already `NEEDS_ATTENTION`/`BLOCKED`.
## Isolation is checkable, so check it
Do not assert isolation from the fact that you did not intend to leak. Look for
a rival's response text inside each candidate's prompt -- that is the shape
leakage actually takes when a join, a retry, or a recovery path rebuilds a
packet from prior artifacts.
Compare task bodies too: every candidate must receive a byte-identical packet
once per-seat addressing (`Handler:`, `Model:`, `Browser model preference:`) is
removed. A packet tailored to one candidate is a rigged competition even when
the tailoring looks harmless.
## A transport failure is not a candidate verdict
A candidate that died before reaching its provider has produced no evidence
about its approach. Record it as a transport blocker with its exact failure,
never as a weak entry.
This matters more than it sounds: on 2026-08-16 every non-`webgpt` browser seat
was failing on a CLI usage error (`unrecognized arguments: --stable-stall-ms`)
before opening a page. Read as candidate quality, that would have "proved"
webgpt superior across every competition ever run on this machine. Always
separate `did not run` from `ran and lost`.
## Charts before and after the run
Compile first and show the human the DAG chart before any multi-candidate
`--execute`: `$ask` prints it at compile and persists it as
`dag-chart.initial.txt` (candidates, join-gate, reviewer lenses, judge, join
all visible). After the run, read `dag-chart.final.txt` — per-node verdicts
on the same topology — and reconcile it with the scorecard: a candidate's
node line must agree with its lane artifacts, and a judge node reading FAIL
or NO_RECEIPT invalidates any winner claim.
The final chart is also the project agent's SELF-CORRECTION instrument: walk
its node lines before writing the scorecard. Every non-PASS node is a work
item -- read that lane's receipts and fix, rerun, or name the blocker before
any verdict is reported. Both artifacts are eval-enforced in `$ask`.
## Scoring Contract
The project agent's scorecard must include:
- candidate status: responded, blocked, stale tab, timed out, rate-limited, or
not run;
- artifact paths for every candidate;
- feature claims accepted, rejected, or unchecked;
- harvest decision for every useful feature considered from every candidate;
- local checks run by the project agent;
- criterion scores with evidence;
- winner, tie, or `NEEDS_ATTENTION`;
- bounded winner-continuation request and status, if a winner exists.
Do not score a candidate higher for sounding confident. Score only what can be
reconciled with the task packet, codebase, skill contracts, and proof artifacts.
## Ask Integration
Use `$ask` for execution. Do not replace competitions with informal subagents or
manual browser prompts.
Compile-only example:
```bash
cd skills/ask
./run.sh compete "Implement the focused patch. Return concrete reusable features as VERIFIED_FEATURE lines only when locally checkable." \
--repo local/agent-skills \
--target ask-competition-example \
--handler webgpt \
--handler webclaude \
--handler gpt-5.5-high \
--handler-project webgpt=tau \
--criterion skill-contract \
--criterion deterministic-proof \
--json
```
Add `--execute` only when live browser/provider calls are authorized and the
required browser-oracle bindings or provider credentials are available.
## Fail-Closed Rules
| Condition | Behavior |
| --- | --- |
| Fewer than two candidates | Request another handler or emit `NEEDS_ATTENTION` |
| Candidate sees another candidate's output or score | Restart; isolation was broken |
| Missing candidate receipt | Mark that candidate `NEEDS_ATTENTION` |
| Candidate feature is not locally checkable | Do not promote it |
| All candidates fail the same gate | Stop the family and report systemic failure |
| No clear evidence-backed winner | Report tie or `NEEDS_ATTENTION` |
| Winner-continuation request exists | Treat it as a request packet, not final proof |
## Common Failure Modes
| Failure mode | Required correction |
| --- | --- |
| Roundtable disguised as competition | Re-run as isolated `$ask compete` |
| Consensus substituted for judging | Use roundtable, or verify features locally |
| Cross-lane coaching | Restart or convert to roundtable with human approval |
| Tool help omitted for a research blocker | Use `$brave-search`, `$github-search`, or `$dogpile` lane-locally |
| Feature harvesting without proof | Remove unchecked features from the winner request |
| Winner picked from style | Score against criteria and artifacts |
| Over-broad final revision | Bound the request to verified features only |
| Ignored losing candidate insight | Preserve rejected and accepted feature reasons |
| Competition continues after clear winner | Close competition and iterate with the winner |
| Browser/API transport hidden | Mark per-candidate transport status explicitly |
## Closure Boundary
A competition can close its selection phase when it has:
- a frozen shared packet;
- isolated candidate receipts;
- a scorecard with local checks;
- explicit accepted/rejected/unchecked feature claims;
- a winner, tie, or `NEEDS_ATTENTION`;
- a bounded winner-continuation request when a winner exists.
It cannot close the user's immutable goal by itself. Final closure requires the
selected implementation to be applied and locally proven against that immutable
goal.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!