Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Upload Method

BSecurity

Process, upload, and register one benchmarked method's results in TabArena. Use this skill whenever a maintainer points at a benchmark run's output directory and wants to turn its raw `results.pkl` files into cached + hosted + registered TabArena artifacts — e.g. "upload this method", "process and upload these results", "host/publish <model>'s results", "register <model> in the leaderboard". By default Claude runs the whole flow itself (inspect → fix `info.py` → process → upload dry-run → rea...

313 stars
0 votes
0 copies
0 views
Added 9/20/2026
documentationpythonrustgoshellbashgitapi

Works with

cliapi

Security Analysis

B75/100
criticalSends environment variables or credentials to an external URL

Scanned 9/20/2026

Install to Claude Code

$npx -y skills add autogluon/tabarena --skill upload-method --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Upload Method?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Upload Method
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/autogluon-upload-method/badge)](https://www.skillsdirectory.com/skills/autogluon-upload-method)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: upload-method
description: Process, upload, and register one benchmarked method's results in TabArena. Use this skill whenever a maintainer points at a benchmark run's output directory and wants to turn its raw `results.pkl` files into cached + hosted + registered TabArena artifacts — e.g. "upload this method", "process and upload these results", "host/publish <model>'s results", "register <model> in the leaderboard". By default Claude runs the whole flow itself (inspect → fix `info.py` → process → upload dry-run → real r2 upload → register in `methods.py`) after first telling the maintainer what it will do; it hands a step off via a command sheet only when the environment can't run it (no `tabarena[benchmark]` venv, no R2 credentials).
argument-hint: <run-data-dir> [<model>]
user-invocable: true
---

# Process & Upload a Benchmarked Method to TabArena

This skill drives the maintainer workflow that turns a benchmark run's **already-present** raw
`results.pkl` files into cached, hosted (r2), and registered TabArena artifacts. There is **no
download and no auto-generation** — the raw data is assumed to be on disk already (e.g. you unzipped
a submission, or it's a fresh run's `output/<run>/data` dir).

The authoritative prose for this flow is **AGENTS.md → "Processing & uploading method artifacts
(maintainers)"** and the two scripts' module docstrings (`scripts/run_process_method.py`,
`scripts/run_upload_results.py`). This skill operationalizes it: **Claude executes the flow itself
by default**, and makes the small deterministic code edits along the way.

## What this skill delivers

1. **A plan, stated first**: which methods, which run dir, which suite, and that the flow ends in a
   real r2 upload — tell the maintainer before executing (no need to wait for approval unless they
   asked to review something first).
2. **Claude-run execution**, in order per method:
   - `inspect` — read the raw data, print inferred fields + a metadata diff.
   - `process` — build + cache `metadata.yaml` + `processed/` + `results/` locally.
   - `upload` (dry-run, then `--no-dry-run`) — push the cached artifacts to r2.
3. **Code edits**:
   - The model's `info.py` `MethodMetadata` — fill in the manual upload fields, fix any
     raw-data mismatches the inspect diff surfaces.
   - The arena collection registration in `methods.py` so the method appears in the benchmark.
   - The run's entry in `packages/tabflow_slurm/BENCHMARK_LOG.md`, if the launch stage never wrote one.
4. **A fallback command sheet** only for steps the environment can't run.

Execution requirements (check before starting; hand off the affected step if missing):
- a venv with **`tabarena[benchmark]`** installed (look under `~/.venvs/tabarena_*`; run scripts with
  its python from the repo root),
- **R2 credentials in the environment** for the real upload (`printenv | grep -o '^R2_[A-Z_]*'`
  should show `R2_ACCOUNT_ID` / `R2_ACCESS_KEY_ID` / `R2_SECRET_ACCESS_KEY`; never pass them as flags).

Processing and uploads are long-running: run them **in the background** (per-method), watch with a
monitor (per-method DONE/FAILED lines + a stall watchdog), and for multi-method batches pipeline the
uploads — start each method's dry-run + real upload as soon as its processing finishes.

## Step 0: Gather inputs

Parse `$ARGUMENTS`. Collect (ask only for what's missing or ambiguous):

| Input | Example | Notes |
|---|---|---|
| `run_data_dir` | `.../output/benchmark_chimeraboost_16062026/data` | Dir of raw `results.pkl` files (searched recursively). Point at the run's `data/`. |
| `model` | `chimeraboost` | The `packages/tabarena/src/tabarena/models/<model>/` folder — or `systems/<system>/` when the method is a whole pipeline (AutoML framework, agent, hosted API). Usually inferable from the run/data dir name; confirm. |
| `suite` | `tabarena-2026-06-30` | The dated run/suite id. **Must differ from `method`** (see Step 3). Default to `tabarena-<run-date>`; ask if unclear. |
| `arena` | `tabarena` | Which arena collection to register in (default `tabarena`; e.g. `beyondarena` has its own). |
| `verified` | `False` until signed off | Whether the results are verified. Default keep `False`; flip to `True` only when the maintainer confirms (Step 3). |

## Step 1: Locate the method's `MethodMetadata`

Read `packages/tabarena/src/tabarena/models/<model>/info.py` (or `systems/<system>/info.py`) and note
the **exact** variable name of its `MethodMetadata` (e.g. `chimeraboost_method_metadata`,
`tabfm_plus_method_metadata`). Both scripts take it as a dotted reference:

```
tabarena.models.<model>.info:<varname>
tabarena.systems.<system>.info:<varname>
```

- **`info.py` exists** (the normal case — the model was added via the `add-model` skill): it already
  carries the raw-inferable fields (`ag_key`, `config_default`, `can_hpo`, `is_bag`, `compute`,
  `method_type`). `validation_protocol` is inferable too but was not authored before the run: a new
  run records it, and `--process` fails until `info.py` declares the inferred key. Claude fills it
  together with the *manual* upload fields in Step 3.
- **No `info.py`** (a raw external submission): the method isn't integrated yet. Run `inspect`
  (Step 2) to get the copy-paste `MethodMetadata.<type>(...)` snippet, then author
  `models/<model>/info.py` from it (use the **`add-model`** skill if the model also needs a wrapper,
  or **`add-system`** if it is a whole pipeline). Processing **requires** an explicit committed
  `MethodMetadata` — it refuses to guess.
- **The method is a system**: the inspector cannot tell a system from any other baseline, since the
  runner records both identically. Use `MethodMetadata.system(...)` rather than `.baseline(...)`, and
  set `tags` — see the `add-system` skill. Getting this wrong types the method as a model on the
  leaderboard and puts it in the models-only entrant pool, where it does not belong.

## Step 2: Inspect the raw data (Claude runs; confirms inference)

```bash
<venv>/bin/python scripts/run_process_method.py <run_data_dir>
```

This prints the fields inferred from the raw data (`method_type`, `ag_key`, `compute`,
`config_default`, `can_hpo`, `is_bag`, `validation_protocol` with a fold histogram,
task/problem-type/metric coverage), a suggested `MethodMetadata` snippet, and — when an `info.py` metadata is passed — an **inferred-vs-provided
diff**. Use it to confirm `info.py` matches the raw data before processing. Two gotchas it catches:

- **`config_default` is compared post-rename**: configs are renamed to the method's prefix during
  processing, so a `config_default` authored with the raw prefix won't match. The snippet/diff shows
  the post-rename value — use that. The most common real mismatch is the infix, not the prefix: the
  TabArena-v0.1 bundle names the first config `<Method>_c1_default_BAG_L1` (the `_default` is the
  preprocessing pipeline name, appended for HPO and default-only models alike), while `info.py` files
  authored before a run, and the BeyondArena bundle, use `<Method>_c1_BAG_L1`. An `info.py` written by
  `add-model` before the run therefore usually needs this one fix. A single-config method
  (`can_hpo=False`) can drop the field instead: an undeclared `config_default` is not checked (the diff
  marks it `not declared`) and `--process` records the lone config in the cached `metadata.yaml`.
- **Only `error` rows gate processing**: the `method` row differs whenever the raw `ag_name` carries
  the `TA-` prefix (warn-only) and `model_key` is never checked (shown for context). A `NO` on either is
  expected and needs no edit; a `NO` on `config_default`, `ag_key`, `compute`, `can_hpo`, `is_bag` or
  `method_type` does.
- **`method != suite`**: `process` fails if they're equal (suite defaults to method when unset).
- **`validation_protocol` must be declared**: a run made after the protocol record existed infers a
  key (`8x1` for TabArena, `system` for a system, the long BeyondArena key); `--process` fails while
  `info.py` leaves it `None` or declares another value. Legacy raw artifacts infer `None` and pass.
  Add `--expect-validation-protocol tabarena` (or `beyondarena`) to also assert that the run used the
  arena's official protocol; a `holdout:`/`outer`/custom key there means the run cannot be uploaded as
  an official result. KNN's `8x1 requested; 1 child via use_child_oof` histogram line is expected.

**Multi-method run dirs**: if the run's `data/` holds several methods' config dirs side by side,
`run_process_method.py` can't split them — write a small gitignored driver in `tmp_scripts/` that
discovers the `results.pkl` paths once, splits them by top-level config-dir prefix, and calls
`_infer_from_raw` / `verify_method_metadata` / `process_raw(file_paths=...)` per method (see the
`multi-method-run-upload` pattern; a `--c1-only` pass over just the `_c1_` dirs makes
`config_default` inferable even for HPO methods with hundreds of configs).

## Step 3: Edit `info.py` (Claude does now)

Fill in the **manual** fields the upload requires. These are not inferable from raw data, so the
maintainer normally hand-edits them — Claude does it now. Read the file first, then `Edit`:

| Field | Set to | Why |
|---|---|---|
| `suite` | the dated run id, e.g. `"tabarena-2026-06-30"` | **Required, must differ from `method`.** Equal pair fails `process`'s `method != suite` check (suite defaults to method when unset). |
| `cache_type` | `"r2"` | **Required for upload** — the upload script is r2-only; `"local"`/`None` has no remote store. |
| `cache_kwargs` | `{"bucket": "tabarena", "prefix": "cache"}` | The r2 location (`bucket` + `prefix`). Required when `cache_type="r2"`. |
| `date` | `"YYYY-MM-DD"` | The run date. Validated as a real calendar date. |
| `verified` | `False` (until signed off), then `True` | Manual trust flag. Keep `False` until the results are verified; flip to `True` once they are (typically the final step). |
| `method_class` / `tags` | systems only | `MethodMetadata.system(...)` sets `method_class`; `tags` (`with-llm`, `closed-source-api`) decide which entrant pools the system competes in. Neither is inferable from raw data. |
| `validation_protocol` | the key the inspect step infers, e.g. `"8x1"` | The inner validation protocol the results record. `--process` requires it for a new run; `run_upload_results.py` refuses an upload whose processed artifact records a key `info.py` does not declare. `MethodMetadata.system(...)` defaults it to `"system"`. |

Leave the other raw-inferable fields (`ag_key`, `config_default`, `can_hpo`, `is_bag`, `compute`,
`method_type`) **as they are** — they came from `add-model`. Only change one if Step 2's diff shows a
genuine mismatch (and then to the inferred value; for a single-config method, deleting a mismatched
`config_default` is equally valid).

Example — making `chimeraboost` upload-ready (the upload fields added; compare `nori`'s `info.py`,
which is already in this shape):

```python
chimeraboost_method_metadata = MethodMetadata.config(
    method="ChimeraBoost",
    suite="tabarena-2026-06-30",                              # added: distinct dated suite
    ag_key="CHIMERA",
    config_default="ChimeraBoost_c1_default_BAG_L1",             # HPO method: post-rename, `_default` = the run's pipeline
    compute="cpu",
    is_bag=False,
    date="2026-06-15",
    reference_url="https://github.com/bbstats/chimeraboost",
    display_name="ChimeraBoost",
    verified=False,                                           # flip to True once verified
    cache_type="r2",                                          # added
    cache_kwargs={"bucket": "tabarena", "prefix": "cache"},   # added
)
```

## Step 4: Execute process → upload (Claude runs, in the background)

Run these with the resolved dotted reference, per method — background + monitor for the slow ones.
The same block doubles as the fallback command sheet if a step must be handed to the maintainer.
Using `chimeraboost` as the example:

```bash
# 1. Inspect (confirm inferred fields match info.py) — already done in Step 2

# 2. (info.py edited: suite + cache_type/cache_kwargs + date, plus any inspect-diff fixes)

# 3. Process: build + cache metadata.yaml + processed/ + results/ locally
<venv>/bin/python scripts/run_process_method.py <run_data_dir> \
    --method-metadata tabarena.models.chimeraboost.info:chimeraboost_method_metadata --process \
    --expect-validation-protocol tabarena  # the run must record TabArena's official protocol

# 4. Upload dry-run: verifies every part exists locally and prints what/where (no creds needed)
<venv>/bin/python scripts/run_upload_results.py \
    --method-metadata tabarena.models.chimeraboost.info:chimeraboost_method_metadata

# 5. Real upload (R2 creds via env, NEVER flags):
<venv>/bin/python scripts/run_upload_results.py \
    --method-metadata tabarena.models.chimeraboost.info:chimeraboost_method_metadata --no-dry-run
```

After the real upload, verify on the bucket (ListObjects via boto3 against the R2 endpoint): each
method should have the full object set under
`cache/artifacts/<suite>/methods/<Method>/` (`metadata.yaml`, `processed.zip`,
`processed/configs_hyperparameters.json`, `raw.zip`, `results/*.parquet`).

Notes:
- `process` requires the explicit `--method-metadata` and verifies it against the raw data first; a
  real mismatch errors (override with `--ignore-metadata-mismatch`, a `method`-name mismatch only
  warns). Raw + HPO trajectories are cached by default (`--no-cache-raw` / `--no-cache-hpo-trajectories`).
- The dry-run is the default for the upload script; it needs no credentials and prints the exact
  `--no-dry-run` command. Raw uploads by default (`--no-upload-raw` to skip).
- For the real upload, set `R2_ACCOUNT_ID` / `R2_ACCESS_KEY_ID` / `R2_SECRET_ACCESS_KEY` in the
  environment (export them or use a `.env`) — never as CLI flags (they'd leak into shell history /
  the process table). The dry-run prints how to obtain them if unset
  (`MethodMetadata.r2_credentials_help()`).

## Step 5: Register in the arena collection (Claude does now)

Add the method to the collection so it appears in the benchmark. For the default `tabarena` arena,
edit `packages/tabarena/src/tabarena/contexts/tabarena/methods.py` (read it first):

1. **Import** the model's `info.py` metadata in the alphabetical import block, matching the file's
   existing style (most recent additions use the plain name, e.g.
   `from tabarena.models.nori.info import nori_method_metadata`; alias only on a name collision).
2. **Add the entry** to `tabarena_method_metadata_collection.method_metadata_lst`, under the matching
   group comment (`# Default tabular models (CPU)` vs `# Neural / GPU / foundation models`), placed by
   compute/type.

It flows into `tabarena_method_metadata_complete_collection` automatically (no separate edit). Other
arenas (e.g. `beyondarena`) register in their own collection's `methods.py`.

**When the upload is a rerun of an already-registered method** (a new library version, a remeasured
run), the entry is a *swap*, not an addition: point the collection at the new metadata and append the
one it replaced to `methods_superseded`, so the predecessor's hosted artifacts stay reachable through
the complete collection. Say in the report that the default leaderboard now carries only the new run.

**Caveat to surface in the report:** the collection entry only resolves to *downloadable* artifacts
once the real upload (Step 4.5) has actually run. The code edit is safe to land now, but the method
won't load for others until uploaded.

## Step 6: Record the run in the benchmark log (Claude does now)

`packages/tabflow_slurm/BENCHMARK_LOG.md` is the committed record of every cluster run (the
`tmp_scripts/run_<model>.py` that launched it is gitignored, so this log is the only surviving copy of
the setup). The `benchmark-model` skill merely *offers* to write the entry at launch time, before any
results exist, so in practice runs reach upload unlogged. Check here, where the run is finished:

1. `grep -n "^## " packages/tabflow_slurm/BENCHMARK_LOG.md | head` — if the run's `benchmark_name`
   is already there, nothing to do.
2. Otherwise add an entry at the **top** of the log (append-only, newest-first), following the
   template in the file's "Conventions" section: `## YYYY-MM-DD — <benchmark_name>`, then
   **Model(s)** / **Git SHA** / **Purpose** / **Notes**, then the verbatim plan.

Two fields need care:

- **Git SHA** is HEAD at *setup* time, not HEAD now. `git reflog --date=iso` shows when HEAD sat on
  which commit; bracket the launch with the run's own timestamps (the earliest `results.pkl` mtime
  under the run's `data/`, or the earliest file in `<workspace>/slurm_out/<benchmark_name>/`) and pick
  the commit that was HEAD then.
- **The python block** is the launch script's `setup()` body copied verbatim, with its module-level
  constants (`BENCHMARK_NAME`, `WORKSPACE`, `PYTHON_PATH`, `MODEL`, `NUM_CONFIGS`) inlined as
  literals and `_path_setup()` expanded, so the snippet stands alone against its SHA. Never refactor
  a neighbouring entry to match the current API.

The launch script's module docstring usually holds the *why* (issue link, what changed versus the
previous run, partition choice, model-specific caveats) — that is the Purpose/Notes material. Add the
processed suite id and the metadata variable name too, so the log ties the run to its artifacts.

## Step 7: Lint touched files

```bash
ruff check <touched-files>
ruff format --check <touched-files>
```

Touched files are the model's `info.py` and `contexts/<arena>/methods.py`. Fix anything reported
(the `from __future__ import annotations` import is already present in both — don't drop it).
`BENCHMARK_LOG.md` is markdown, so ruff does not apply to it.

## Step 8: Report

Tell the maintainer:

- **What was executed and verified**: inspect diff results, processing outcome per method, and the
  r2 destinations confirmed after the real upload — plus any data-quality warnings from processing
  (e.g. `Not close TEST` prediction-fidelity lines, with affected datasets and severity).
- **Edits Claude made**: the `info.py` upload fields (suite / cache_type / cache_kwargs / date /
  verified, plus any inspect-diff fixes), the `methods.py` import + collection entry (and the
  `methods_superseded` append, for a rerun), and the `BENCHMARK_LOG.md` entry (with the SHA it
  recorded and how it was determined, since that one is inferred).
- **Steps handed off** (only if the env couldn't run one): the exact command(s) from Step 4.
- **Open decisions / TODOs**: whether to flip `verified` to `True` (only after sign-off), committing
  the working-tree edits, and — if the method should appear on the website — that `update-leaderboard`
  is the next lifecycle step.

When asked to open the PR, use `.github/pull_request_template.md` and delete its "Model or system
submission" section (that one is for external submissions). The summary names the method and what
was uploaded; the collapsed Details block carries the method key, suite and date, the r2
destinations, the `verified` state, and whether `update-leaderboard` is the next step. Do not paste
this Report into the PR body.

Attribution

autogluonautogluon
View sourceMore from autogluon →
SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Your tool, in front of Claude Code builders.

3 founder slots · $299/mo · GSC-verified traffic · sponsors can never buy grades.

See placements

Related Skills

Context Fundamentals

Understand the components, mechanics, and constraints of context in agent systems. Use when designing agent architectures, debugging context-related failures, or optimizing context usage.

179001 votes

release-notes

Draft release notes and changelog entries from git history or merged PRs between two refs (tags/SHAs/branches), including breaking changes, migrations, and upgrade steps. Use when the user asks for release notes, changelog updates, or a GitHub Release draft.

1301 votes

docs-style-guide

Documentation style guide enforcer by @planetabhi. Applies and reviews the writing style guide when authoring or editing product documentation and tutorials. Use to check prose for voice, tense, word choice, inclusive language, formatting, code block, UI, Markdown, and number/date conventions.

11 votes

Caveman Help

Quick-reference card for all caveman modes, skills, and commands. One-shot display, not a persistent mode. Trigger: /caveman-help, "caveman help", "what caveman commands", "how do I use caveman".

1023330 votes

How It Works

Explain how claude-mem captures observations, when memory injection kicks in, and where data lives. Use when the user asks "how does claude-mem work?" or "what is this thing doing?".

929660 votes
View all in documentation →