Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Toolformer Pattern

ASecurity

Teach language models to decide when to call tools — API invocation learned from self-supervised execution feedback.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgoapi

Works with

api

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill toolformer-pattern --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Toolformer Pattern?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Toolformer Pattern
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-toolformer-pattern/badge)](https://www.skillsdirectory.com/skills/aicodedecode-toolformer-pattern)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: toolformer-pattern
description: Teach language models to decide when to call tools — API invocation learned from self-supervised execution feedback.
category: ai-research
---

## Overview

The Toolformer pattern trains a language model to use external tools — calculators,
search engines, calendars, code interpreters — by deciding for itself when a tool
call would help. The original method: annotate text with candidate API calls,
execute them, and keep only the calls whose results actually improve the model's
predictions. The model learns a policy over tool use from this filtered data: not
just how to format calls, but when they're worth making.

As a design pattern (beyond the original paper), "Toolformer-style" means any
system where tool use is learned or prompted as a decision, not hardcoded. The
model sees tool descriptions, generates calls interleaved with text using special
tokens or structured formats, receives execution results, and continues. Modern
agent frameworks implement this with function calling, but the underlying
questions are Toolformer's: which tools, when to call, how to handle results,
what to do when calls fail.

The pattern's enduring lesson: tool use is a decision under uncertainty. The model
should call tools when the expected information gain exceeds the cost — and that
judgment can be learned, not just scripted.

## When to use

- Building agents that need external capabilities: computation, retrieval,
  real-time data, actions in the world.
- Deciding between hardcoded tool orchestration (your code decides) vs.
  model-driven tool use (the model decides).
- Training or fine-tuning a model to use a specific API reliably.
- Evaluating whether tool augmentation actually helps a task vs. adding failure
  modes.
- Designing the tool interface itself — Toolformer-style thinking clarifies what
  the model needs to know about each tool.
- Auditing an existing agent's tool use: is it calling tools when they help, or
  out of habit?

## Core concepts

- **The tool-use decision**: the core learned behavior. Good tool use means calling
  when the tool adds information the model lacks (current facts, exact
  computation) and skipping when it doesn't. Over-calling wastes latency and
  money; under-calling leaves capability on the table.
- **Self-supervised filtering**: the original Toolformer insight — generate
  candidate calls, execute them, keep the ones that reduce loss on the true
  continuation. This creates training data for tool use without human annotation
  of "correct" calls.
- **Call format**: special tokens (`<API>calculator(2+2)</API>`), JSON function
  calls, or markdown code blocks. The format must be unambiguous to parse and hard
  for the model to emit accidentally in normal text.
- **Result incorporation**: tool outputs re-enter the context. Design result
  schemas for model consumption: concise, structured, with errors clearly marked.
  A 10KB raw API dump is not a result; it's a context bomb.
- **Failure handling**: tools fail — timeouts, bad arguments, empty results. The
  policy needs a failure branch: retry with fixed args, try an alternative tool,
  or proceed without the tool. Train or prompt for this explicitly.
- **Tool descriptions**: the model's only knowledge of what tools do. Write them
  like API docs for a smart but literal junior: purpose, when to use, argument
  semantics, what the output looks like, known limitations.
- **Learned vs. prompted**: prompted tool use (instructions + examples) is fastest
  to iterate; learned tool use (fine-tuned on filtered trajectories) is most
  reliable at scale. Most production systems start prompted and graduate the
  critical decisions to learned.
- **Tool-use metrics**: call precision (were calls useful?), argument correctness,
  failure recovery rate, and the counterfactual — task success with vs. without
  tools. Measure the decisions, not just the outcomes.

## Practical workflow

1. **Inventory candidate tools.** List every external capability the task might
   need. For each: is the information truly unavailable to the model? If the model
   can answer from weights, the tool is overhead.
2. **Write model-facing tool docs.** One paragraph per tool: what it does, when to
   call it, exact argument format with an example, output shape, failure modes.
3. **Choose the decision mechanism.** Prompted (instructions + examples in context
   — fastest to iterate) vs. fine-tuned (Toolformer-style data — most reliable at
   scale). Start prompted; fine-tune the decision policy once the toolset
   stabilizes.
4. **Build the execution loop.** Parse calls → validate arguments → execute with
   timeouts → format results → continue generation. Log every call: what was
   requested, what returned, whether it helped.
5. **Create filtered training data (for fine-tuning).** Sample tasks, let the
   model attempt tool calls, execute them, and keep trajectories where tool use
   improved the outcome. This is the self-supervised signal.
6. **Evaluate tool-use quality separately.** Metrics: call precision (were calls
   useful?), argument correctness, failure recovery rate, and end-task success with
   vs. without tools. A model that calls tools constantly but no better than
   baseline has learned theater, not tool use.
7. **Harden the failure paths.** Test timeouts, malformed arguments, empty
   results, and rate limits explicitly. The failure branch needs as much design
   as the happy path.

Checklist for shipping model-driven tool use:
- Tool docs reviewed by someone who didn't write the tools.
- Argument validation and timeouts on every call path.
- Failure branch tested: bad args, timeouts, empty results.
- Call logs reviewed for over/under-calling patterns.
- Tool-use metrics (precision, recovery) tracked in production.

## Common pitfalls

- **Tool sprawl.** Twenty overlapping tools paralyze the decision. Consolidate to
  the minimal set; the model chooses better among five clear tools than twenty
  fuzzy ones.
- **Vague tool descriptions.** "Searches the web" is not a spec. The model needs
  to know what the tool returns and when it's better than guessing.
- **No failure branch.** The first production timeout will hang or crash the
  agent. Design the "tool failed, now what" path before the happy path.
- **Result context bombs.** Dumping raw API responses into context wastes tokens
  and confuses the model. Summarize and structure results.
- **Evaluating only end-task success.** A task can succeed despite terrible tool
  use (the model knew the answer anyway). Measure the tool-use decisions directly.
- **Hardcoding what should be learned (and vice versa).** If the tool choice is
  deterministic given the state, hardcode it — don't burn inference on a decision
  with one right answer. Reserve model-driven choice for genuinely ambiguous
  cases.
- **Prompted forever.** Staying on prompted tool use after the toolset stabilized,
  paying the reliability cost indefinitely. Graduate to fine-tuned policies for
  high-volume decisions.
- **No cost accounting.** Tool calls cost latency and money (and sometimes real
  API fees). Budget per-task tool spend and alert on over-calling regressions.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →