Teach language models to decide when to call tools — API invocation learned from self-supervised execution feedback.
Scanned 9/29/2026
npx -y skills add aicodedecode/awesome-muse-skills --skill toolformer-pattern --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Toolformer Pattern?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/aicodedecode-toolformer-pattern)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: toolformer-pattern
description: Teach language models to decide when to call tools — API invocation learned from self-supervised execution feedback.
category: ai-research
---
## Overview
The Toolformer pattern trains a language model to use external tools — calculators,
search engines, calendars, code interpreters — by deciding for itself when a tool
call would help. The original method: annotate text with candidate API calls,
execute them, and keep only the calls whose results actually improve the model's
predictions. The model learns a policy over tool use from this filtered data: not
just how to format calls, but when they're worth making.
As a design pattern (beyond the original paper), "Toolformer-style" means any
system where tool use is learned or prompted as a decision, not hardcoded. The
model sees tool descriptions, generates calls interleaved with text using special
tokens or structured formats, receives execution results, and continues. Modern
agent frameworks implement this with function calling, but the underlying
questions are Toolformer's: which tools, when to call, how to handle results,
what to do when calls fail.
The pattern's enduring lesson: tool use is a decision under uncertainty. The model
should call tools when the expected information gain exceeds the cost — and that
judgment can be learned, not just scripted.
## When to use
- Building agents that need external capabilities: computation, retrieval,
real-time data, actions in the world.
- Deciding between hardcoded tool orchestration (your code decides) vs.
model-driven tool use (the model decides).
- Training or fine-tuning a model to use a specific API reliably.
- Evaluating whether tool augmentation actually helps a task vs. adding failure
modes.
- Designing the tool interface itself — Toolformer-style thinking clarifies what
the model needs to know about each tool.
- Auditing an existing agent's tool use: is it calling tools when they help, or
out of habit?
## Core concepts
- **The tool-use decision**: the core learned behavior. Good tool use means calling
when the tool adds information the model lacks (current facts, exact
computation) and skipping when it doesn't. Over-calling wastes latency and
money; under-calling leaves capability on the table.
- **Self-supervised filtering**: the original Toolformer insight — generate
candidate calls, execute them, keep the ones that reduce loss on the true
continuation. This creates training data for tool use without human annotation
of "correct" calls.
- **Call format**: special tokens (`<API>calculator(2+2)</API>`), JSON function
calls, or markdown code blocks. The format must be unambiguous to parse and hard
for the model to emit accidentally in normal text.
- **Result incorporation**: tool outputs re-enter the context. Design result
schemas for model consumption: concise, structured, with errors clearly marked.
A 10KB raw API dump is not a result; it's a context bomb.
- **Failure handling**: tools fail — timeouts, bad arguments, empty results. The
policy needs a failure branch: retry with fixed args, try an alternative tool,
or proceed without the tool. Train or prompt for this explicitly.
- **Tool descriptions**: the model's only knowledge of what tools do. Write them
like API docs for a smart but literal junior: purpose, when to use, argument
semantics, what the output looks like, known limitations.
- **Learned vs. prompted**: prompted tool use (instructions + examples) is fastest
to iterate; learned tool use (fine-tuned on filtered trajectories) is most
reliable at scale. Most production systems start prompted and graduate the
critical decisions to learned.
- **Tool-use metrics**: call precision (were calls useful?), argument correctness,
failure recovery rate, and the counterfactual — task success with vs. without
tools. Measure the decisions, not just the outcomes.
## Practical workflow
1. **Inventory candidate tools.** List every external capability the task might
need. For each: is the information truly unavailable to the model? If the model
can answer from weights, the tool is overhead.
2. **Write model-facing tool docs.** One paragraph per tool: what it does, when to
call it, exact argument format with an example, output shape, failure modes.
3. **Choose the decision mechanism.** Prompted (instructions + examples in context
— fastest to iterate) vs. fine-tuned (Toolformer-style data — most reliable at
scale). Start prompted; fine-tune the decision policy once the toolset
stabilizes.
4. **Build the execution loop.** Parse calls → validate arguments → execute with
timeouts → format results → continue generation. Log every call: what was
requested, what returned, whether it helped.
5. **Create filtered training data (for fine-tuning).** Sample tasks, let the
model attempt tool calls, execute them, and keep trajectories where tool use
improved the outcome. This is the self-supervised signal.
6. **Evaluate tool-use quality separately.** Metrics: call precision (were calls
useful?), argument correctness, failure recovery rate, and end-task success with
vs. without tools. A model that calls tools constantly but no better than
baseline has learned theater, not tool use.
7. **Harden the failure paths.** Test timeouts, malformed arguments, empty
results, and rate limits explicitly. The failure branch needs as much design
as the happy path.
Checklist for shipping model-driven tool use:
- Tool docs reviewed by someone who didn't write the tools.
- Argument validation and timeouts on every call path.
- Failure branch tested: bad args, timeouts, empty results.
- Call logs reviewed for over/under-calling patterns.
- Tool-use metrics (precision, recovery) tracked in production.
## Common pitfalls
- **Tool sprawl.** Twenty overlapping tools paralyze the decision. Consolidate to
the minimal set; the model chooses better among five clear tools than twenty
fuzzy ones.
- **Vague tool descriptions.** "Searches the web" is not a spec. The model needs
to know what the tool returns and when it's better than guessing.
- **No failure branch.** The first production timeout will hang or crash the
agent. Design the "tool failed, now what" path before the happy path.
- **Result context bombs.** Dumping raw API responses into context wastes tokens
and confuses the model. Summarize and structure results.
- **Evaluating only end-task success.** A task can succeed despite terrible tool
use (the model knew the answer anyway). Measure the tool-use decisions directly.
- **Hardcoding what should be learned (and vice versa).** If the tool choice is
deterministic given the state, hardcode it — don't burn inference on a decision
with one right answer. Reserve model-driven choice for genuinely ambiguous
cases.
- **Prompted forever.** Staying on prompted tool use after the toolset stabilized,
paying the reliability cost indefinitely. Graduate to fine-tuned policies for
high-volume decisions.
- **No cost accounting.** Tool calls cost latency and money (and sometimes real
API fees). Budget per-task tool spend and alert on over-calling regressions.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!