Skip to content
Back to skills

Costing A Model Vs An Api

ASecurity

Decide whether running your own model actually costs less than paying a hosted API, and at what volume the two cross over. Gathers the real inputs, token counts per request, today's provider pricing, hosting or hardware cost, and one-off training cost, then computes the breakeven point and reports it with the assumptions visible, so the number can be challenged. Use when someone asks whether self-hosting is cheaper, what a fine-tune would save, when owning a model pays for itself, or wants to...

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agentspythonrustgobashreactgitapi

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned September 19, 2026

npx -y skills add ErtasAI/open-model-skills --skill costing-a-model-vs-an-api --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Costing A Model Vs An Api?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Costing A Model Vs An Api
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ertasai-costing-a-model-vs-an-api/badge)](https://www.skillsdirectory.com/skills/ertasai-costing-a-model-vs-an-api)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: costing-a-model-vs-an-api
description: >-
  Decide whether running your own model actually costs less than paying a
  hosted API, and at what volume the two cross over. Gathers the real inputs,
  token counts per request, today's provider pricing, hosting or hardware
  cost, and one-off training cost, then computes the breakeven point and
  reports it with the assumptions visible, so the number can be challenged.
  Use when someone asks whether self-hosting is cheaper, what a fine-tune
  would save, when owning a model pays for itself, or wants to justify or
  reject a move off a hosted API. Not for deciding whether the model is good
  enough, which is evaluating-a-tuned-model, and not for choosing a
  deployment target.
license: Apache-2.0
metadata:
  version: "0.1.0"
  author: "Edward Xi Yang, Ertas AI"
  stage: "cost"
  previous_skill: "evaluating-a-tuned-model"
---

# Costing a model vs an API

This skill exists to be trustworthy, not persuasive. It answers one question
honestly: at this project's actual usage, does owning the model cost less
than paying an API per API call, and if not yet, at what volume would it.
The honest answer is often "no, stay on the API." Most projects never reach
the volume where owning pays for itself, and there is no shame in that. A
crossover calculation that only ever concludes in favour of owning is not
measuring anything, it is decorating a decision that was already made.

## Gather the real inputs before computing anything

Six numbers drive the whole result. Guessing any of them produces a
confident-looking number that means nothing:

1. **Tokens in and out per request**, from this project's actual traffic, and
   counting the whole payload: system prompt, tool schemas, retrieved passages,
   injected state, and history, not only the user's message.
2. **Current API pricing** for the model actually being compared against,
   per million input and output tokens.
3. **Monthly hosting cost**, if the owned model runs on rented
   infrastructure. **Zero if it runs on the user's own device**, where there
   is no hosting bill and no per request inference cost at all.
4. **One off training cost**, if a fine-tune is part of the plan. Usually far
   smaller than people assume: a parameter-efficient run on a small model is
   often single-digit dollars, not hundreds.
5. **The amortisation window**, how many months this version of the model
   will serve before it is retrained or replaced by the next version.
6. **Expected or actual monthly request volume**, to know where this
   project actually sits relative to the breakeven point.

**Ask for these before computing, and offer a default for each.** Someone who
has not built the thing yet will not know their payload size or which API they
are comparing against, and telling them to go and measure first leaves them with
no number at all. Ask the questions, take whatever answers they have, and fill
the rest from the defaults below.

| Ask | If they do not know |
|---|---|
| Which model or API are you comparing against, and at what price? | Use a current mid-tier frontier model's published rate, listed in `references/cost-inputs.md`, and name the model and the date in the assumptions |
| How much text goes in and comes back per request? | Measure it from the project with `references/request-payload.md` rather than asking again. Failing that, use the closest payload profile in `references/cost-inputs.md` and name which one |
| Does the app add context per request: retrieval, memory, or injected state? | Look for it in the project before assuming there is none. It is usually the largest line in the payload and the one most often left out |
| Is there a stable system prompt long enough to cache? | Measure it, then check it against the provider's minimum cacheable length. Below that it silently does not cache |
| Where will the model run: the user's device, or a server you pay for? | Ask this one directly rather than defaulting. It moves the answer by orders of magnitude and everyone knows which they intend |
| What did or will the training run cost? | Single-digit dollars for a parameter-efficient run on a small model, or $0 where a free or included allowance covers it |
| How long will this version of the model serve before you retrain or replace it? | 12 months, stated as an assumption |
| How many requests a month do you expect? | Compute the breakeven anyway and tell them what volume they would need to reach it |

**Every default has to be visible in the answer.** Say which numbers came from
the user and which came from a default, both when you present the result and in
the assumptions block of `COST-CROSSOVER.md`. A crossover built on defaults is a
starting point for a conversation. A crossover built on defaults and presented
as measured is the invented-inputs failure this skill exists to prevent.
`references/cost-inputs.md` covers how to get each one honestly: pulling
token counts from real request logs instead of guessing, finding a
provider's own current pricing page instead of remembering a number that has
since moved, recognising an on-device deployment where hosting is zero,
sizing a hosting instance for a given model and throughput, what is
defensible to assume when the hardware is already owned, what a training run
actually costs, how to choose the amortisation window, and how prompt caching
changes the effective input rate.

**`references/request-payload.md` is how to get input 1 properly.** It reads the
project's own code to find where each part of the prompt is assembled, counts
each part, and sorts them into the parts that repeat verbatim between requests
and the parts that do not. Run it whenever the payload is unknown or the app has
a retrieval, memory, or state-injection layer, because those layers usually
dominate the token count and are the most commonly omitted line in an estimate.
The same measurement answers questions for **scoping-a-custom-model** (whether
retrieval already exists), **evaluating-a-tuned-model** (the request shape an
eval has to reproduce), and the `shipping-a-model-in-*` skills (how much context
the model has to hold on the device). A crossover
built on invented inputs is worse than no crossover, because it looks like
evidence.

## Run the calculation

From this skill's own directory:

```bash
python3 scripts/crossover.py \
  --api-in <dollars per 1M input tokens> \
  --api-out <dollars per 1M output tokens> \
  --tokens-in <input tokens per request> \
  --tokens-out <output tokens per request> \
  --hosting <monthly hosting cost, 0 if already owned> \
  --training <one off training cost, 0 if none> \
  --months <amortisation window> \
  --requests <actual or expected monthly volume, optional>
```

Python 3.9 or newer, standard library only, no network access. The script
does not know any provider's price and never fetches one; every number above
is supplied by the caller, gathered per `references/cost-inputs.md`.

**If Python is unavailable, or the script errors,** compute by hand. The
whole calculation is four steps:

1. Cost per API call = (tokens in / 1,000,000 x price per 1M input tokens)
   + (tokens out / 1,000,000 x price per 1M output tokens).
2. Amortised training cost per month = training cost / months. This is the
   step most hand calculations skip, and skipping it is what tilts a
   comparison toward owning: a training run that never shows up in the
   monthly number makes owning look free the month after it finishes.
3. Fixed monthly cost if owned = monthly hosting + amortised training cost
   per month. This is the hosting floor: what owning costs in a month with
   zero requests, not zero.
4. Breakeven requests per month = fixed monthly cost if owned / cost per API
   call, rounded up. Below that volume the API costs less overall; above it,
   owning does.

If the script raises an error on bad input, for example a negative price or
an amortisation window of zero months, it says exactly which value is wrong
rather than a raw traceback, and the fix is almost always a mistyped flag.

## Write `COST-CROSSOVER.md` with the assumptions first

Write this file into the user's project root. **The assumptions block goes
at the top, before the answer.** A cost model whose inputs are buried below
the headline number is a sales tool wearing a spreadsheet's clothes; putting
them first is what makes the number something a reader can actually argue
with, agree with, or correct.

```markdown
# Cost crossover: <project name>

## Assumptions
Mark every line MEASURED, STATED (the user told you) or DEFAULT (this skill
supplied it). A reader who cannot tell which is which cannot judge the number.

- API: <provider, model>, $<x> / 1M input tokens, $<y> / 1M output tokens
  (MEASURED | STATED | DEFAULT. Source: <pricing page URL>, checked <date>.
  Re-check before relying on this; API pricing moves.)
- Tokens per request: <n> in, <n> out (MEASURED from real traffic | STATED |
  DEFAULT: <which payload profile>)
- Deployment target: on-device | rented infrastructure | already owned hardware
- Monthly hosting: $<n> (MEASURED | STATED | DEFAULT. $0 on-device, and this
  line should say so rather than being left blank)
- Training cost: $<n>, one off (MEASURED from the run | STATED | DEFAULT.
  Source: <how this was arrived at>)
- Amortisation window: <n> months, because <how long this version of the model
  is expected to serve before it is retrained or replaced>
- Current or expected volume: <n> requests per month (MEASURED | STATED |
  DEFAULT: none, breakeven reported without it)

## Breakeven
<n> API calls per month

## Cost at this project's volume
| Monthly volume | API cost | Owned cost |
|---|---|---|
| <stated volume> | $<n> | $<n> |

## Verdict
<the script's verdict text, verbatim>

## Costs this model does not include
<see "Count the costs people forget" below; list which ones apply here>
```

## Count the costs people forget

The fixed monthly figure the script produces, hosting plus amortised
training, is a floor, not the full picture. Before treating it as the whole
cost of owning, check whether these apply and add them in by hand:

- **Idle hosting time.** A dedicated instance billed by the month costs the
  same at 2am with no traffic as it does at peak. The monthly hosting figure
  already reflects this; it is why `owned_monthly` in the script does not
  fall to zero at zero requests.
- **The engineering hours to run inference infrastructure.** Someone has to
  deploy, monitor, and restart the serving stack. A rough hourly rate times
  expected hours a month is a real cost even when nobody bills a line item
  for it.
- **Retraining as data or requirements drift.** The one off training cost is
  rarely truly one off. If the task or the underlying data changes enough to
  need a second training run, that is a second amortised cost, not a
  rounding error.
- **On-device deployments trade hosting for distribution and device
  budget.** The per request cost is genuinely zero, and the costs that
  replace it are a larger application download, storage on the user's
  device, battery and thermal headroom during inference, and the engineering
  work to ship model updates through an app release rather than a server
  deploy. None of these belong in the arithmetic. All of them belong in the
  written record next to it.
- **Hardware that is "already owned" is not free.** It is close to free at
  the margin, the honest way to describe it, but power draw and the
  opportunity cost of not running something else on that hardware are real.
  Treating already owned hardware as exactly $0 is only defensible when it
  would otherwise sit fully idle.

## Read the crossover honestly

**Below the breakeven volume, the API is cheaper, plainly and without
qualification.** Say so directly in the verdict; do not soften it into
"potentially comparable" or bury it under caveats about future growth.
Most projects, measured honestly, sit below their own crossover point. That
is the expected outcome, not a disappointing one.

**Above it, owning is cheaper on the cost axis alone.** That does not
automatically make owning the right decision: it still has to clear the
quality bar (**evaluating-a-tuned-model**) and someone still has to run the
infrastructure. Cost is one input to that decision, not the whole decision.

**The two options differ in predictability as well as in total.** The owned
side is close to flat: hosting plus amortised training costs the same in a
quiet month as in a busy one, and on-device it stays flat at the amortised
training figure however much the app is used. The API side scales with every
API call, so the bill tracks traffic, including traffic nobody forecast. For
an application with steady, sustained volume that sits near or above the
crossover, that makes the owned side's monthly cost easier to forecast and
budget, which is a real advantage worth naming alongside the total.

Report it as predictability rather than as a fixed cost. Hosting tiers
change when throughput outgrows an instance, retraining recurs, and a
managed platform's usage still moves with the work it does. "More
predictable at steady volume" is defensible; "fixed" overstates it and is
easy for a reader to falsify. Below the crossover the API is both cheaper
and perfectly adequate, so predictability does not rescue an owned model
that loses on cost.

**If the projected volume is close to the breakeven line,** say that
plainly too, and flag which input the answer is most sensitive to (usually
tokens per request or the API's own price, both of which move). A crossover
that flips with a 10% change in one input is a genuinely uncertain result,
not a firm verdict either way, and reporting it as firm either way would be
its own kind of tilt. Given `--requests`, the script applies this rule
itself: when the two monthly figures land within 5% of each other it
reports a wash and names no winner at that volume, so the verdict text is
safe to paste verbatim near the line as well as far from it.

## Worked example

Illustrative inputs, not a claim about any specific project's actual
numbers:

- API: OpenAI's GPT-4o mini list pricing, $0.15 per 1M input tokens and
  $0.60 per 1M output tokens (source:
  `developers.openai.com/api/docs/pricing`, checked 2026-07-29. This is
  exactly the kind of number that moves; re-check it before using it for a
  real decision).
- A request that sends 500 input tokens and receives 300 output tokens,
  round numbers standing in for a modest classification or extraction task.
- Hosting: $180 a month, illustrating roughly 180 hours on a mid-tier rented
  GPU at approximately $1 an hour (source: a GPU rental marketplace's own
  on-demand pricing page, checked 2026-07-29; the actual rate and the hours
  a given model needs both vary by provider, GPU, and traffic pattern, see
  `references/cost-inputs.md`).
- Training: $5, standing in for a parameter-efficient fine-tune of a small
  model that finishes in about an hour on a single mid-tier GPU. A real
  number depends on the base model size and the GPU hours the run actually
  takes, and a managed platform may cover it inside an existing plan or free
  allowance; do not reuse this figure for a real decision.
- Amortised over 12 months.

Running `scripts/crossover.py` on these inputs gives a cost of $0.000255 per
API call, a fixed monthly cost of $180.42 if owned ($0.42 of which is the
amortised training cost), and a breakeven of 707,517 API calls a month.

At 400,000 API calls a month, the API costs $102.00 and owning costs $180.42:
**the API wins here, clearly**, because this volume sits well below the
breakeven line. At 1,000,000 API calls a month, the API costs $255.00 and
owning costs $180.42: owning wins, because this volume sits above it. Both
outcomes come from the same inputs at different volumes; neither is the
"right" answer to ship, the volume decides it.

Notice what dominates. Hosting is $180 of the $180.42, so the breakeven is
set almost entirely by the serving instance and barely at all by the
training run. That is the usual shape once the training figure is realistic,
and it is why the next example changes the answer so completely.

## Worked example, on-device

Same task, same API comparison, same token counts, same $5 training run.
The one change is that the model ships inside the application and runs on
the user's own hardware, so `--hosting 0`.

Cost per API call is unchanged at $0.000255. The fixed monthly cost of
owning falls to **$0.42**, all of it amortised training, and the breakeven
lands at **1,634 API calls a month**.

At 1,000 API calls a month the API still costs less, $0.25 against $0.42, so
even here the honest verdict below the line is the API. At 50,000 API calls
a month the API costs $12.75 and owning costs $0.42: **owning is cheaper by
96.7%**, and the gap keeps widening with volume, because the owned side
stays flat at $0.42 no matter how many API calls run through it.

That flatness is the whole point of the on-device case. There is no per
request inference charge to whoever ships the app, so the owned curve does
not climb. A hosted deployment crosses over in the high hundreds of
thousands of API calls a month; the same model shipped on-device crosses
over in the low thousands, roughly two and a half orders of magnitude
earlier. Anyone comparing an on-device deployment against a rented-GPU
breakeven figure is answering a question they did not ask.

The costs that replace hosting are real and belong in "costs this model does
not include": a larger application download, device storage, and battery and
thermal budget on the user's hardware. A per request bill is not among
them.

## Hand off to

- The model has not been shown to actually beat its base yet, only priced:
  **evaluating-a-tuned-model**
- The crossover favours owning and it needs to actually run somewhere:
  **shipping-a-model-in-a-react-native-app**,
  **shipping-a-model-in-an-ios-app**, **shipping-a-model-in-an-android-app**,
  or **shipping-a-model-in-a-flutter-app**
- The crossover clearly favours the API, or it is no longer obvious a custom
  model is worth building at all: **scoping-a-custom-model**

Files in this skill

  • SKILL.md17.8 KB
  • references/cost-inputs.md19.1 KB
  • references/request-payload.md5.9 KB
  • scripts/crossover.py10.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…