Skip to content
INFRO

What are you spending on Claude?

Anthropic's line-up spans a wide price range — the flagship costs an order of magnitude more per token than the small model — so on Claude workloads the model-selection question usually dominates the vendor question. Put your spend against each model to see where it actually concentrates.

Build your basket

Cost auditMonthly · USD

Your monthly AI spend

2 models
  • −25%
  • −30%

Enter what you spend today at each provider's direct price. Savings are computed per model from the rates published on this page — models we price the same as the provider show no saving, and are left in rather than quietly dropped to flatter the total.

Your estimate

Current spend

$4,800

per month

Estimated with INFRO

$3,560

per month

Potential savings

$1,240

26% blended across your basket

Over twelve months

$14,880

same basket, same rates

26% of current spend

Get a free cost auditOr request early access →

An estimate, not a quote. It prices your basket at the reference and INFRO rates published on this page — your real mix of prompt lengths, cached tokens, and image sizes will move it.

What moves this bill more than the vendor does.

A calculator prices a basket. These are the things that make a real invoice differ from one — worth knowing before you treat any estimate, including ours, as a number.

The tier gap is the biggest lever you have
Fable 5, Opus 5, Sonnet 5, and Haiku 4.5 are separated by large multiples per token. Before comparing vendors, it is worth knowing what share of your traffic is running on a frontier model because it was the default rather than because the task needs it — classification and extraction rarely do.
Thinking tokens are output tokens
Extended reasoning is billed at the output rate, and reasoning-heavy prompts can generate far more of it than the visible answer suggests. A workload whose output token count looks surprising is usually a reasoning budget, not a bug.
Prompt caching compounds on long system prompts
Agent and RAG workloads that resend a large stable prefix benefit most. Cached reads are billed well below the standard input rate, and INFRO bills the provider's cached rate through rather than charging you the uncached price — so the discount compounds instead of being absorbed.

Every Anthropic model we carry

Direct price beside ours, per unit. Rows priced at parity are left in rather than dropped.

ModelTaskDirectINFROSave
Claude Fable 5anthropic/claude-fable-5Frontier reasoning & agents$20$16 /1M tok20%
Claude Opus 5anthropic/claude-opus-5Coding & complex work$10$8.00 /1M tok20%
Claude Sonnet 5anthropic/claude-sonnet-5Production workhorse$6.00$4.50 /1M tok25%
Claude Haiku 4.5anthropic/claude-haiku-4.5Fast, low-cost chat$2.00$1.40 /1M tok30%

USD, illustrative launch rates reconciled on 2026-08-24. Live rates come from GET /v1/models once your key is active.

An estimate is a floor. Want the real number?

Send a usage export and we will price your actual traffic — with the arithmetic shown, and the caveats named.