AI Model Cost Calculator

What will AI actually cost your business?

Compare live API pricing across Claude, GPT, Gemini, and DeepSeek. Adjust the inputs — the numbers update as you go.

Your workload

Requests / month50K/mo
10010K1M10M

Average tokens per request. Defaults are typical for the selected use case.

Total tokens / month38M

Estimated monthly cost

Sorted cheapest first · USD
  • GPT-5.6 LunabudgetCHEAPEST
    $20.00
    $0.20 in · $1.20 out / 1M$0.00040 / request
  • DeepSeek V4 Probudget
    $21.75
    $0.43 in · $0.87 out / 1M$0.00044 / request
  • Claude Haiku 4.5budget
    $87.50
    $1.00 in · $5.00 out / 1M$0.00175 / request
  • Gemini 3.6 Flashmid
    $131
    $1.50 in · $7.50 out / 1M$0.00263 / request
  • Claude Sonnet 5premium
    $175
    $2.00 in · $10.00 out / 1M$0.00350 / request
  • Gemini 3.1 Propremium
    $200
    $2.00 in · $12.00 out / 1M$0.00400 / request
  • GPT-5.6 Solpremium
    $500
    $5.00 in · $30.00 out / 1M$0.010 / request
  • TYPICAL ROUTED SETUP80% GPT-5.6 Luna · 20%
    $51.00
    Cheapest budget tier + your chosen premium model, blended$0.00102 / request
P
Our take

The flagship models here cost roughly 25.0× the efficient tier for this workload. Most teams start on an efficient model, measure quality on real traffic, and escalate only the requests that need it — a routing pattern that usually saves 40–70%.

Want a real estimate for your use case?

These are back-of-envelope numbers. We build production AI systems — routing, caching, evals, and the infrastructure around them — and can model your actual workload.

Frequently asked questions

How accurate are these cost estimates?

They're planning-level estimates based on each provider's own published per-token list price (see the sources linked in the footer). Actual bills depend on your real token counts, any volume discounts you've negotiated, and provider pricing changes — use this to budget, not as an invoice.

What is prompt caching, and how does it change the price?

Several providers let you reuse parts of a prompt — a system prompt or long context, for example — across requests at a steep discount versus reprocessing it every time. Turn on the caching toggle to see the effect: it applies each model's own cached-input rate to whatever share of input tokens you set with the caching slider.

What does the 'Typical routed setup' row show?

It models a common cost-optimization pattern: sending most requests to the cheapest budget-tier model in the comparison, and only routing the harder share to a premium model you pick. Production systems often route this way to cut spend without giving up quality on the requests that actually need a stronger model.

How often is the pricing updated?

The date in the footer shows when these prices were last checked directly against each provider's own pricing page. AI API pricing changes often — if a number looks off, follow the source link next to it.

Which providers and models are included?

The comparison spans premium, mid, and budget-tier models from Anthropic, OpenAI, Google, and DeepSeek — see the exact current lineup, tiers, and prices in the table above.

Is this calculator free to use?

Yes — it's a free tool from Pacifiq Labs. If you want help modeling your actual production workload, routing strategy, or caching setup, use the contact link above.