induwara.lk
induwara.lkAI · Developer tools

LLM API Pricing Comparison

Enter your monthly input and output token volume and see the exact per-month API cost of 26 text models — GPT, Claude, Gemini, Grok, DeepSeek, Mistral and Llama — ranked cheapest-first, with a USD→LKR toggle. No signup, no ads.

By Induwara AshinsanaUpdated Jul 17, 2026
Compare LLM API prices26 models
Prices verified 2026-07-17

1,000,000 tokens/month — Prompt tokens you send per month (whole number).

250,000 tokens/month — Completion tokens the model returns per month.

Example workloads

0 = no caching. 100 = every input token cached.

Provider rates are in USD; LKR uses the rate on the right.

Rs

CBSL indicative default — edit to your bank's rate.

Tiers
Cheapest for your workload
Mistral Ministral 8B$0.125/mo
#ModelIn $/1MOut $/1MMonthly cost
1
Ministral 8BBest value
Mistral · Budget
$0.10$0.10$0.125
2
GPT-5 nano
OpenAI · Budget
$0.05$0.40$0.15
3
Mistral Small
Mistral · Budget
$0.10$0.30$0.175
4
Llama 4 Scout
Meta · Open-weight
$0.11$0.34$0.195
5
Gemini 2.5 Flash-Lite
Google · Budget
$0.10$0.40$0.20
6
DeepSeek V4 Flash
DeepSeek · Budget
$0.14$0.28$0.21
7
Grok 4 Fast
xAI · Budget
$0.20$0.50$0.325
8
Llama 4 Maverick
Meta · Open-weight
$0.20$0.60$0.35
9
Qwen3 235B
Alibaba · Open-weight
$0.20$0.60$0.35
10
DeepSeek V4
DeepSeek · Budget
$0.28$0.42$0.385
11
GPT-5.4 mini
OpenAI · Mid-tier
$0.25$2.00$0.75
12
Mistral Medium
Mistral · Mid-tier
$0.40$2.00$0.90
13
Gemini 2.5 Flash
Google · Mid-tier
$0.30$2.50$0.925
14
DeepSeek R2
DeepSeek · Mid-tier
$0.55$2.19$1.0975
15
Llama 3.3 70B
Meta · Open-weight
$0.88$0.88$1.10
16
Gemini 3 Flash
Google · Mid-tier
$0.50$3.00$1.25
17
Claude Haiku 4.5
Anthropic · Mid-tier
$1.00$5.00$2.25
18
Qwen3 Max
Alibaba · Frontier
$1.20$6.00$2.70
19
Mistral Large
Mistral · Mid-tier
$2.00$6.00$3.50
20
Command A
Cohere · Mid-tier
$2.50$10.00$5.00
21
Gemini 3.1 Pro
Google · Frontier
$2.00$12.00$5.00
22
GPT-5.4
OpenAI · Frontier
$2.50$15.00$6.25
23
Claude Sonnet 4.5
Anthropic · Frontier
$3.00$15.00$6.75
24
Grok 4.1
xAI · Frontier
$3.00$15.00$6.75
25
Claude Opus 4.8
Anthropic · Frontier
$5.00$25.00$11.25
26
GPT-5.5
OpenAI · Frontier
$5.00$30.00$12.50

Rates are each provider's published USD price per 1,000,000 tokens for standard synchronous text, last verified 2026-07-17. Figures exclude batch, fine-tuning, image, audio and cache-write charges. Always confirm against the provider's pricing page before committing — see sources below.

How it works

The calculator answers one question — which LLM API is cheapest for my workload— by applying each provider's published per-token rate to the exact volume you enter. Every model's price is stored as USD per 1,000,000 tokens for standard synchronous text, split into a separate input (prompt) rate and output (completion) rate, because providers bill the two differently and output is usually the larger cost.

For a model with input price pin, output price pout and cache-read price pcache, given monthly input tokens I, output tokens O and a cached input share s:

  1. Split input into cached and uncached tokens: uncached = I × (1 − s) and cached = I × s.
  2. Input cost = (uncached ÷ 1,000,000 × pin) + (cached ÷ 1,000,000 × pcache). Models with no published cache rate use pcache = pin (no discount).
  3. Output cost = O ÷ 1,000,000 × pout.
  4. Monthly cost (USD) = input cost + output cost. The table sorts ascending by this figure and flags the minimum as Best value.
  5. If you switch to LKR, every figure is multiplied by your editable USD→LKR rate (default Rs 305, the Central Bank of Sri Lanka indicative rate).

The result is cross-checked by an independent, algebraically-equivalent blended-rate formula that folds the cached and uncached input rates into one effective per-token price — effInput = pin × (1 − s) + pcache × s — and produces the same total to the cent. Prices are a static, verified table (last checked 2026-07-17); no live network calls are made, so the tool stays fast and never breaks on a failed price fetch. Because model prices change often, always confirm the current rate on the provider's page before committing — every source is linked below.

Worked examples

Support chatbot — 2,000,000 input + 500,000 output tokens/mo, 0% cached

  1. GPT-5.5: (2 × $5) + (0.5 × $30) = $10.00 + $15.00 = $25.00/mo
  2. Claude Sonnet: (2 × $3) + (0.5 × $15) = $6.00 + $7.50 = $13.50/mo
  3. Gemini 2.5 Flash: (2 × $0.30) + (0.5 × $2.50) = $0.60 + $1.25 = $1.85/mo
  4. DeepSeek V4 Flash: (2 × $0.14) + (0.5 × $0.28) = $0.28 + $0.14 = $0.42/mo
  5. Cheapest → dearest: DeepSeek $0.42 · Gemini $1.85 · Sonnet $13.50 · GPT-5.5 $25.00
  6. In LKR @305: DeepSeek ≈ Rs 128.10/mo · GPT-5.5 ≈ Rs 7,625.00/mo

Coding assistant — 5,000,000 input + 3,000,000 output tokens/mo, 0% cached

  1. Gemini 3.1 Pro: (5 × $2) + (3 × $12) = $10.00 + $36.00 = $46.00/mo
  2. GPT-5.4: (5 × $2.50) + (3 × $15) = $12.50 + $45.00 = $57.50/mo
  3. Claude Opus 4.8: (5 × $5) + (3 × $25) = $25.00 + $75.00 = $100.00/mo
  4. Ranked: Gemini 3.1 Pro $46.00 · GPT-5.4 $57.50 · Claude Opus 4.8 $100.00
  5. Cheapest verdict: Gemini 3.1 Pro ≈ Rs 14,030/mo at 305

Cached edge case — 1,000,000 input, 0 output, 100% cached (DeepSeek V4 Flash)

  1. Cache-read rate: $0.014 per 1M (vs $0.14 standard input)
  2. uncached = 1,000,000 × (1 − 1.00) = 0 tokens
  3. cached = 1,000,000 × 1.00 = 1,000,000 tokens
  4. Input cost: (0 × $0.14) + (1 × $0.014) = $0.014/mo
  5. Output cost: 0 → Monthly cost = $0.014/mo (full cache discount applied)

Frequently asked questions

Sources & references

Every rate in the price table was cross-checked against these official pricing pages on 2026-07-17. Provider-hosted open-weight models (Llama, Qwen) are priced by the serving provider, not the model author. LLM prices change frequently — confirm the current figure on the provider's page before you commit to a model.

Related tools

Rate this tool
Be the first to rate

Comments & feedback

Spotted a bug or want an improvement? Tell us — our team reviews every comment, and good ideas get built. Comments are public and anonymous.

Spotted a stale price, missing model, or edge case?

Email me at [email protected] — most fixes ship within 24 hours.