LLM API Cost Calculator
Enter your monthly request volume and average token counts to compare what 14 frontier models cost at official list prices — GPT-5.6, Claude, Gemini, Grok, DeepSeek, Kimi and more. Prices verified 2026-09-03, each with a dated source.
Presets:
| Model | Per request | Per month | Relative |
|---|---|---|---|
| DeepSeek V4 FlashDeepSeekcheapest | $0.0005 | $15.96 | |
| GPT-5.6 LunaOpenAI | $0.0011 | $32.40 | |
| Gemini 3.5 Flash-LiteGoogle | $0.0019 | $57.00 | |
| Gemini 3.8 FlashGoogle | $0.0037 | $112.50 | |
| Meta Muse Spark 1.1Meta | $0.0054 | $163.50 | |
| Gemini 3.6 FlashGoogle | $0.0075 | $225.00 | |
| Grok 4.5xAI | $0.0084 | $252.00 | |
| GPT-5.6 TerraOpenAI | $0.0108 | $324.00 | |
| Claude Sonnet 5Anthropic | $0.0150 | $450.00 | |
| Kimi K3Moonshot AI | $0.0150 | $450.00 | |
| Grok 4.5 FastxAI | $0.0192 | $576.00 | |
| Claude Opus 5Anthropic | $0.0250 | $750.00 | |
| GPT-5.6 SolOpenAI | $0.0270 | $810.00 | |
| Claude Fable 5Anthropic | $0.0500 | $1,500 |
List prices, cache misses, no batch discounts. Models differ in how many tokens they spend to finish the same task, so blended real-world costs can rank differently.
The list prices behind the math
USD per million tokens, cache-miss rates. Every row links to the CodingSalt coverage that documents it.
| Model | Input $/1M | Output $/1M | Verified | Notes |
|---|---|---|---|---|
| DeepSeek V4 Flash (DeepSeek) | $0.14 | $0.28 | 2026-08-02 | Cache hit input $0.0028; peak-hours 2x multiplier documented but not yet in effect |
| GPT-5.6 Luna (OpenAI) | $0.20 | $1.20 | 2026-08-01 | 80% cut on July 30, 2026 |
| Gemini 3.5 Flash-Lite (Google) | $0.30 | $2.50 | 2026-07-28 | — |
| Gemini 3.8 Flash (Google) | $0.75 | $3.75 | 2026-09-03 | Introductory — doubles to $1.50/$7.50 on Jan 1, 2027 |
| Meta Muse Spark 1.1 (Meta) | $1.25 | $4.25 | 2026-07-28 | Cached input $0.15 |
| Gemini 3.6 Flash (Google) | $1.50 | $7.50 | 2026-07-28 | Uses ~17% fewer output tokens per task than 3.5 Flash (vendor-reported) |
| Grok 4.5 (xAI) | $2.00 | $6.00 | 2026-07-28 | — |
| GPT-5.6 Terra (OpenAI) | $2.00 | $12.00 | 2026-08-01 | 20% cut on July 30, 2026 |
| Claude Sonnet 5 (Anthropic) | $3.00 | $15.00 | 2026-09-01 | Standard rate since Sept 1, 2026 (intro was $2/$10) |
| Kimi K3 (Moonshot AI) | $3.00 | $15.00 | 2026-07-28 | Cached input $0.30; open weights |
| Grok 4.5 Fast (xAI) | $4.00 | $18.00 | 2026-07-28 | Lower latency, higher per-token cost |
| Claude Opus 5 (Anthropic) | $5.00 | $25.00 | 2026-07-25 | Fast mode bills 2x |
| GPT-5.6 Sol (OpenAI) | $5.00 | $30.00 | 2026-08-01 | Unchanged in the July 30 cut; gained a Fast mode |
| Claude Fable 5 (Anthropic) | $10.00 | $50.00 | 2026-07-28 | Cache reads ~$1/M |
Frequently asked questions
- How is the monthly LLM API cost calculated?
- Cost per request = (average input tokens × input price + average output tokens × output price) ÷ 1,000,000, multiplied by your monthly request count. Prices are official vendor list rates per million tokens, assuming cache misses and no batch discounts.
- Which LLM API is cheapest in late 2026?
- On list price, DeepSeek V4 Flash is the cheapest at $0.14 per million input tokens and $0.28 per million output tokens, followed by GPT-5.6 Luna at $0.20/$1.20 after OpenAI's July 30, 2026 price cut. Real costs depend on how many tokens each model spends to finish your task.
- Where do the prices in this calculator come from?
- Every price is the vendor's official list rate, verified on the date shown in the table and documented in a linked CodingSalt article. The table was last updated on 2026-09-03.
- Do these prices include prompt caching or batch discounts?
- No — the calculator assumes cache-miss input pricing and synchronous requests. Caching and batch APIs can cut costs substantially, but the rules differ per vendor: cache reads, cache writes and batch rates are each priced separately.
Prices move — we publish the changes the day they happen. Follow the RSS feed or read the full pricing comparison.