codingsalt

LLM API Cost Calculator

Enter your monthly request volume and average token counts to compare what 14 frontier models cost at official list prices — GPT-5.6, Claude, Gemini, Grok, DeepSeek, Kimi and more. Prices verified 2026-09-03, each with a dated source.

Presets:
ModelPer requestPer monthRelative
DeepSeek V4 FlashDeepSeekcheapest$0.0005$15.96
GPT-5.6 LunaOpenAI$0.0011$32.40
Gemini 3.5 Flash-LiteGoogle$0.0019$57.00
Gemini 3.8 FlashGoogle$0.0037$112.50
Meta Muse Spark 1.1Meta$0.0054$163.50
Gemini 3.6 FlashGoogle$0.0075$225.00
Grok 4.5xAI$0.0084$252.00
GPT-5.6 TerraOpenAI$0.0108$324.00
Claude Sonnet 5Anthropic$0.0150$450.00
Kimi K3Moonshot AI$0.0150$450.00
Grok 4.5 FastxAI$0.0192$576.00
Claude Opus 5Anthropic$0.0250$750.00
GPT-5.6 SolOpenAI$0.0270$810.00
Claude Fable 5Anthropic$0.0500$1,500

List prices, cache misses, no batch discounts. Models differ in how many tokens they spend to finish the same task, so blended real-world costs can rank differently.

The list prices behind the math

USD per million tokens, cache-miss rates. Every row links to the CodingSalt coverage that documents it.

ModelInput $/1MOutput $/1MVerifiedNotes
DeepSeek V4 Flash (DeepSeek)$0.14$0.282026-08-02Cache hit input $0.0028; peak-hours 2x multiplier documented but not yet in effect
GPT-5.6 Luna (OpenAI)$0.20$1.202026-08-0180% cut on July 30, 2026
Gemini 3.5 Flash-Lite (Google)$0.30$2.502026-07-28
Gemini 3.8 Flash (Google)$0.75$3.752026-09-03Introductory — doubles to $1.50/$7.50 on Jan 1, 2027
Meta Muse Spark 1.1 (Meta)$1.25$4.252026-07-28Cached input $0.15
Gemini 3.6 Flash (Google)$1.50$7.502026-07-28Uses ~17% fewer output tokens per task than 3.5 Flash (vendor-reported)
Grok 4.5 (xAI)$2.00$6.002026-07-28
GPT-5.6 Terra (OpenAI)$2.00$12.002026-08-0120% cut on July 30, 2026
Claude Sonnet 5 (Anthropic)$3.00$15.002026-09-01Standard rate since Sept 1, 2026 (intro was $2/$10)
Kimi K3 (Moonshot AI)$3.00$15.002026-07-28Cached input $0.30; open weights
Grok 4.5 Fast (xAI)$4.00$18.002026-07-28Lower latency, higher per-token cost
Claude Opus 5 (Anthropic)$5.00$25.002026-07-25Fast mode bills 2x
GPT-5.6 Sol (OpenAI)$5.00$30.002026-08-01Unchanged in the July 30 cut; gained a Fast mode
Claude Fable 5 (Anthropic)$10.00$50.002026-07-28Cache reads ~$1/M

Frequently asked questions

How is the monthly LLM API cost calculated?
Cost per request = (average input tokens × input price + average output tokens × output price) ÷ 1,000,000, multiplied by your monthly request count. Prices are official vendor list rates per million tokens, assuming cache misses and no batch discounts.
Which LLM API is cheapest in late 2026?
On list price, DeepSeek V4 Flash is the cheapest at $0.14 per million input tokens and $0.28 per million output tokens, followed by GPT-5.6 Luna at $0.20/$1.20 after OpenAI's July 30, 2026 price cut. Real costs depend on how many tokens each model spends to finish your task.
Where do the prices in this calculator come from?
Every price is the vendor's official list rate, verified on the date shown in the table and documented in a linked CodingSalt article. The table was last updated on 2026-09-03.
Do these prices include prompt caching or batch discounts?
No — the calculator assumes cache-miss input pricing and synchronous requests. Caching and batch APIs can cut costs substantially, but the rules differ per vendor: cache reads, cache writes and batch rates are each priced separately.

Prices move — we publish the changes the day they happen. Follow the RSS feed or read the full pricing comparison.