whichAI
Data reviewed 2026-07-17

AI API Pricing Calculator 2026

Estimate monthly API spend by workload: coding agents, research, RAG support bots, realtime apps, and low-cost automation across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and more.

← Subscription plansAPI vs subscription optimizer β†’
Cheapest overall
DeepSeek V4 Flash
$0.14 / $0.28 per MTok
Best value general
Gemini 2.5 Flash
$0.30 / $2.50 per MTok
Best for coding
Mistral Codestral
$0.30 / $0.90 per MTok
Most powerful
Claude Opus 4.8
$5.00 / $25.00 per MTok

πŸ“‹ Not sure about token counts? Pick one or more use cases:

🌍 Your country:
No Sales Tax (varies)
ℹ️ No federal VAT. State sales tax on digital services varies (0–10%). Most AI providers do n…

⚠️ Tax rates and exchange rates are approximate and may vary. Actual charges depend on your payment method, billing country, and current exchange rates. Prices shown are estimates only β€” verify with your provider before subscribing.

πŸ’° Monthly Cost Calculator

~750 words β‰ˆ 1K tokens
~375 words β‰ˆ 500 tokens
= 900/month
Your usage estimate
0.90Minput tokens/mo
0.45Moutput tokens/mo
900requests/mo
⚠️ This estimate excludes:

Taxes, tool/function call costs, web search fees, image/audio input costs, cached write costs, long-context surcharges, and minimum request unit rounding.

Prompt caching can reduce input costs up to 90% for repeated prompts. Actual costs may be 20–50% higher depending on tool/search/image usage.

What this usage means

Quick read before you scan every model row. Costs assume standard pricing only. Discounts (batch, cached input) not applied unless toggled above.

Full API vs subscription guide β†’
API likely cheaper

Your cheapest API option is about $0.08/mo, below a typical $20 subscription.

Moderate workload

900 requests/month using 1.35M total tokens.

Best price in view

Llama 3.1 8B Instant (Groq) at $0.08/mo.

Workload:

Show every token-priced model in the calculator.

Recommended APIs for this workload

These cards use the selected workload and provider filters, then rank models by monthly cost, quality fit, and production fit.

Read the API pricing guide
Cheapest
Lowest bill
Llama 3.1 8B Instant
Groq Β· $0.08/mo

Start here when the workload is repetitive, simple, or high volume.

Quality
Best output quality
Gemini 3.1 Pro Preview
Google (Gemini API) Β· $7/mo

Use this when accuracy, reasoning, writing quality, or hard coding tasks matter more than raw price.

Production
Best app backend
Gemini 2.5 Flash-Lite
Google (Gemini API) Β· $0.27/mo

A practical candidate for RAG, routing, support chat, agents, or cached/batch workloads.

ModelProviderUse case fitEst. Monthly Costvs $20 SubInput / 1MOutput / 1MCached / 1MContextBest for
Llama 3.1 8B Instant
20M input tokens per dollar. Fastest inference available for this model size.
Groq
RealtimeLowest cost
$0.08
/month
No sub
$0.05
/1M
$0.08
/1M
β€”128KUltra-fast simple tasks at 500+ tok/s β€” latency-sensitive classification and routing
Command R7B
Cheapest flagship-quality API as of June 2026. 4x cheaper than GPT-5.4 Nano on i…
Cohere
RAGLowest costRealtime
$0.10
/month
No sub
$0.0375
/1M
$0.15
/1M
β€”128KHigh-volume RAG, classification, routing β€” cheapest first-party production API
Llama 4 Scout (DeepInfra)
DeepInfra hosted meta-llama/Llama-4-Scout-17B-16E-Instruct.
Meta Llama (hosted via DeepInfra)
Lowest costRealtimeResearch
$0.23
/month
No sub
$0.1
/1M
$0.3
/1M
β€”327,680Efficient MoE model β€” fast and cheap for classification, chat, light coding tasks
Llama 3.3 70B Turbo (DeepInfra)
Current DeepInfra Turbo endpoint; replaces the deprecated non-Turbo endpoint pre…
Meta Llama (hosted via DeepInfra)
Lowest costRealtimeCoding
$0.23
/month
No sub
$0.1
/1M
$0.32
/1M
β€”128KCheapest hosted 70B model β€” best cost/quality for general tasks
DeepSeek V4 Flash
One of the cheapest frontier-class APIs available. Strong on coding benchmarks.
DeepSeek
Lowest costCodingRealtime
$0.25
/month
No sub
$0.14
/1M
$0.28
/1M
$0.0028
cached
1MBudget coding tasks, high-volume text generation, agentic pipelines where cost matters most
Gemini 2.5 Flash-Lite
Cheapest model in Google lineup. Flat pricing on all context lengths.
Google (Gemini API)
Lowest costRealtimeRAG
$0.27
/month
API cheaper
save $20/mo
$0.1
/1M
$0.4
/1M
$0.01
cached
1MUltra-high-volume simple tasks: classification, routing, lightweight extraction
Mistral Small 4
Strong price-performance for EU-based applications.
Mistral AI
Lowest costRealtime
$0.41
/month
API cheaper
save $20/mo
$0.15
/1M
$0.6
/1M
β€”256KEU deployments, budget general tasks β€” French company with EU data residency
Command R
Native grounding and citation generation. Strong for retrieval-heavy workloads.
Cohere
RAGLowest costResearch
$0.41
/month
No sub
$0.15
/1M
$0.6
/1M
β€”128KRAG pipelines, cost-efficient enterprise chat with native grounding and tool use
GPT-OSS 120B
Current GroqCloud replacement for the retired Kimi K2 endpoint.
Groq
RealtimeLowest cost
$0.41
/month
No sub
$0.15
/1M
$0.6
/1M
β€”128KFast open-weight reasoning and agentic workloads on GroqCloud
Llama 4 Maverick FP8 (DeepInfra)
DeepInfra hosted meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8.
Meta Llama (hosted via DeepInfra)
CodingLowest costResearch
$0.54
/month
No sub
$0.2
/1M
$0.8
/1M
β€”1MBest open-weight model for coding and reasoning β€” strong competitor to GPT-5.4
Codestral
Purpose-built coding model. Very competitive pricing vs OpenAI/Anthropic for cod…
Mistral AI
CodingLowest cost
$0.68
/month
API cheaper
save $19/mo
$0.3
/1M
$0.9
/1M
β€”128KCode completion, FIM (fill-in-the-middle), IDE autocomplete, coding-specific tasks β€” purpose-built for code
GPT-5.4 Nano
Cheapest OpenAI model for production. Official short-context standard pricing.
OpenAI
Lowest costRealtimeRAG
$0.74
/month
API cheaper
save $19/mo
$0.2
/1M
$1.25
/1M
$0.02
cached
400KUltra-high-volume simple tasks: classification, intent detection, short-form extraction
DeepSeek V4 Pro
Current official price from DeepSeek API docs. Legacy deepseek-chat and deepseek…
DeepSeek
CodingLowest cost
$0.78
/month
No sub
$0.435
/1M
$0.87
/1M
$0.003625
cached
1MComplex coding and reasoning at significantly lower cost than OpenAI/Anthropic equivalents
Llama 3.3 70B Versatile
~250 tok/s. Undercuts GPT-4o mini on output cost while offering flagship-class c…
Groq
RealtimeLowest costCoding
$0.89
/month
No sub
$0.59
/1M
$0.79
/1M
β€”128KFlagship-quality at low cost β€” coding, QA, summarization on Groq LPU infrastructure
Gemini 3.1 Flash-Lite
Current cost-efficient Gemini 3 model. Batch/Flex standard text pricing is $0.12…
Google (Gemini API)
Lowest costRealtimeRAG
$0.90
/month
API cheaper
save $19/mo
$0.25
/1M
$1.5
/1M
$0.025
cached
1MHigh-volume agentic tasks, translation, lightweight data processing
Mistral Large 3
Flagship Mistral model. Good alternative to GPT-4o for EU deployments.
Mistral AI
QualityResearchLowest cost
$1
/month
API cheaper
save $19/mo
$0.5
/1M
$1.5
/1M
β€”256KComplex reasoning, EU compliance-sensitive tasks, multilingual European deployments
Sonar
Search included at no extra charge. ~$5-12 per 1,000 requests additional based o…
Perplexity (Sonar API)
ResearchRAG
$1
/month
No sub
$1
/1M
$1
/1M
β€”127KWeb-grounded Q&A with citations β€” cheapest model with built-in real-time search
Gemini 2.5 Flash
Best price-performance in current Gemini lineup. Free tier available.
Google (Gemini API)
RealtimeRAGLowest cost
$1
/month
API cheaper
save $19/mo
$0.3
/1M
$2.5
/1M
$0.03
cached
1MHigh-volume general tasks: summarization, extraction, coding assistance, rapid prototyping
grok-build-0.1
Early-access coding model listed on xAI API pricing.
xAI (Grok API)
CodingRealtime
$2
/month
API cheaper
save $18/mo
$1
/1M
$2
/1M
$0.2
cached
256KAgentic coding and build tasks on xAI infrastructure
Kimi K2.5
Official Kimi API price: cache hit $0.10, cache miss $0.60, output $3.00 per MTo…
Moonshot AI (Kimi)
CodingLowest costResearch
$2
/month
API cheaper
save $18/mo
$0.6
/1M
$3
/1M
$0.1
cached
256000Budget coding and multimodal tasks β€” cheapest vision-capable model in this comparison
Grok 4.3
Current xAI flagship text model.
xAI (Grok API)
QualityResearchRealtime
$2
/month
API cheaper
save $18/mo
$1.25
/1M
$2.5
/1M
$0.2
cached
1MGeneral reasoning, tool calling, long-context applications
Kimi K2.6
Official Kimi API price: cache hit $0.16, cache miss $0.95, output $4.00 per MTo…
Moonshot AI (Kimi)
CodingResearch
$3
/month
API cheaper
save $17/mo
$0.95
/1M
$4
/1M
$0.16
cached
256000Coding and agentic tasks at 4Γ— lower cost than Claude Sonnet β€” strong coding benchmark scores
GPT-5.4 Mini
Mid-range GPT-5.4 family model. Official short-context standard pricing.
OpenAI
CodingResearch
$3
/month
API cheaper
save $17/mo
$0.75
/1M
$4.5
/1M
$0.075
cached
400KLower-cost coding, long-form summarization, and medium-complexity agent steps
Claude Haiku 4.5
Cheapest current-gen Claude. Replaces deprecated Haiku 3 ($0.25/$1.25).
Anthropic
RealtimeRAG
$3
/month
API cheaper
save $17/mo
$1
/1M
$5
/1M
$0.1
cached
200KClassification, routing, extraction, summarization, high-volume workloads
GPT-5.6 Luna
OpenAI
RealtimeCoding
$4
/month
API cheaper
save $16/mo
$1
/1M
$6
/1M
$0.1
cached
1.05MFastest and cheapest GPT-5.6 β€” good for high-volume, latency-sensitive tasks
Moonshot V1 (128K)
Legacy model. Official price is $2 input/$5 output per MTok; platform sunset exp…
Moonshot AI (Kimi)
Research
$4
/month
API cheaper
save $16/mo
$2
/1M
$5
/1M
β€”128000Legacy model β€” superseded by K2.5 and K2.6. Listed for reference only.
Gemini 2.5 Pro
Best value for complex tasks in Gemini lineup. 2x surcharge beyond 200K.
Google (Gemini API)
ResearchCodingQuality
$6
/month
API cheaper
save $14/mo
$1.25
/1M
$10
/1M
$0.125
cached
1MComplex reasoning, coding, long-document analysis β€” strong at 1M context tasks
Claude Sonnet 5
Introductory pricing through 2026-08-31; standard pricing is $3/$15 per MTok the…
Anthropic
CodingResearch
$6
/month
API cheaper
save $14/mo
$2
/1M
$10
/1M
$0.2
cached
1MHigh-performance coding, agents, and general production workloads during introductory pricing
Command R+
Same price as GPT-5.4 on input, 33% cheaper on output. Purpose-built for RAG bea…
Cohere
RAGResearchQuality
$7
/month
No sub
$2.5
/1M
$10
/1M
β€”128KComplex agentic RAG, enterprise document intelligence, multi-step grounded reasoning
Gemini 3.1 Pro Preview
Preview model. Standard price for prompts up to 200K tokens; prompts above 200K …
Google (Gemini API)
QualityResearchCoding
$7
/month
API cheaper
save $13/mo
$2
/1M
$12
/1M
$0.2
cached
1MHigh-quality multimodal reasoning, agentic workflows, vibe-coding
GPT-5.4
Strong balance of capability and cost. Official short-context standard pricing; …
OpenAI
CodingResearch
$9
/month
API cheaper
save $11/mo
$2.5
/1M
$15
/1M
$0.25
cached
1.05MCoding (debugging, code gen, refactoring), content generation, analysis β€” solid all-rounder
GPT-5.6 Terra
OpenAI
CodingResearch
$9
/month
API cheaper
save $11/mo
$2.5
/1M
$15
/1M
$0.25
cached
1.05MBalanced quality and cost β€” good daily driver for most tasks
Claude Sonnet 4.6
Legacy/current comparison point. Sonnet 5 is cheaper during introductory pricing…
Anthropic
CodingResearchRAG
$9
/month
API cheaper
save $11/mo
$3
/1M
$15
/1M
$0.3
cached
1MGeneral coding, analysis, writing, RAG pipelines, agentic tasks
Sonar Pro
$6-14 per 1,000 requests additional. Use when citation accuracy and source depth…
Perplexity (Sonar API)
ResearchRAG
$9
/month
No sub
$3
/1M
$15
/1M
β€”127KDeep research with multi-source citations β€” complex questions requiring current web data
Claude Opus 4.8
Current Opus model. Fast Mode is 2x standard pricing.
Anthropic
QualityCodingResearch
$16
/month
API cheaper
save $4/mo
$5
/1M
$25
/1M
$0.5
cached
1MComplex reasoning, agentic workflows, code review, long-context analysis
GPT-5.5
Flagship model. Official short-context standard pricing; long-context and Priori…
OpenAI
QualityCodingResearch
$18
/month
Break-even
$5
/1M
$30
/1M
$0.5
cached
1.05MComplex multi-step workflows, frontier writing quality, tool-use intensive agents
GPT-5.6 Sol
OpenAI
QualityCodingResearch
$18
/month
Break-even
$5
/1M
$30
/1M
$0.5
cached
1.05MFlagship β€” highest quality for complex reasoning, coding, and research
Claude Fable 5
Top-tier Claude model listed in official pricing. More expensive than Opus 4.8.
Anthropic
QualityResearchCoding
$32
/month
Sub cheaper
$12/mo over
$10
/1M
$50
/1M
$1
cached
1MFrontier reasoning, long-context agentic workflows, high-stakes analysis
Claude Mythos 5
Limited availability. Official Claude Platform pricing is $10/$50 per MTok.
Anthropic
QualityResearch
$32
/month
Sub cheaper
$12/mo over
$10
/1M
$50
/1M
$1
cached
1MAdvanced reasoning and long-context work where limited availability is acceptable
GPT-5.5 Pro
Premium model. Official pricing is $30/$180 per MTok; no cached-input discount.
OpenAI
QualityResearch
$108
/month
Sub cheaper
$88/mo over
$30
/1M
$180
/1M
β€”1.05MMost demanding frontier tasks where quality matters more than cost

Prices shown are standard per-token rates. Batch, cached, priority, and long-context rates vary. Verify current pricing at each provider's official API documentation.

Provider notes & discounts

verified 2026-07-17

Output tokens are typically priced 5x input on Claude models. Batch API: 50% off. Prompt caching read is 90% off cached input. Sonnet 5 is listed with introductory $2/$10 pricing through 2026-08-31; standard pricing is $3/$15 thereafter.

βœ“ batch: 50% off (24hr async processing)βœ“ prompt cache: 90% off cached input tokensβœ“ long context surcharge: 2x input price beyond 200K tokens on some models
verified 2026-07-17

Standard short-context pricing shown for flagship models. OpenAI lists separate long-context and Priority pricing for GPT-5.5/GPT-5.4 families. Batch and Flex pricing may differ by endpoint; verify current rates before production use.

βœ“ batch: 50% off (24hr async)βœ“ prompt cache: 90% off cached input (GPT-5.5 and 5.4 families)βœ“ priority: Premium pricing for guaranteed faster processing; see official table

Google (Gemini API)

Official pricing β†—
verified 2026-07-17

Free tier available on many Gemini models. Batch and Flex are often discounted. Context caching is priced separately by model. Gemini 3 pricing includes 5,000 free grounded prompts per month shared across Gemini 3 models.

βœ“ batch: 50% off (24hr async)βœ“ context cache: 90% off cached inputβœ“ free tier: Free tier available on many Gemini modelsβœ“ long context surcharge: 2x input/output beyond 200K tokens on Gemini 2.5 Pro standard pricing
verified 2026-07-17

Current xAI API text pricing is focused on Grok 4.3 and grok-build-0.1. Image, video, and voice APIs use per-image, per-second, or per-hour pricing.

βœ“ new user credits: $25 freeβœ“ data sharing: $150/month additional via data sharing program
verified 2026-07-17

DeepSeek V4 Flash and V4 Pro both support 1M context and 384K max output. Privacy considerations: Chinese company. Web chat is free.

verified 2026-07-17

EU-based provider, strong GDPR compliance. Open-source models available. Good for EU privacy-conscious deployments.

verified 2026-07-17

Purpose-built for enterprise RAG (Retrieval-Augmented Generation). Command R7B is cheapest first-party production API in 2026. Complete RAG stack: Embed v3 ($0.10/M), Rerank 3.5 ($2/1K queries), Command generation. Free trial key available.

βœ“ free trial: Rate-limited trial key, no credit card requiredβœ“ enterprise: Custom pricing on AWS, Azure, Oracle, self-hosted
verified 2026-07-17

GroqCloud hosted pricing. Llama 3.1 8B and Llama 3.3 70B are scheduled for developer-tier shutdown on 2026-08-16; GPT-OSS 120B is the recommended current replacement.

βœ“ free tier: All models free with rate limitsβœ“ batch: 50% offβœ“ prompt cache: 50% off cached tokens

Perplexity (Sonar API)

Official pricing β†—
verified 2026-07-17

Sonar prices include token charges plus a separate per-request search-context fee. Low/medium/high context fees differ by model; token-only estimates understate total cost.

βœ“ vs gpt search: Typically 3-5x cheaper than GPT-5.4 + web search for research workloads

Meta Llama (hosted via DeepInfra)

Official pricing β†—
verified 2026-07-17

Meta does not sell a direct Llama API. These are current DeepInfra hosted endpoints and prices, verified from DeepInfra's first-party model-list API.

βœ“ self hosted: Free if self-hosted β€” only pay for GPU infrastructureβœ“ deepinfra: Current listed endpoints: Scout $0.10/$0.30, Maverick FP8 $0.20/$0.80, Llama 3.3 70B Turbo $0.10/$0.32

FLUX API (Black Forest Labs)

Official pricing β†—
verified 2026-07-17

Per-image pricing, not per-token. 1 credit = $0.01. Prices vary by model, operation, and output resolution.

βœ“ aggregators: FAL.ai and Replicate often 30-50% cheaper than BFL direct for older versionsβœ“ self hosted: FLUX.2 [klein] 4B free with ~13GB VRAM

Moonshot AI (Kimi)

Official pricing β†—
verified 2026-07-17

Prices exclude applicable taxes. Kimi K2.6 and K2.5 prices were extracted from the official page source because the rendered table parser omitted the numeric cells. Moonshot V1 platform sunset is expected on 2026-08-31.

API vs subscription: which is cheaper for you?

At ~2,000–2,200 interactions/month, Claude Sonnet API and Claude Pro subscription cost roughly the same. Below that, API wins. Above it, the flat $20/month Pro subscription is cheaper. Read our full breakdown:

API vs Subscription: When does pay-per-token save money? β†’

All 12 providers and 42 listed models rechecked against first-party pricing pages on 2026-07-17. Corrected OpenAI, Google, Anthropic, Mistral, Meta-hosted, Groq, and FLUX records; model-level verification metadata is complete. Workload tags were added on 2026-07-17 as a first-pass editorial classification for API pricing UX. Cost calculator uses standard pricing; batch and caching discounts apply separately. Actual costs may vary based on model routing, context length, and feature usage.