Cheapest AI Models
Ranked by input/output pricing per million tokens, prompt caching savings, and high-volume batch processing discounts.
Most Cost-Efficient API Models
Ranked by verified official input token cost per million tokens.
Gemini 2.0 Flash
Google's high-speed multimodal workhorse with 1M token context, native tool use, and real-time audio/video streaming.
1M tokens
Not Disclosed
$0.10 / $0.40 / MTok
Llama 3.3 70B Instruct
Meta's open-weights 70B flagship matching Llama 3.1 405B capabilities on industry benchmarks at 1/5th the compute.
128k tokens
70B
$0.12 / $0.30 / MTok
DeepSeek-V3
Frontier open-weights 671B MoE base and chat model with multi-head latent attention (MLA) and dual-pipe training.
64k tokens
671B (37B active)
$0.14 / $0.28 / MTok
GPT-4o mini
High-speed, ultra-affordable multimodal small model designed to replace GPT-3.5 Turbo at 60% lower cost.
128k tokens
Not Disclosed
$0.15 / $0.60 / MTok
Codestral 2501
Mistral AI's dedicated code generation model with 256k context window and fill-in-the-middle (FIM) capabilities.
256k tokens
22B
$0.30 / $0.90 / MTok
DeepSeek-R1
Open-weights frontier reasoning model trained via large-scale reinforcement learning without supervised cold start.
64k tokens
671B (37B active)
$0.55 / $2.19 / MTok
Interactive Multi-Model Token Calculator
Token Cost Estimator
Calculate projected inference spend with prompt caching
Quick Workload Presets
$0.2000
2.00M input tokens
$0.2000
0.50M output tokens
$0.40
Avg: $0.00040 / req
Token Pricing FAQ
Which frontier AI model has the lowest token pricing?
Gemini 2.0 Flash ($0.10/MTok input, $0.40/MTok output) and DeepSeek-V3 ($0.14/MTok input, $0.28/MTok output) currently offer the lowest official API pricing for frontier-class capabilities.
How does batch API pricing work?
OpenAI, Anthropic, and Mistral offer 50% discounts on token rates for asynchronous queries submitted through their Batch API endpoints that complete within 24 hours.
Data Accuracy & Verification Notice
AI model specifications, pricing records, and benchmark metrics published on this platform are compiled directly from authoritative sources (official provider documentation, research papers, and verified evaluation harnesses). Benchmark results reflect specific test harnesses and prompting methodologies; scores are not directly comparable across differing evaluation setups.