24 models, list price
Dollars per million tokens, read from each vendor's published page. Sort by what you're optimising for.
Sort: blended · input · output · cache · context · name
ModelProvider
InputOutput
Cache readContext
Ternary Bonsai 27B
Prism-ML/Ternary-Bonsai-27B
LFM2.5 8B-A1B
LiquidAI/LFM2.5-8B-A1B
Gemma 3n E4B IT
google/gemma-3n-E4B-it
GPT-OSS 20B
openai/gpt-oss-20b
Qwen3.5 9B
Qwen/Qwen3.5-9B
DeepSeek V4 Flash (0731)
deepseek-ai/DeepSeek-V4-Flash-0731
Qwen2.5 7B Instruct Turbo
Qwen/Qwen2.5-7B-Instruct-Turbo
GPT-OSS 120B
openai/gpt-oss-120b
Gemma 4 31B Instruct (Pearl AI)
pearl-ai/gemma-4-31b-it
Gemma 4 31B IT
google/gemma-4-31B-it
MiniMax M3
MiniMaxAI/MiniMax-M3
Inkling Small
thinkingmachines/Inkling-Small
Llama 3.3 70B Instruct Turbo
meta-llama/Llama-3.3-70B-Instruct-Turbo
Qwen3.7 Plus
Qwen/Qwen3.7-Plus
Cogito v2.1 671B
deepcogito/cogito-v2-1-671b
Qwen3.6 Plus
Qwen/Qwen3.6-Plus
Nemotron 3 Ultra 550B-A55B
nvidia/nemotron-3-ultra-550b-a55b
DeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-Pro
Qwen3.7 Max
Qwen/Qwen3.7-Max
Kimi K2.7 Code
moonshotai/Kimi-K2.7-Code
Inkling
thinkingmachines/Inkling
GLM-5.2
zai-org/GLM-5.2
Kimi K2.6
moonshotai/Kimi-K2.6
Kimi K3
moonshotai/Kimi-K3