Every AI API, priced from the source.
14 providers, input and output price per million tokens, each figure read off the vendor's own page and stamped with the date it was read — catalogue pass 2026-08-04, with high-change examples rechecked separately. Treat each page's date as the evidence boundary.
Pricing is provider evidence. Runtime performance is a different evidence class: the benchmark page publishes the protocol and withholds latency or quality rankings until a reproducible run is complete.
The reference API most tooling targets first, and the shape every other provider on this page imitates.
Claude models, with prompt caching priced separately and long-context tiers billed at a different rate.
Gemini through the AI Studio API, distinct from the Vertex AI path used inside Google Cloud.
European provider with open-weight and hosted models on the same API.
Grok models behind an OpenAI-compatible endpoint.
Enterprise-oriented models with retrieval and reranking as first-class endpoints.
Low list prices on reasoning-capable models, OpenAI-compatible.
The Llama family, also resold by most inference providers further down this page.
Kimi models, long context, OpenAI-compatible.
Qwen models through Model Studio.
Inference provider built for latency — it serves other people's open-weight models, fast.
Inference on wafer-scale hardware; the pitch is tokens per second, not model choice.
Broad catalogue of open-weight models, with fine-tuning on the same platform.
Inference provider serving open-weight models on its own silicon.