Every AI API, priced from the source.
14 providers, input and output price per million tokens, each figure read off the vendor's own page and stamped with the date it was read — last pass 2026-08-04. No estimate, no “starting from”, no number kept because it used to be true.
Price is the easy half. The other half — how fast the first token actually arrives, what throughput survives concurrency, whether function calling holds up — is measured, not quoted.
The reference API most tooling targets first, and the shape every other provider on this page imitates.
Claude models, with prompt caching priced separately and long-context tiers billed at a different rate.
Gemini through the AI Studio API, distinct from the Vertex AI path used inside Google Cloud.
European provider with open-weight and hosted models on the same API.
Grok models behind an OpenAI-compatible endpoint.
Enterprise-oriented models with retrieval and reranking as first-class endpoints.
Low list prices on reasoning-capable models, OpenAI-compatible.
The Llama family, also resold by most inference providers further down this page.
Kimi models, long context, OpenAI-compatible.
Qwen models through Model Studio.
Inference provider built for latency — it serves other people's open-weight models, fast.
Inference on wafer-scale hardware; the pitch is tokens per second, not model choice.
Broad catalogue of open-weight models, with fine-tuning on the same platform.
Inference provider serving open-weight models on its own silicon.