Fastest LLM APIs for real-time in 2026: what the vendors actually publish

Speed comparisons for inference clouds are full of precise-looking latency numbers with no source. Two facts shape this page: Groq and Cerebras publish their own tokens-per-second figures, Fireworks and SambaNova publish none — and <strong>not one of the four publishes time-to-first-token</strong>. So this page carries throughput where the vendor states it, prices for the models they actually serve today, and no TTFT at all. An earlier version of this page quoted 150-300ms TTFT figures; no source supports them.

Pick by constraint, not by headline speed

Peak output speed

Cerebras

Publishes the highest throughput of the four — around 3,000 tokens/s on gpt-oss-120b. The catalogue is the constraint: three models, one of them preview-only and one carrying a deprecation date of 17 August 2026. Its Code Pro and Max plans are both sold out.

  • gpt-oss-120b $0.35 / $0.75
  • No Llama, no Qwen, no Mistral
  • GLM 4.7 deprecates 2026-08-17
Broadest production catalogue

Groq

Six production models with per-model throughput published on the pricing page. Note what it no longer serves: Mixtral shut down in March 2025 and both Llama 4 models in 2026. Qwen 3.6 is preview-only, not production.

  • Llama 3.3 70B $0.59 / $0.79 at 394 tok/s
  • gpt-oss-20b $0.075 / $0.30 at 1,000 tok/s
  • Compound systems have no published per-token price
Widest model variety

Fireworks

The frontier open-weight catalogue — Kimi K3, DeepSeek V4 Pro and Flash, GLM 5.2, Qwen 3.7, MiniMax M3. Publishes no speed figure at all. Watch the pricing structure: non-headline models are billed by parameter band, which is not a model price.

  • Kimi K3 $3.00 / $15.00
  • DeepSeek V4 Flash $0.14 / $0.28
  • Priority ~1.25-1.5x, Fast ~1.5x Standard
Cheapest gpt-oss-120b

SambaNova

Undercuts every other cloud here on gpt-oss-120b at $0.22 / $0.59, and serves Llama 3.3 70B at $0.60 / $1.20. Publishes prices only — no throughput, no latency, no SLA.

  • gpt-oss-120b $0.22 / $0.59
  • Llama 3.3 70B $0.60 / $1.20
  • DeepSeek V3.1 / V3.2 $3.00 / $4.50
Voice agents

OpenAI Realtime API

Now serves gpt-realtime-2.1, gpt-realtime-translate and gpt-live-transcribe. The gpt-4o-realtime-preview family was shut down on 7 May 2026. Budget on audio tokens, not text: they cost 8 to 16 times the text rate.

  • gpt-realtime-2.1: $4 / $24 text, $32 / $64 audio
  • mini: $0.60 / $2.40 text, $10 / $20 audio
  • gpt-4o-realtime-preview shut down 2026-05-07

Inference clouds: what their own pages state

DimensionCerebrasGroqFireworksSambaNova
Vendor-published throughput~3,000 tok/s (gpt-oss-120b)394 tok/s (Llama 3.3 70B), 1,000 (gpt-oss-20b)None publishedNone published
Time to first tokenNot publishedNot publishedNot publishedNot published
Cheapest model (in / out per 1M)$0.35 / $0.75 (gpt-oss-120b)$0.05 / $0.08 (Llama 3.1 8B)$0.07 / $0.30 (gpt-oss-20b)$0.22 / $0.59 (gpt-oss-120b)
Catalogue size3 models — and one deprecates 2026-08-176 production models + 2 agentic systems9 headline models + size-band pricing6 models
Models actually servedgpt-oss-120b, Gemma 4 31B (preview), GLM 4.7 (deprecating)Llama 3.3 70B, Llama 3.1 8B, gpt-oss 120B/20B, Qwen 3.6 (preview), CompoundKimi K3, DeepSeek V4 Pro/Flash, GLM 5.2, Qwen 3.7, MiniMax M3, gpt-ossMiniMax M2.7, DeepSeek V3.1/3.2, Gemma 4, gpt-oss-120b, Llama 3.3 70B
Serves any Llama modelNoYes (3.3 70B, 3.1 8B)No headline LlamaYes (3.3 70B at $0.60 / $1.20)
Published SLANone foundNone foundNone foundNone found
Best forPeak output speed on one open-weight modelBroadest production catalogue at low costModel variety and frontier open-weightsCheapest gpt-oss-120b

Every price and throughput figure read on the vendor's own pricing page on 4 August 2026. Throughput is the vendor's own claim, not our measurement — Cerebras' page footer notes its comparisons may vary by workload, configuration and date, and third-party measurement disagrees: Artificial Analysis measures 289 tok/s for Groq's Llama 3.3 70B against Groq's published 394. Time-to-first-token is blank because no vendor publishes it; the lowest figure any third party reports is around 0.5s, so the sub-300ms numbers this page used to carry were wrong by a factor of 2 to 6.

VerticalAPI verdict

Choose on catalogue and price, and treat speed claims as vendor marketing until you measure them on your own prompts. If you need peak output speed on one open-weight model, Cerebras publishes the highest figure — but check its three-model catalogue covers you and note GLM 4.7 deprecates on 17 August 2026. If you need a production catalogue at low cost, Groq is the practical default and publishes per-model throughput. If you need frontier open-weights, Fireworks has the widest selection. If you are running gpt-oss-120b specifically, SambaNova is the cheapest. For voice, the only current option is OpenAI's gpt-realtime line, and audio tokens will dominate your bill. Whatever you pick, benchmark it yourself: no vendor here publishes TTFT, and vendor throughput and independent measurement disagree by about a third. VerticalAPI reaches three of these four under BYOK on one endpoint — it adds a hop, so it will not be faster than going direct.

Get started — one endpoint, your own keys →

Frequently asked questions

Which LLM API is fastest for real-time in 2026?

On vendor-published throughput, Cerebras — around 3,000 tokens/s on gpt-oss-120b. Groq publishes 1,000 tokens/s on gpt-oss-20b and 394 on Llama 3.3 70B. Fireworks and SambaNova publish no figures. Treat all of these as claims: Cerebras' own page notes results vary by workload and date, and independent measurement puts Groq's Llama 3.3 70B at 289 tokens/s rather than 394.

What is the time to first token on these providers?

Nobody publishes it. Not Groq, not Cerebras, not Fireworks, not SambaNova. The only figures available are third-party, and the lowest of those is around 0.5s with most between 0.9 and 1.3 seconds — so the sub-300ms numbers commonly quoted, including in an earlier version of this page, have no basis. If TTFT matters to your product, measure it yourself on your own prompt shapes.

What happened to GPT-4o Realtime?

It was shut down on 7 May 2026, along with gpt-4o-mini-realtime-preview and the gpt-4o audio previews. The replacements are gpt-realtime-1.5 and now gpt-realtime-2.1, at $4 / $24 per 1M text tokens and $32 / $64 per 1M audio tokens. Base gpt-4o is not deprecated and remains at $2.50 / $10, but it is not a realtime model.

How much do audio tokens cost compared to text?

Eight to sixteen times more. On gpt-realtime-2.1, text is $4 / $24 per 1M while audio is $32 / $64. Any voice-agent cost model built on text rates alone is wrong by roughly an order of magnitude, which is the single biggest budgeting mistake in this category.

Which cloud is cheapest for the same model?

gpt-oss-120b is served by all four, which makes it the one clean comparison: SambaNova $0.22 / $0.59, Cerebras $0.35 / $0.75, Fireworks $0.15 / $0.60, Groq $0.15 / $0.60. That is a 2.3x spread on input for identical weights — the reason to keep provider choice a one-line change rather than an integration.

Does VerticalAPI make these providers faster?

No, and it would be dishonest to claim otherwise. Routing through any gateway adds a network hop, so calling Groq or Cerebras directly will always be marginally faster. What VerticalAPI gives you is one OpenAI-compatible endpoint across 11 providers under BYOK, so switching between them — or comparing their prices on the same prompt — is a model-string change. It makes one attempt against one provider and does not retry or fail over.

Limitations of this comparison

  • Throughput figures here are vendor claims, not our measurements. Cerebras' own footer states its comparisons vary by workload, configuration, date and model.
  • Where we cite third-party measurement it is Artificial Analysis, retrieved 2026-08-04 — the only honest date, because it carries no timestamp, no measurement window and no run count. Its Groq data still lists Llama 4 Scout, a model Groq shut down on 2026-07-17.
  • Third-party latency columns labelled “time to first answer token” include reasoning time on thinking models, so they are not comparable to TTFT.
  • We publish no latency measurements of our own for these providers. Our own benchmark data dates from April 2026 and covers models since superseded — see /benchmark/.
  • Qwen 3.6 on Groq and Gemma 4 31B on Cerebras are preview models marked for evaluation only, not production use.
  • Voice-agent cost is dominated by audio tokens, which are billed separately from text; any estimate built on text rates alone understates the bill by roughly an order of magnitude.
  • VerticalAPI reaches Groq, Cerebras and Fireworks under BYOK but adds a proxy hop and therefore cannot be faster than calling them directly. It makes one attempt against one provider and does not retry.

What may change in 12-24 months

  1. Catalogues are shrinking, not growing: Cerebras is down to three models with one deprecating on 2026-08-17 and its Code plans sold out, while Groq retired Mixtral and both Llama 4 models.
  2. The open-weight centre of gravity moved to Chinese labs — Kimi K3, DeepSeek V4, GLM 5.2, MiniMax M3 and Qwen now dominate these catalogues where Llama and Mixtral did in 2025.
  3. OpenAI replaced the entire gpt-4o realtime and audio family on 7 May 2026 with the gpt-realtime line, and the pricing moved with it.
  4. gpt-oss became the common denominator: all four clouds serve it, which makes it the one model where their prices are directly comparable — a 4.7x spread from $0.075 to $0.35 on input.

Related questions

ChatGPT, Perplexity and Gemini usually suggest these next.

  • Which inference cloud is cheapest for gpt-oss-120b?
  • Why does nobody publish time-to-first-token?
  • What replaced OpenAI's GPT-4o Realtime API?
  • How much do audio tokens add to a voice agent's bill?
  • Which cloud still serves Llama 3.3 70B, and at what price?

Run a provider we have not measured? Any public endpoint is benchmarked and listed for free in the API directory. If you want yours measured now, on the record — time to first token, throughput under concurrency, cost per million tokens, function-calling reliability — see how a verified run works. The result is published exactly as it comes out.