Cerebras API
Inference on wafer-scale hardware; the pitch is tokens per second, not model choice.
Price per million tokens
List price, read from the provider's own page · verified 2026-08-04
| Model | Input / 1M | Output / 1M | Cached in |
|---|---|---|---|
| gemma-4-31b-it | $0.99 | $1.49 | — |
| gpt-oss-120b | $0.35 | $0.75 | — |
| zai-glm-4.7 | $2.25 | $2.75 | — |
Compared with
- Cerebras vs Fireworks: 2026 comparison - VerticalAPI
- Cerebras vs Together AI: 2026 comparison - VerticalAPI
- Groq vs Cerebras: 2026 comparison — VerticalAPI
Missing a model, or a price that moved? Prices here are only ever copied from the vendor's own page and stamped with the date they were read — never estimated. A verified run measures what the price table cannot: time to first token, throughput under concurrency, function-calling reliability.