Cerebras API

Inference on wafer-scale hardware; the pitch is tokens per second, not model choice.

Official docs ›

Price per million tokens

List price, read from the provider's own page · verified 2026-08-04

ModelInput / 1MOutput / 1MCached in
gemma-4-31b-it$0.99$1.49
gpt-oss-120b$0.35$0.75
zai-glm-4.7$2.25$2.75

Compared with

Missing a model, or a price that moved? Prices here are only ever copied from the vendor's own page and stamped with the date they were read — never estimated. A verified run measures what the price table cannot: time to first token, throughput under concurrency, function-calling reliability.