Fireworks API
Broad catalogue of open-weight models, with fine-tuning on the same platform.
Price per million tokens
List price, read from the provider's own page · verified 2026-08-04
| Model | Input / 1M | Output / 1M | Cached in |
|---|---|---|---|
| deepseek-v4-flash | $0.14 | $0.28 | $0.028 |
| deepseek-v4-pro | $1.74 | $3.48 | $0.145 |
| glm-5.2 | $1.4 | $4.4 | $0.14 |
| gpt-oss-120b | $0.15 | $0.6 | $0.015 |
| gpt-oss-20b | $0.07 | $0.3 | $0.035 |
| kimi-k2.7-code | $0.95 | $4 | $0.19 |
| kimi-k3 | $3 | $15 | $0.3 |
| minimax-m3 | $0.3 | $1.2 | $0.06 |
| qwen-3.7-plus | $0.4 | $1.6 | $0.08 |
Compared with
- Cerebras vs Fireworks: 2026 comparison - VerticalAPI
- Fireworks vs Replicate: open-weight inference (2026) — VerticalAPI
- Groq vs Fireworks: 2026 comparison — VerticalAPI
- Lepton vs Fireworks: enterprise LLM inference (2026) — VerticalAPI
- OctoAI vs Fireworks: 2026 comparison - VerticalAPI
- Together AI vs Fireworks: open-weight inference (2026) — VerticalAPI
Missing a model, or a price that moved? Prices here are only ever copied from the vendor's own page and stamped with the date they were read — never estimated. A verified run measures what the price table cannot: time to first token, throughput under concurrency, function-calling reliability.