Groq API
Inference provider built for latency — it serves other people's open-weight models, fast.
Price per million tokens
List price, read from the provider's own page · verified 2026-08-04
| Model | Input / 1M | Output / 1M | Cached in |
|---|---|---|---|
| groq/compound | — | — | — |
| groq/compound-mini | — | — | — |
| llama-3.1-8b-instant | $0.05 | $0.08 | — |
| llama-3.3-70b-versatile | $0.59 | $0.79 | — |
| openai/gpt-oss-120b | $0.15 | $0.6 | — |
| openai/gpt-oss-20b | $0.075 | $0.3 | — |
| openai/gpt-oss-safeguard-20b | $0.075 | $0.3 | — |
| qwen/qwen3.6-27b | $0.6 | $3 | — |
Compared with
- Groq vs Cerebras: 2026 comparison — VerticalAPI
- Groq vs DeepInfra: 2026 comparison - VerticalAPI
- Groq vs Fireworks: 2026 comparison — VerticalAPI
- Groq vs Together AI: 2026 comparison — VerticalAPI
Missing a model, or a price that moved? Prices here are only ever copied from the vendor's own page and stamped with the date they were read — never estimated. A verified run measures what the price table cannot: time to first token, throughput under concurrency, function-calling reliability.