Groq API

Inference provider built for latency — it serves other people's open-weight models, fast.

Official docs ›

Price per million tokens

List price, read from the provider's own page · verified 2026-08-04

ModelInput / 1MOutput / 1MCached in
groq/compound
groq/compound-mini
llama-3.1-8b-instant$0.05$0.08
llama-3.3-70b-versatile$0.59$0.79
openai/gpt-oss-120b$0.15$0.6
openai/gpt-oss-20b$0.075$0.3
openai/gpt-oss-safeguard-20b$0.075$0.3
qwen/qwen3.6-27b$0.6$3

Compared with

Missing a model, or a price that moved? Prices here are only ever copied from the vendor's own page and stamped with the date they were read — never estimated. A verified run measures what the price table cannot: time to first token, throughput under concurrency, function-calling reliability.