Fixed workloads
Versioned prompts, modalities, input/output targets, temperature and tool settings for every provider.
This page defines the benchmark protocol and publishes a small source-checked pricing snapshot. Latency, throughput and quality leaderboards are intentionally withheld until a complete, reproducible run is available.
No current performance winner is claimed. Earlier illustrative values and statements about 1,000 calls or weekly refreshes did not have completed production evidence in this repository and have been withdrawn.
Standard USD rates per 1M text tokens. Batch, cache, fast-mode, regional, tool and other surcharges are excluded unless noted.
| Model | Provider | Input | Output | Condition | Checked | Source |
|---|---|---|---|---|---|---|
| GPT-5.6 Terra | OpenAI | $2 | $12 | Standard, short context | 2026-08-23 | Official ↗ |
| Claude Sonnet 5 | Anthropic | $2 | $10 | Introductory through 2026-08-31; then $3 / $15 | 2026-08-23 | Official ↗ |
| Gemini 3.1 Pro Preview | $2 | $12 | Prompt ≤200k; $4 / $18 above 200k | 2026-08-23 | Official ↗ |
This is a focused current snapshot, not complete market coverage. Use the model directory and pricing comparison for broader research; check each page's evidence date.
The protocol is a preregistration, not proof that the measurements have been completed.
Versioned prompts, modalities, input/output targets, temperature and tool settings for every provider.
Model ID, API tier, region, concurrency, retries and provider route recorded for each request.
Timestamped request metadata and calculation code retained so p50, p95 and throughput can be recomputed.
Coverage gaps, time-of-day effects, judge bias and gateway overhead published beside the result.
Each dataset block must name its evidence state.
Pricing and documented limits can be published when the first-party source, date and conditions are recorded.
Latency, throughput and quality remain unavailable until the complete run and provenance pass validation.
Recommendations must cite their constraints and cannot be presented as a universal benchmark winner.
The repository marked the values as illustrative placeholders pending a real harness run, while other parts of the page described them as measured production data. Keeping both statements would violate the evidence policy.
After a complete run has versioned inputs, request-level provenance, calculation code, repeatability checks and a written limitations section. No date is promised until that evidence exists.
No. They are provider-published standard list prices checked against official sources. Pricing evidence and runtime performance measurements are different data types.
Start with the comparison directory, shortlist providers that meet the documented constraints, then benchmark those providers with your own prompts, region, account tier and concurrency.