Home / Benchmark methodology
Benchmark / validation status

Measure first. Publish only what can be reproduced.

This page defines the benchmark protocol and publishes a small source-checked pricing snapshot. Latency, throughput and quality leaderboards are intentionally withheld until a complete, reproducible run is available.

Updated 2026-08-23 · VerticalAPI research · Coverage gaps are explicit
MEASURED RESULTS PENDING

No current performance winner is claimed. Earlier illustrative values and statements about 1,000 calls or weekly refreshes did not have completed production evidence in this repository and have been withdrawn.

Verified now

Provider-published pricing snapshot.

Standard USD rates per 1M text tokens. Batch, cache, fast-mode, regional, tool and other surcharges are excluded unless noted.

ModelProviderInputOutputConditionCheckedSource
GPT-5.6 TerraOpenAI$2$12Standard, short context2026-08-23Official ↗
Claude Sonnet 5Anthropic$2$10Introductory through 2026-08-31; then $3 / $152026-08-23Official ↗
Gemini 3.1 Pro PreviewGoogle$2$12Prompt ≤200k; $4 / $18 above 200k2026-08-23Official ↗

This is a focused current snapshot, not complete market coverage. Use the model directory and pricing comparison for broader research; check each page's evidence date.

Measurement protocol

What a publishable run must include.

The protocol is a preregistration, not proof that the measurements have been completed.

01

Fixed workloads

Versioned prompts, modalities, input/output targets, temperature and tool settings for every provider.

02

Comparable access

Model ID, API tier, region, concurrency, retries and provider route recorded for each request.

03

Raw provenance

Timestamped request metadata and calculation code retained so p50, p95 and throughput can be recomputed.

04

Limitations review

Coverage gaps, time-of-day effects, judge bias and gateway overhead published beside the result.

Evidence states

What readers can—and cannot—infer.

Each dataset block must name its evidence state.

VERIFIED SOURCE

Provider facts

Pricing and documented limits can be published when the first-party source, date and conditions are recorded.

NOT PUBLISHED

Measured performance

Latency, throughput and quality remain unavailable until the complete run and provenance pass validation.

EDITORIAL ANALYSIS

Workload fit

Recommendations must cite their constraints and cannot be presented as a universal benchmark winner.

FAQ

Benchmark boundaries.

Why remove the old performance tables?

The repository marked the values as illustrative placeholders pending a real harness run, while other parts of the page described them as measured production data. Keeping both statements would violate the evidence policy.

When will latency and quality results return?

After a complete run has versioned inputs, request-level provenance, calculation code, repeatability checks and a written limitations section. No date is promised until that evidence exists.

Are the prices a VerticalAPI measurement?

No. They are provider-published standard list prices checked against official sources. Pricing evidence and runtime performance measurements are different data types.

How should I choose a provider today?

Start with the comparison directory, shortlist providers that meet the documented constraints, then benchmark those providers with your own prompts, region, account tier and concurrency.