Evidence
Grouped by task, because that is the only grouping that means anything. Results from different datasets and protocols are never averaged, ranked against each other, or reduced to a single score.
- Tier A means the artifacts to re-run the benchmark are public. It does not mean anyone has re-run it — no run here has been independently reproduced yet.
- Latency is reported as the source measured it, with the network floor it contains. Anything we subtract is labelled derived and is not model compute time: the residual still contains queueing, gateway hops, TLS and serialisation.
- A run whose source never stated a date says so. An undated result cannot be placed against a model version or a price period.
- Confidence intervals appear where the source reported them. Where they are missing, the sample is usually small enough that the point estimate should not carry much weight.
A comparison page exists only where a single run measured both models on the same data under the same protocol, and both sides have verified pricing. No head-to-head, no page.
Run by a third party with no vendor involvement. Code and dataset or protocol are public, so the result could be re-run. This does not mean anyone has re-run it — see `independentlyReproduced`, which nothing in the dataset sets yet.
Run by a third-party researcher or outlet, but reproduction material is partial or absent.
Published by the vendor of the model being measured. Recorded for completeness, never treated as independent evidence.
A single hands-on session, blog post or social media report. Directional only.
Tiers grade provenance, not result strength. A small study can be tier A; its size is recorded in the run’s sample count and limitations, not in its tier.