Evidence

Grouped by task, because that is the only grouping that means anything. Results from different datasets and protocols are never averaged, ranked against each other, or reduced to a single score.

How to read these
  • Tier A means the artifacts to re-run the benchmark are public. It does not mean anyone has re-run it — no run here has been independently reproduced yet.
  • Latency is reported as the source measured it, with the network floor it contains. Anything we subtract is labelled derived and is not model compute time: the residual still contains queueing, gateway hops, TLS and serialisation.
  • A run whose source never stated a date says so. An undated result cannot be placed against a model version or a price period.
  • Confidence intervals appear where the source reported them. Where they are missing, the sample is usually small enough that the point estimate should not carry much weight.
Security detection
Classification
Intent & routing
Agent tool-call risk
Agent workloads
Direct comparisons

A comparison page exists only where a single run measured both models on the same data under the same protocol, and both sides have verified pricing. No head-to-head, no page.

Jev vs Claude Haiku 4.5 Jev vs GPT-5.4 nano — pending pricingJev vs GPT-5.6 Terra — pending pricingJev vs BGE-small-en-v1.5 + logistic regression — pending pricingJev vs TF-IDF + logistic regression — pending pricingJev vs Jev + TF-IDF ensemble — pending pricingJev vs Qwen 3.8 27B — pending pricingJev vs Needle 3 — pending pricing
Evidence tiers
A
Independent · Reproduction artifacts available

Run by a third party with no vendor involvement. Code and dataset or protocol are public, so the result could be re-run. This does not mean anyone has re-run it — see `independentlyReproduced`, which nothing in the dataset sets yet.

12 runs
B
Independent · Published

Run by a third-party researcher or outlet, but reproduction material is partial or absent.

0 runs
C
Vendor-reported

Published by the vendor of the model being measured. Recorded for completeness, never treated as independent evidence.

0 runs
D
Anecdotal

A single hands-on session, blog post or social media report. Directional only.

0 runs

Tiers grade provenance, not result strength. A small study can be tier A; its size is recorded in the run’s sample count and limitations, not in its tier.