The same six gates on graded product relevance: four of six fail
Rank query–product pairs by a Jev probability against human relevance grades, using the same pre-registered gates as the easy corpus.
| Model | Accuracy | Recall | FPR | AUROC | ECE | Latency | Cost |
|---|---|---|---|---|---|---|---|
Jev | — | — | — | — | 0.242 | — | not recorded |
Jev | — | — | — | — | — | — | not recorded |
Jev | — | — | — | — | 0.279 | — | not recorded |
The same six pre-registered gate conditions and thresholds as the newsgroups corpus, applied to graded product relevance rather than topic membership. 296 requests; the author reports about $0.015 of spend for this probe, covering all call modes together. Recorded as its own run because it is a different dataset and a different judgment.
- n=306 pairs from only 30 queries, so the effective sample of independent judgments is smaller than the row count suggests.
- No baseline model was run on this probe, so this shows how Jev behaves on graded relevance, not how it compares.
- No confidence intervals for the primary metrics.
- No latency figures are reported.
- No per-call cost is recorded: the stated spend covers all call modes together.
- Graded relevance is a harder and more subjective judgment than topic membership, which is the point of the probe but also means the human grades carry their own disagreement.
What this benchmark does not prove
- 01That Jev is poorly calibrated. It is poorly calibrated on graded product relevance; the same gates passed on topic membership.
- 02That Jev is unsuitable for search ranking. One probe of 306 pairs from 30 queries, with no baseline, does not establish that.
- 03That a competing model would do better here. None was run.
- 04That the gate thresholds are the right ones for your application. They were the author's pre-registered choices, not an industry standard.
jev-orderby-bench — is ORDER BY over a Jev probability defensible? — yodablocks, verified 2026-09-19
Reproduction artifacts: github.com/yodablocks/jev-orderby-bench
Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet
Six pre-registered gate conditions with stated thresholds, run against an easy corpus and a deliberately hard probe. Author states API costs were paid by the author and used TypeSafe's published $0.042/MTok rate for cost estimates.