← All evidence

The same six gates on graded product relevance: four of six fail

AIndependent · Reproduction artifacts availableClassificationAmazon ESCI (hard probe) · n = 3062026-09-18pre-registeredverdict: FAIL — four of six gates

Rank query–product pairs by a Jev probability against human relevance grades, using the same pre-registered gates as the easy corpus.

Accuracy
Every recorded metric
ModelAccuracyRecallFPRAUROCECELatencyCost
Jev
0.242not recorded
Jev
not recorded
Jev
0.279not recorded
Protocol

The same six pre-registered gate conditions and thresholds as the newsgroups corpus, applied to graded product relevance rather than topic membership. 296 requests; the author reports about $0.015 of spend for this probe, covering all call modes together. Recorded as its own run because it is a different dataset and a different judgment.

Dataset
Amazon ESCI (hard probe)
n = 306
306 query–product pairs drawn from 30 queries that carry all four human grades
Limitations (6)
  • n=306 pairs from only 30 queries, so the effective sample of independent judgments is smaller than the row count suggests.
  • No baseline model was run on this probe, so this shows how Jev behaves on graded relevance, not how it compares.
  • No confidence intervals for the primary metrics.
  • No latency figures are reported.
  • No per-call cost is recorded: the stated spend covers all call modes together.
  • Graded relevance is a harder and more subjective judgment than topic membership, which is the point of the probe but also means the human grades carry their own disagreement.
Note on this run
The most direct evidence in the dataset that Jev's calibration is task-dependent. Same author, same gates, same thresholds, same model version — and an ECE of 0.242 against 0.0453 on the easy corpus.

What this benchmark does not prove

  • 01That Jev is poorly calibrated. It is poorly calibrated on graded product relevance; the same gates passed on topic membership.
  • 02That Jev is unsuitable for search ranking. One probe of 306 pairs from 30 queries, with no baseline, does not establish that.
  • 03That a competing model would do better here. None was run.
  • 04That the gate thresholds are the right ones for your application. They were the author's pre-registered choices, not an industry standard.
Source and artifacts

jev-orderby-bench — is ORDER BY over a Jev probability defensible? yodablocks, verified 2026-09-19

Reproduction artifacts: github.com/yodablocks/jev-orderby-bench

Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet

Six pre-registered gate conditions with stated thresholds, run against an easy corpus and a deliberately hard probe. Author states API costs were paid by the author and used TypeSafe's published $0.042/MTok rate for cost estimates.