Experimental record / run 009 / experiment VE-006 / rev A

Jev against the open decision models

[ FIRST-PARTY ][ PRE-REGISTERED ][ 3 OPEN MODELS ][ 0 NEW JEV REQUESTS ]

Within two weeks of Jev, open models with the same shape appeared: text and typed questions in, calibrated probabilities out, one forward pass. Their cards compare themselves with Jev on benchmarks their authors chose. This record runs OpenDecider-nano, Laya and open-jev on a dataset none of them lists, the same 2,974 MASSIVE commands Jev answered in RUN 005, with the identical question and option descriptions and no examples. The method and how it would be read were committed first.

JEV, ENGLISH
81.7%
BEST OPEN MODEL, ENGLISH
54.4%
JEV, CHINESE
79.3%
LAYA-MULTILINGUAL, CHINESE
34.8%
The pre-registered headline
On MASSIVE, where none of them lists the data, Jev scored 81.7% in English; OpenDecider-nano scored 54.4%, open-jev 52.9%, Laya-multilingual 43.9% and Laya 37.3%. In Chinese Jev scored 79.3% and Laya-multilingual 34.8%. Every gap is large and none changes when the 41 commands found in training text are left out. This measures these models as released and without examples, on one task: it says nothing about their own benchmarks.
01 / Accuracy
0%25%50%75%100%JevJev, English: 81.7% (80.3%–83.1%)ENGLISH81.7%Jev, Chinese: 79.3% (77.8%–80.7%)CHINESE79.3%OpenDecider-nanoOpenDecider-nano, English: 54.4% (52.6%–56.2%)ENGLISH54.4%CHINESE NOT RUN: ITS CARD CLAIMS ENGLISH ONLYopen-jev (DeBERTa-v3-large)open-jev (DeBERTa-v3-large), English: 52.9% (51.1%–54.7%)ENGLISH52.9%CHINESE NOT RUN: ITS CARD CLAIMS ENGLISH ONLYLaya-multilingualLaya-multilingual, English: 43.9% (42.2%–45.7%)ENGLISH43.9%Laya-multilingual, Chinese: 34.8% (33.1%–36.5%)CHINESE34.8%Laya (English checkpoint)Laya (English checkpoint), English: 37.3% (35.5%–39.0%)ENGLISH37.3%CHINESE NOT RUN: ITS CARD CLAIMS ENGLISH ONLY
MODELLANGUAGEACCURACY95% (WILSON)WITHOUT THE 41 OVERLAP ITEMSANSWERS OUTSIDE THE 60
JevEnglish81.7%80.3%–83.1%81.6%0
JevChinese79.3%77.8%–80.7%79.2%0
OpenDecider-nanoEnglish54.4%52.6%–56.2%54.1%0
open-jevEnglish52.9%51.1%–54.7%52.8%0
Laya-multilingualEnglish43.9%42.2%–45.7%43.9%0
Laya-multilingualChinese34.8%33.1%–36.5%34.7%0
Laya, English checkpointEnglish37.3%35.5%–39.0%37.3%0
Laya, English, library defaults (secondary)English42.3%40.5%–44.1%42.3%0

A model is run only in languages its card claims, so OpenDecider-nano and open-jev, both English-only, were not run on Chinese. Laya’s card says that with 50 or more options its defaults degrade and tells users to raise two settings; those are the primary arm. At the recommended settings it scored 37.3%; at its defaults, run as a secondary arm with no verdict, 42.3%.

02 / The five comparisons
COMPARISONDIFFERENCE95% (PAIRED BOOTSTRAP)ONLY THE OPEN MODEL RIGHTONLY JEV RIGHTREADINGWITHOUT OVERLAP ITEMS
OpenDecider-nano − Jev (English)−27.3 pp−29.1 pp to −25.6 pp73886model worse−27.4 pp, model worse
Laya, English checkpoint − Jev (English)−44.5 pp−46.3 pp to −42.6 pp331,355model worse−44.3 pp, model worse
open-jev − Jev (English)−28.8 pp−30.6 pp to −26.9 pp83939model worse−28.7 pp, model worse
Laya-multilingual − Jev (English)−37.8 pp−39.7 pp to −35.9 pp751,198model worse−37.7 pp, model worse
Laya-multilingual − Jev (Chinese)−44.5 pp−46.5 pp to −42.5 pp801,404model worse−44.5 pp, model worse

Each comparison is on the same items. It reads “better” or “worse” only when its interval excludes zero, as the protocol fixed. The open models got right 33 to 83 commands that Jev missed; Jev got right 886 to 1,404 that they missed.

03 / Calibration, on the same quantity
MODELERROR AUROC (TOP PROBABILITY)AUROC, AS REPORTEDMEAN TOP PROBABILITYACCURACYECE (10 BINS)
Jev (English)0.8510.85689.7%81.7%0.080
OpenDecider-nano (English)0.8870.88743.5%54.4%0.157
open-jev (DeBERTa-v3-large) (English)0.8470.84729.7%52.9%0.233
Laya-multilingual (English)0.8020.80369.5%43.9%0.255
Laya (English checkpoint) (English)0.8770.87557.1%37.3%0.199
  • Telling its right answers from its wrong ones. By error AUROC (0.5 is chance), OpenDecider-nano (0.887) and Laya (0.877) score above Jev (0.851); open-jev (0.847) is level with it; Laya-multilingual (0.802) is below.
  • Saying how sure it is. Jev’s expected calibration error (0.080) is the lowest here; the open models’ run from 0.157 to 0.255. OpenDecider-nano and open-jev state less probability on their answer than they earn (43.5% against 54.4%, 29.7% against 52.9%); Laya states more (57.1% against 37.3%), and so does Jev, less so (89.7% against 81.7%).
  • One caution about Jev’s figure. On 57% of its English answers Jev puts at least 99% of the probability on one option, so its ECE depends heavily on how confident its few wrong answers are.
  • The confidence field each library returns is not comparable across them (OpenDecider’s is the top probability, Laya’s an entropy-based quantity, Jev’s a concentration statistic); it is the “as reported” column.
04 / Speed, side by side
WHERE IT RUNSTIME PER COMMAND (P50 / P95)PRICE
JevTypeSafe API, over the network113 ms / 160 ms round trip$0.054 per 1,000 at list price
OpenDecider-nanoYour own machine (here: 4 CPU threads)1.85 s / 2.13 s local computeNo API fee; hardware not priced
open-jev (DeBERTa-v3-large)Your own machine (here: 4 CPU threads)1.19 s / 1.41 s local computeNo API fee; hardware not priced
Laya-multilingualYour own machine (here: 4 CPU threads)617 ms / 733 ms local computeNo API fee; hardware not priced
Laya (English checkpoint)Your own machine (here: 4 CPU threads)1.61 s / 1.82 s local computeNo API fee; hardware not priced

Timed one command at a time on a machine doing nothing else, with 60 options per request. The two columns measure different things (a network round trip from a cloud runner in an unrecorded region against local CPU time), so they are shown beside each other and not as a ratio. The open models’ cards report GPU latencies, which were not measured here.

05 / Reading it fairly
  • This is not a verdict on those projects. OpenDecider-nano’s card reports 0.796 against Jev’s 0.754 on its own typed-decisions benchmark, and says the model was fine-tuned on that benchmark’s training split. Laya’s card calls its released checkpoints a fast base to specialise rather than a zero-shot decision engine. This record runs them as released, with no examples, on a different task. A result here that differs from their cards is not evidence the cards are wrong.
  • For scale, RUN 008’s plain embedding model BGE-M3, given the same 60 option descriptions, scored 48.3% in English: above Laya-multilingual and Laya, below OpenDecider-nano and open-jev. That is context, not a pre-registered comparison.
  • The question was written for Jev. It is the plain English instruction with intent names as options that Jev was given in RUN 005, and it was not tuned for the open models. Options written for each model, or a short fine-tune on the labels, are the obvious next step and are not tested here.
06 / Method

Pre-registered before any open model saw a test item, with a dated addendum on the interrupted and resumed runs. OpenDecider-nano, Laya (at the settings its card recommends for 50+ options, and at defaults as a secondary arm), Laya-multilingual and open-jev — all Apache-2.0, pinned to commit hashes, run locally on CPU — each answer the question Jev was given in RUN 005: its English instructions and the 60 intent names as option descriptions, one utterance as state, no examples. Models whose card claims a language are run in it. Jev's answers are RUN 005's published rows; no new requests were sent. Comparisons are paired, with intervals from 10,000 bootstrap resamples of items. Speed comes from separate one-at-a-time timing runs. 17,844 local answers across 6 arms, no API fees.

07 / Limitations
  • Jev's answers are one run from RUN 005, reused. RUN 004 found about 1% of Jev's labels change between identical runs.
  • One question for every model: the plain English instructions and intent-name options Jev was given in RUN 005. It was not adapted to the open models, and none was shown examples. Options written for each model, or a short fine-tune, could change their scores.
  • The models are small (about 320M to 440M parameters), and their cards describe them partly as bases to specialise. OpenDecider's medium and large variants and the 27B-and-larger open models need GPUs and were not run.
  • 41 English test commands appear verbatim in CLINC-OOS (two also in GLiClass), which OpenDecider-nano's card lists as training data. All results are also computed without them and no conclusion changes. Seven other datasets on that card, Laya's undisclosed training data and open-jev's listed Banking77, SST-5 and BoolQ were not checked for overlap.
  • Speed is CPU time for these models: one request at a time, 4 threads, with 60 options. Their cards report GPU latencies, which were not measured. Jev's is a network round trip from a cloud runner whose region was not recorded, so the two are shown side by side and not as a ratio.
  • At the settings Laya's card recommends for many options it scored lower than at its defaults. The recommended setting was fixed in advance as the primary arm; both are reported.
  • The full runs were interrupted once, when the container was suspended, and resumed from their own row files with no row changed; the protocol's addendum records it.
08 / What this run does not prove
  • It does not show which model is best in general: one task, one question, short English and Chinese commands.
  • It does not show that the models' own reported results are wrong. Those come from different tasks, and from checkpoints fine-tuned on those tasks' training data.
  • It does not show how larger open variants, fine-tuned versions or other open decision models would score; none was run.
  • It is not an independent result in the sense of a third-party reproduction: JevBench chose the models and the question, wrote the protocol, ran the models and wrote the analysis.
09 / Data, models and licences

MASSIVE: FitzGerald et al. (2023), ACL 2023, licensed CC BY 4.0. All three models are Apache-2.0 and pinned to commit hashes: OpenDecider-nano 6ac21e98; open-jev (DeBERTa-v3-large) 188ee67a; Laya-multilingual 55cf4c4e; Laya (English checkpoint) 55cf4c4e. Two environments over one PyTorch, because the models disagree about transformers; the protocol’s addendum records the interrupted and resumed runs.