Jev against the open decision models
Within two weeks of Jev, open models with the same shape appeared: text and typed questions in, calibrated probabilities out, one forward pass. Their cards compare themselves with Jev on benchmarks their authors chose. This record runs OpenDecider-nano, Laya and open-jev on a dataset none of them lists, the same 2,974 MASSIVE commands Jev answered in RUN 005, with the identical question and option descriptions and no examples. The method and how it would be read were committed first.
- JEV, ENGLISH
- 81.7%
- BEST OPEN MODEL, ENGLISH
- 54.4%
- JEV, CHINESE
- 79.3%
- LAYA-MULTILINGUAL, CHINESE
- 34.8%
| MODEL | LANGUAGE | ACCURACY | 95% (WILSON) | WITHOUT THE 41 OVERLAP ITEMS | ANSWERS OUTSIDE THE 60 |
|---|---|---|---|---|---|
| Jev | English | 81.7% | 80.3%–83.1% | 81.6% | 0 |
| Jev | Chinese | 79.3% | 77.8%–80.7% | 79.2% | 0 |
| OpenDecider-nano | English | 54.4% | 52.6%–56.2% | 54.1% | 0 |
| open-jev | English | 52.9% | 51.1%–54.7% | 52.8% | 0 |
| Laya-multilingual | English | 43.9% | 42.2%–45.7% | 43.9% | 0 |
| Laya-multilingual | Chinese | 34.8% | 33.1%–36.5% | 34.7% | 0 |
| Laya, English checkpoint | English | 37.3% | 35.5%–39.0% | 37.3% | 0 |
| Laya, English, library defaults (secondary) | English | 42.3% | 40.5%–44.1% | 42.3% | 0 |
A model is run only in languages its card claims, so OpenDecider-nano and open-jev, both English-only, were not run on Chinese. Laya’s card says that with 50 or more options its defaults degrade and tells users to raise two settings; those are the primary arm. At the recommended settings it scored 37.3%; at its defaults, run as a secondary arm with no verdict, 42.3%.
| COMPARISON | DIFFERENCE | 95% (PAIRED BOOTSTRAP) | ONLY THE OPEN MODEL RIGHT | ONLY JEV RIGHT | READING | WITHOUT OVERLAP ITEMS |
|---|---|---|---|---|---|---|
| OpenDecider-nano − Jev (English) | −27.3 pp | −29.1 pp to −25.6 pp | 73 | 886 | model worse | −27.4 pp, model worse |
| Laya, English checkpoint − Jev (English) | −44.5 pp | −46.3 pp to −42.6 pp | 33 | 1,355 | model worse | −44.3 pp, model worse |
| open-jev − Jev (English) | −28.8 pp | −30.6 pp to −26.9 pp | 83 | 939 | model worse | −28.7 pp, model worse |
| Laya-multilingual − Jev (English) | −37.8 pp | −39.7 pp to −35.9 pp | 75 | 1,198 | model worse | −37.7 pp, model worse |
| Laya-multilingual − Jev (Chinese) | −44.5 pp | −46.5 pp to −42.5 pp | 80 | 1,404 | model worse | −44.5 pp, model worse |
Each comparison is on the same items. It reads “better” or “worse” only when its interval excludes zero, as the protocol fixed. The open models got right 33 to 83 commands that Jev missed; Jev got right 886 to 1,404 that they missed.
| MODEL | ERROR AUROC (TOP PROBABILITY) | AUROC, AS REPORTED | MEAN TOP PROBABILITY | ACCURACY | ECE (10 BINS) |
|---|---|---|---|---|---|
| Jev (English) | 0.851 | 0.856 | 89.7% | 81.7% | 0.080 |
| OpenDecider-nano (English) | 0.887 | 0.887 | 43.5% | 54.4% | 0.157 |
| open-jev (DeBERTa-v3-large) (English) | 0.847 | 0.847 | 29.7% | 52.9% | 0.233 |
| Laya-multilingual (English) | 0.802 | 0.803 | 69.5% | 43.9% | 0.255 |
| Laya (English checkpoint) (English) | 0.877 | 0.875 | 57.1% | 37.3% | 0.199 |
- Telling its right answers from its wrong ones. By error AUROC (0.5 is chance), OpenDecider-nano (0.887) and Laya (0.877) score above Jev (0.851); open-jev (0.847) is level with it; Laya-multilingual (0.802) is below.
- Saying how sure it is. Jev’s expected calibration error (0.080) is the lowest here; the open models’ run from 0.157 to 0.255. OpenDecider-nano and open-jev state less probability on their answer than they earn (43.5% against 54.4%, 29.7% against 52.9%); Laya states more (57.1% against 37.3%), and so does Jev, less so (89.7% against 81.7%).
- One caution about Jev’s figure. On 57% of its English answers Jev puts at least 99% of the probability on one option, so its ECE depends heavily on how confident its few wrong answers are.
- The
confidencefield each library returns is not comparable across them (OpenDecider’s is the top probability, Laya’s an entropy-based quantity, Jev’s a concentration statistic); it is the “as reported” column.
| WHERE IT RUNS | TIME PER COMMAND (P50 / P95) | PRICE | |
|---|---|---|---|
| Jev | TypeSafe API, over the network | 113 ms / 160 ms round trip | $0.054 per 1,000 at list price |
| OpenDecider-nano | Your own machine (here: 4 CPU threads) | 1.85 s / 2.13 s local compute | No API fee; hardware not priced |
| open-jev (DeBERTa-v3-large) | Your own machine (here: 4 CPU threads) | 1.19 s / 1.41 s local compute | No API fee; hardware not priced |
| Laya-multilingual | Your own machine (here: 4 CPU threads) | 617 ms / 733 ms local compute | No API fee; hardware not priced |
| Laya (English checkpoint) | Your own machine (here: 4 CPU threads) | 1.61 s / 1.82 s local compute | No API fee; hardware not priced |
Timed one command at a time on a machine doing nothing else, with 60 options per request. The two columns measure different things (a network round trip from a cloud runner in an unrecorded region against local CPU time), so they are shown beside each other and not as a ratio. The open models’ cards report GPU latencies, which were not measured here.
- This is not a verdict on those projects. OpenDecider-nano’s card reports 0.796 against Jev’s 0.754 on its own typed-decisions benchmark, and says the model was fine-tuned on that benchmark’s training split. Laya’s card calls its released checkpoints a fast base to specialise rather than a zero-shot decision engine. This record runs them as released, with no examples, on a different task. A result here that differs from their cards is not evidence the cards are wrong.
- For scale, RUN 008’s plain embedding model BGE-M3, given the same 60 option descriptions, scored 48.3% in English: above Laya-multilingual and Laya, below OpenDecider-nano and open-jev. That is context, not a pre-registered comparison.
- The question was written for Jev. It is the plain English instruction with intent names as options that Jev was given in RUN 005, and it was not tuned for the open models. Options written for each model, or a short fine-tune on the labels, are the obvious next step and are not tested here.
Pre-registered before any open model saw a test item, with a dated addendum on the interrupted and resumed runs. OpenDecider-nano, Laya (at the settings its card recommends for 50+ options, and at defaults as a secondary arm), Laya-multilingual and open-jev — all Apache-2.0, pinned to commit hashes, run locally on CPU — each answer the question Jev was given in RUN 005: its English instructions and the 60 intent names as option descriptions, one utterance as state, no examples. Models whose card claims a language are run in it. Jev's answers are RUN 005's published rows; no new requests were sent. Comparisons are paired, with intervals from 10,000 bootstrap resamples of items. Speed comes from separate one-at-a-time timing runs. 17,844 local answers across 6 arms, no API fees.
- Jev's answers are one run from RUN 005, reused. RUN 004 found about 1% of Jev's labels change between identical runs.
- One question for every model: the plain English instructions and intent-name options Jev was given in RUN 005. It was not adapted to the open models, and none was shown examples. Options written for each model, or a short fine-tune, could change their scores.
- The models are small (about 320M to 440M parameters), and their cards describe them partly as bases to specialise. OpenDecider's medium and large variants and the 27B-and-larger open models need GPUs and were not run.
- 41 English test commands appear verbatim in CLINC-OOS (two also in GLiClass), which OpenDecider-nano's card lists as training data. All results are also computed without them and no conclusion changes. Seven other datasets on that card, Laya's undisclosed training data and open-jev's listed Banking77, SST-5 and BoolQ were not checked for overlap.
- Speed is CPU time for these models: one request at a time, 4 threads, with 60 options. Their cards report GPU latencies, which were not measured. Jev's is a network round trip from a cloud runner whose region was not recorded, so the two are shown side by side and not as a ratio.
- At the settings Laya's card recommends for many options it scored lower than at its defaults. The recommended setting was fixed in advance as the primary arm; both are reported.
- The full runs were interrupted once, when the container was suspended, and resumed from their own row files with no row changed; the protocol's addendum records it.
- It does not show which model is best in general: one task, one question, short English and Chinese commands.
- It does not show that the models' own reported results are wrong. Those come from different tasks, and from checkpoints fine-tuned on those tasks' training data.
- It does not show how larger open variants, fine-tuned versions or other open decision models would score; none was run.
- It is not an independent result in the sense of a third-party reproduction: JevBench chose the models and the question, wrote the protocol, ran the models and wrote the analysis.
MASSIVE: FitzGerald et al. (2023), ACL 2023, licensed CC BY 4.0. All three models are Apache-2.0 and pinned to commit hashes: OpenDecider-nano 6ac21e98; open-jev (DeBERTa-v3-large) 188ee67a; Laya-multilingual 55cf4c4e; Laya (English checkpoint) 55cf4c4e. Two environments over one PyTorch, because the models disagree about transformers; the protocol’s addendum records the interrupted and resumed runs.