Is Jev better than a free model?
Every earlier record says no other model was run. This one runs one: BGE-M3, a free, openly licensed multilingual model that runs on an ordinary CPU, on the same 2,974 MASSIVE commands Jev answered in RUN 005, in English and Chinese. Once with only what Jev was given, once trained on MASSIVE’s labelled examples. The design was committed first.
- JEV, NO EXAMPLES
- 81.7%
- FREE MODEL, NO EXAMPLES
- 48.3%
- FREE MODEL + 11,514 EXAMPLES
- 88.1%
- NEW API REQUESTS
- 0
ENGLISH ACCURACY. CHINESE BELOW.
| METHOD | TRAINING DATA | ENGLISH | 95% | CHINESE | 95% | CHINESE, CONFIRMED ITEMS | ERROR AUROC (EN / ZH) |
|---|---|---|---|---|---|---|---|
| Jev | no examples | 81.7% | 80.3%–83.1% | 79.3% | 77.8%–80.7% | 80.0% | 0.856 / 0.846 |
| Free open model (BGE-M3) | no examples | 48.3% | 46.5%–50.1% | 46.1% | 44.3%–47.9% | 46.5% | 0.775 / 0.723 |
| BGE-M3 + logistic regression | 11,514 labelled examples | 88.1% | 86.9%–89.2% | 85.2% | 83.8%–86.4% | 85.7% | 0.881 / 0.871 |
“Confirmed items” are the 2,944 Chinese utterances MASSIVE’s own reviewers agreed kept their intent. Error AUROC measures how well each method’s confidence separates its right answers from its wrong ones; 0.5 is chance.
| COMPARISON | DIFFERENCE | 95% (PAIRED BOOTSTRAP) | ONLY OPEN MODEL RIGHT | ONLY JEV RIGHT | READING |
|---|---|---|---|---|---|
| Free model, no examples − Jev (English) | −33.4 pp | −35.4 pp to −31.4 pp | 113 | 1,107 | open model worse |
| Free model, no examples − Jev (Chinese) | −33.2 pp | −35.2 pp to −31.3 pp | 131 | 1,119 | open model worse |
| Free model + 11,514 examples − Jev (English) | +6.4 pp | +5.0 pp to +7.7 pp | 313 | 123 | open model better |
| Free model + 11,514 examples − Jev (Chinese) | +5.9 pp | +4.5 pp to +7.2 pp | 311 | 137 | open model better |
Each comparison is on the same items. It reads “better” or “worse” only when its interval excludes zero, as the protocol fixed.
- Without labelled data, Jev is far ahead here. Given only the command and 60 plain intent names, the free model was right on 48.3% of English commands; Jev on 81.7%. Jev got 1,107 commands right that the free model missed, and missed 113 it got.
- With labelled data, the free model passes Jev. Trained on 11,514 examples, it scored +6.4 pp above Jev in English and +5.9 pp in Chinese, and its confidence flagged its errors better (AUROC 0.881 against 0.856). The catch is the data: someone had to label those examples, and a new intent needs new ones. Jev needs only a name for it.
- Chinese costs everyone about the same. From English to Chinese, Jev lost 2.4 points, the free model 2.2 points without examples and 2.9 points with them. Neither open setup’s loss differs clearly from Jev’s (intervals −1.8 pp to +1.5 pp and −0.8 pp to +1.9 pp).
| JEV | FREE OPEN MODEL | |
|---|---|---|
| Where it runs | TypeSafe API, over the network | Your own machine (here: 4 CPU threads) |
| Price | $0.054 per 1,000 commands at list price | No API fee; hardware not priced |
| Time per command (p50 / p95) | 113 / 160 ms round trip | 50.1 / 67.3 ms local compute |
| Needs labelled examples | No | No for zero-shot; yes for the trained setup |
The two time columns measure different things — a network round trip from a cloud runner in an unrecorded region, and local CPU time — so they are shown side by side and not as a ratio. The open model’s speed will differ on other hardware; its accuracy should not.
Pre-registered before the open model saw any test item. BAAI/bge-m3 (MIT, pinned revision, ONNX on CPU) classifies RUN 005's frozen MASSIVE test items two ways: zero-shot, choosing the option description most similar to the utterance, using the same 60 descriptions Jev was given; and trained, a logistic regression on the same embeddings fitted to MASSIVE's train split in the same language, with C chosen on the dev split. Jev's answers are RUN 005's published rows; no new requests were sent. Comparisons are paired, with intervals from 10,000 bootstrap resamples of items. The trained classifier chose C=10 in both languages (dev accuracy 88.4% English, 86.1% Chinese).
- Jev's answers are one run from RUN 005, reused. RUN 004 found about 1% of Jev's labels change between identical runs.
- One free model, chosen in advance as a reasonable option, not the strongest possible one. Larger or differently trained open models may do better.
- The option descriptions are the intent names with underscores replaced by spaces, the same plain text Jev got. Better-written descriptions could help either side.
- The trained classifier's C was chosen from {0.1, 1, 10} on the dev split; it picked 10, the top of the range, in both languages, so a wider search could score slightly higher.
- Latency is not comparable across the two sides: Jev's is a network round trip to an API from a cloud runner whose region was not recorded; the open model's is local CPU time on this project's cloud container, one utterance at a time.
- Jev's cost is derived from recorded input tokens at list price. The open model has no API fee, but the hardware that runs it is not free and is not priced here.
- It does not show how Jev compares with open models in general, or with commercial LLMs; one free embedding model was run.
- It does not show how Jev compares on any task other than short virtual-assistant commands in English and Chinese.
- The trained comparison does not show that the open model is better than Jev: the classifier used 11,514 labelled examples per language that Jev never saw.
- It is not an independent result: JevBench chose the models, wrote the protocol, ran them and wrote the analysis.
MASSIVE: FitzGerald et al. (2023), ACL 2023, licensed CC BY 4.0. BGE-M3: Chen et al. (2024), BAAI, MIT licence, revision 5617a9f61b028005a4858fdac845db406aefb181, run with onnxruntime 1.20.1 and scikit-learn 1.6.1.