Experimental record / run 008 / experiment VE-005 / rev A

Is Jev better than a free model?

[ FIRST-PARTY ][ PRE-REGISTERED ][ FIRST COMPARISON ][ 0 NEW JEV REQUESTS ]

Every earlier record says no other model was run. This one runs one: BGE-M3, a free, openly licensed multilingual model that runs on an ordinary CPU, on the same 2,974 MASSIVE commands Jev answered in RUN 005, in English and Chinese. Once with only what Jev was given, once trained on MASSIVE’s labelled examples. The design was committed first.

JEV, NO EXAMPLES
81.7%
FREE MODEL, NO EXAMPLES
48.3%
FREE MODEL + 11,514 EXAMPLES
88.1%
NEW API REQUESTS
0

ENGLISH ACCURACY. CHINESE BELOW.

The pre-registered headline
Given only what Jev was given, a free open model scored 48.3% (English) and 46.1% (Chinese) against Jev’s 81.7% and 79.3%. Trained on MASSIVE’s 11,514 labelled examples per language, it scored 88.1% and 85.2%. Without examples Jev is far ahead; with enough labelled data, a cheap trained classifier passes it. The trained comparison is not like-for-like: Jev saw none of those examples.
01 / Accuracy
0%25%50%75%100%SAME INFORMATION AS JEVWITH LABELLED TRAINING DATAFree open model (BGE-M3)no examplesFree open model (BGE-M3), English: 48.3% (46.5%–50.1%)ENGLISH48.3%Free open model (BGE-M3), Chinese: 46.1% (44.3%–47.9%)CHINESE46.1%Jevno examplesJev, English: 81.7% (80.3%–83.1%)ENGLISH81.7%Jev, Chinese: 79.3% (77.8%–80.7%)CHINESE79.3%BGE-M3 + logistic regression11,514 labelled examplesBGE-M3 + logistic regression, English: 88.1% (86.9%–89.2%)ENGLISH88.1%BGE-M3 + logistic regression, Chinese: 85.2% (83.8%–86.4%)CHINESE85.2%
METHODTRAINING DATAENGLISH95%CHINESE95%CHINESE, CONFIRMED ITEMSERROR AUROC (EN / ZH)
Jevno examples81.7%80.3%–83.1%79.3%77.8%–80.7%80.0%0.856 / 0.846
Free open model (BGE-M3)no examples48.3%46.5%–50.1%46.1%44.3%–47.9%46.5%0.775 / 0.723
BGE-M3 + logistic regression11,514 labelled examples88.1%86.9%–89.2%85.2%83.8%–86.4%85.7%0.881 / 0.871

“Confirmed items” are the 2,944 Chinese utterances MASSIVE’s own reviewers agreed kept their intent. Error AUROC measures how well each method’s confidence separates its right answers from its wrong ones; 0.5 is chance.

02 / The four comparisons
COMPARISONDIFFERENCE95% (PAIRED BOOTSTRAP)ONLY OPEN MODEL RIGHTONLY JEV RIGHTREADING
Free model, no examples − Jev (English)−33.4 pp−35.4 pp to −31.4 pp1131,107open model worse
Free model, no examples − Jev (Chinese)−33.2 pp−35.2 pp to −31.3 pp1311,119open model worse
Free model + 11,514 examples − Jev (English)+6.4 pp+5.0 pp to +7.7 pp313123open model better
Free model + 11,514 examples − Jev (Chinese)+5.9 pp+4.5 pp to +7.2 pp311137open model better

Each comparison is on the same items. It reads “better” or “worse” only when its interval excludes zero, as the protocol fixed.

03 / What it means
  • Without labelled data, Jev is far ahead here. Given only the command and 60 plain intent names, the free model was right on 48.3% of English commands; Jev on 81.7%. Jev got 1,107 commands right that the free model missed, and missed 113 it got.
  • With labelled data, the free model passes Jev. Trained on 11,514 examples, it scored +6.4 pp above Jev in English and +5.9 pp in Chinese, and its confidence flagged its errors better (AUROC 0.881 against 0.856). The catch is the data: someone had to label those examples, and a new intent needs new ones. Jev needs only a name for it.
  • Chinese costs everyone about the same. From English to Chinese, Jev lost 2.4 points, the free model 2.2 points without examples and 2.9 points with them. Neither open setup’s loss differs clearly from Jev’s (intervals −1.8 pp to +1.5 pp and −0.8 pp to +1.9 pp).
04 / Cost and speed, side by side
JEVFREE OPEN MODEL
Where it runsTypeSafe API, over the networkYour own machine (here: 4 CPU threads)
Price$0.054 per 1,000 commands at list priceNo API fee; hardware not priced
Time per command (p50 / p95)113 / 160 ms round trip50.1 / 67.3 ms local compute
Needs labelled examplesNoNo for zero-shot; yes for the trained setup

The two time columns measure different things — a network round trip from a cloud runner in an unrecorded region, and local CPU time — so they are shown side by side and not as a ratio. The open model’s speed will differ on other hardware; its accuracy should not.

05 / Method

Pre-registered before the open model saw any test item. BAAI/bge-m3 (MIT, pinned revision, ONNX on CPU) classifies RUN 005's frozen MASSIVE test items two ways: zero-shot, choosing the option description most similar to the utterance, using the same 60 descriptions Jev was given; and trained, a logistic regression on the same embeddings fitted to MASSIVE's train split in the same language, with C chosen on the dev split. Jev's answers are RUN 005's published rows; no new requests were sent. Comparisons are paired, with intervals from 10,000 bootstrap resamples of items. The trained classifier chose C=10 in both languages (dev accuracy 88.4% English, 86.1% Chinese).

06 / Limitations
  • Jev's answers are one run from RUN 005, reused. RUN 004 found about 1% of Jev's labels change between identical runs.
  • One free model, chosen in advance as a reasonable option, not the strongest possible one. Larger or differently trained open models may do better.
  • The option descriptions are the intent names with underscores replaced by spaces, the same plain text Jev got. Better-written descriptions could help either side.
  • The trained classifier's C was chosen from {0.1, 1, 10} on the dev split; it picked 10, the top of the range, in both languages, so a wider search could score slightly higher.
  • Latency is not comparable across the two sides: Jev's is a network round trip to an API from a cloud runner whose region was not recorded; the open model's is local CPU time on this project's cloud container, one utterance at a time.
  • Jev's cost is derived from recorded input tokens at list price. The open model has no API fee, but the hardware that runs it is not free and is not priced here.
07 / What this run does not prove
  • It does not show how Jev compares with open models in general, or with commercial LLMs; one free embedding model was run.
  • It does not show how Jev compares on any task other than short virtual-assistant commands in English and Chinese.
  • The trained comparison does not show that the open model is better than Jev: the classifier used 11,514 labelled examples per language that Jev never saw.
  • It is not an independent result: JevBench chose the models, wrote the protocol, ran them and wrote the analysis.
08 / Data, model and licences

MASSIVE: FitzGerald et al. (2023), ACL 2023, licensed CC BY 4.0. BGE-M3: Chen et al. (2024), BAAI, MIT licence, revision 5617a9f61b028005a4858fdac845db406aefb181, run with onnxruntime 1.20.1 and scikit-learn 1.6.1.