Experimental record / run 005 / experiment VE-002 / rev A

Does Jev understand Chinese?

[ FIRST-PARTY ][ PRE-REGISTERED ][ HUMAN LABELS ]

Every JevBench record so far was in English. This one takes 2,974 short virtual-assistant commands from Amazon’s MASSIVE test set — each in English and in a professionally localized Chinese version, with the same human intent label — and asks jev-1.13.0 to pick one of 60 intents, three ways. The questions and how each answer would be read were committed before the first request.

ENGLISH INPUT
81.7%
CHINESE INPUT
79.3%
DIFFERENCE
−2.4 pp
CHINESE-WRITTEN QUESTION
−0.1 pp
What the pre-registered rules say
Chinese input: lower. With the question unchanged, Jev scored −2.4 pp on Chinese (95% paired-bootstrap interval −3.5 pp to −1.3 pp), and the same conclusion holds without the 30 Chinese items MASSIVE's own reviewers doubted. Chinese-written question: no measurable difference. Translating the question as well moved accuracy by −0.1 pp (−0.6 pp to +0.4 pp) and used 16% more input tokens per request: more cost for no measurable gain.
01 / Three arms, same items
ARMACCURACY95% INTERVALERROR AUROCKEPT AT 0.90ACCURACY KEPTP50TOKENS / REQ$ / 1K
A · English input · English question81.7%80.3%–83.1%0.85671.0%93.9%113 ms1,295$0.054
B · Chinese input · English question79.3%77.8%–80.7%0.84667.7%92.5%113 ms1,298$0.055
C · Chinese input · Chinese question79.2%77.7%–80.6%0.84969.2%92.5%113 ms1,502$0.063

Error AUROC is how well confidence separates right answers from wrong ones (0.5 is chance). Confidence flagged errors about as well in Chinese as in English. Absolute accuracy is a floor: the options are the bare intent names, deliberately unedited.

02 / Paired comparisons
COMPARISONITEMSDIFFERENCE95% BOOTSTRAPMcNEMAR pONLY FIRST RIGHTONLY SECOND RIGHTVERDICT
A → B2,974−2.4 pp−3.5 pp to −1.3 pp< 0.00116998lower
B → C2,974−0.1 pp−0.6 pp to +0.4 pp0.8133734no measurable difference
A → C2,974−2.5 pp−3.6 pp to −1.3 pp< 0.001181107lower
A → B · confirmed only2,944−1.9 pp−3.0 pp to −0.9 pp< 0.00115396lower
B → C · confirmed only2,944−0.1 pp−0.6 pp to +0.5 pp0.8133734no measurable difference
A → C · confirmed only2,944−2.0 pp−3.2 pp to −1.0 pp< 0.001165105lower

Rules fixed before the run: a difference counts only if its interval excludes zero; “no measurable difference” only if the interval sits within ±2 pp. “Confirmed only” drops the 30 items where at least two of MASSIVE’s reviewers said the Chinese sentence did not match the intent.

03 / Where Chinese lost

The gap is not spread evenly. Per-scenario figures are descriptive — the protocol tests only the overall difference — and small scenarios move a lot with a few items.

ENGLISH INPUTCHINESE INPUT · SAME ENGLISH QUESTIONPP40%60%80%100%ACCURACY BY SCENARIOcalendar (402 items): English 72.1%, Chinese 64.7%calendar · 402−7.5social (106 items): English 81.1%, Chinese 75.5%social · 106−5.7music (81 items): English 72.8%, Chinese 67.9%music · 81−4.9weather (156 items): English 91.7%, Chinese 87.2%weather · 156−4.5iot (220 items): English 92.3%, Chinese 88.2%iot · 220−4.1audio (62 items): English 88.7%, Chinese 85.5%audio · 62−3.2news (124 items): English 89.5%, Chinese 86.3%news · 124−3.2transport (124 items): English 88.7%, Chinese 86.3%transport · 124−2.4play (387 items): English 88.1%, Chinese 85.8%play · 387−2.3datetime (103 items): English 92.2%, Chinese 90.3%datetime · 103−1.9general (189 items): English 42.9%, Chinese 41.8%general · 189−1.1lists (142 items): English 89.4%, Chinese 88.7%lists · 142−0.7email (271 items): English 81.5%, Chinese 81.5%email · 271±0.0recommendation (94 items): English 81.9%, Chinese 81.9%recommendation · 94±0.0takeaway (57 items): English 77.2%, Chinese 77.2%takeaway · 57±0.0qa (288 items): English 87.8%, Chinese 88.5%qa · 288+0.7alarm (96 items): English 96.9%, Chinese 99.0%alarm · 96+2.1cooking (72 items): English 56.9%, Chinese 62.5%cooking · 72+5.6
04 / Right in English, wrong in Chinese: calendar

calendar lost most (72.1% → 64.7%). The first 8 items there that Jev got right in English and wrong in Chinese, under the identical English question:

ENGLISHCHINESELABELENGLISH ANSWERCHINESE ANSWER
what's my upcoming week look like我下周会是什么样子calendar_querycalendar_query 1.00general_quirky 0.52
mark today as the start of my diet把今天作为我节食的开始calendar_setcalendar_set 0.91lists_createoradd 0.38
is there anything i should be reminded about有什么事情是我需要做的吗calendar_querycalendar_query 0.43general_greet 0.55
cancel the breakfast at tiffany's house取消在一点利的早饭calendar_removecalendar_remove 1.00alarm_remove 0.46
erase all the events清除所有事项calendar_removecalendar_remove 1.00lists_remove 0.95
set reminder for noon to get lunch with boss设置一个提醒中午跟老板吃午餐calendar_setcalendar_set 0.51alarm_set 0.66
make sure there is no events on my calendar确保我的日历上没有事项calendar_removecalendar_remove 0.88calendar_query 0.72
can you remind my next meeting with my boss one hour before it will happen你能在下一个和我老板的会议发生一小时前提醒我吗calendar_setcalendar_set 0.65alarm_set 0.62
05 / Method

Pre-registered before any request. Every MASSIVE test item was sent once under three arms, interleaved: A English utterance with an English question; B the professionally localized Chinese utterance with the identical English question; C the Chinese utterance with the question written in Chinese. One Choice over all 60 MASSIVE intents, option descriptions produced mechanically from the intent names, no examples. Paired comparisons use a 10,000-resample bootstrap and an exact McNemar test. 8,922 requests, $0.51 at list price, all resolved to jev-1.13.0.

06 / Limitations
  • Option descriptions are the intent names with underscores replaced by spaces, deliberately plain and identical across arms A and B; absolute accuracy is therefore a floor for this task rather than a ceiling.
  • The Chinese question in arm C was written by Claude, which also writes this project's code. Some intent names are opaque, so translating them meant choosing a meaning; C versus B measures a Chinese-written question as a whole, not language alone.
  • MASSIVE's own reviewers doubted that 30 of the Chinese utterances kept their intent; they are included in the headline figures and removed in a pre-registered sensitivity analysis.
  • Utterances are short, spoken-style virtual-assistant commands in Mainland Simplified Chinese. Written, domain-specific or longer Chinese text may behave differently.
  • Each arm ran once. A repeat of a different task (RUN 004) found about 1% of labels change between identical runs.
  • Latency is client-observed round-trip time from a cloud runner whose region was not recorded; all three arms were interleaved, so they are comparable with each other but not with other runs.
  • Cost is derived from recorded input tokens at list price, not reconciled to an invoice.
07 / What this run does not prove
  • It does not show how Jev compares with any other model on Chinese. No other model was run.
  • It does not show how Jev handles Chinese beyond short virtual-assistant commands, or other Chinese varieties such as Traditional Chinese.
  • It does not show that MASSIVE's labels are correct; its label noise is shared by every arm.
  • It is not an independent result: JevBench chose the dataset, wrote the protocol, ran Jev and wrote the analysis.
08 / Data and licence

MASSIVE is © Amazon.com, Inc. or its affiliates, released under CC BY 4.0 and localized from SLURP (CC BY 4.0). FitzGerald et al. (2023), MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages, ACL 2023, doi:10.18653/v1/2023.acl-long.235. JevBench paired the en-US and zh-CN test split by id and added a flag from MASSIVE’s own reviewer scores; text and labels are unchanged.