Experimental record / run 001 / rev A

Banking77 × Jev

[ VERIFIED ]

Complete official test split, sent as zero-shot 77-way typed Choice decisions to jev-1.13.0. “Verified” means the run completed and its public manifest, protocol and summary agree; it does not mean an independent party reproduced it.

TEST SAMPLES
3,080
ACCURACY
80.3%
CORRECT / WRONG
2,472 / 608
P50 / P95
310 / 896 ms
EST. API COST
$0.222
01 / Experimental setup
DATASET
BANKING77 official test split / CC-BY-4.0
DATASET COMMIT
57ec275d8078af65b7731c2a98be812d844a6d6b
SAMPLES / LABELS
3,080 / 77
MODEL
jev-1.13.0
TASK
zero-shot 77-way intent classification
RUN DATE
2026-09-20
PROMPT PROTOCOL
One Choice question; label descriptions normalized only from dataset label names. No examples, tuning or fallback.
02 / Results
2,472
correct
608
wrong
19.7%
error rate
Cost is an estimate derived from 5,296,450 recorded input tokens at the $0.042/million list price observed 2026-09-20. Latency is client-observed request time and includes network and service overhead; it is not model compute time.
03 / Latency distribution
<250
63
250–349
2241
350–499
423
500–749
156
750–999
150
≥1000
47

BANDS IN MILLISECONDS / COUNTS AT RIGHT

04 / Confidence vs correctness
0.0–0.1
0.1–0.2
0.0%
0.2–0.3
10.0%
0.3–0.4
25.0%
0.4–0.5
28.1%
0.5–0.6
35.4%
0.6–0.7
51.6%
0.7–0.8
63.2%
0.8–0.9
70.4%
0.9–1.0
92.6%

BAR = SAMPLE COUNT / RIGHT = OBSERVED ACCURACY

05 / Failure archive
CASE 1039 / WRONG

On my last transaction it seem that my top-up was not successful.

EXPECTED
top_up_reverted
PREDICTED
top_up_failed
CONFIDENCE
1.00
LATENCY
326 ms
CASE 2545 / WRONG

My card is just not working at this time.

EXPECTED
virtual_card_not_working
PREDICTED
card_not_working
CONFIDENCE
1.00
LATENCY
290 ms
CASE 1935 / WRONG

My ATM got stuck and I'm not sure what to do.

EXPECTED
card_swallowed
PREDICTED
atm_support
CONFIDENCE
1.00
LATENCY
262 ms
[ VIEW ALL 608 FAILURE CASES ]
06 / Against the collected evidence

JevBench already held a third-party run on this dataset: Banking77 intent classification: Jev vs nano, frontier and a supervised encoder, a 208-row paired sample scoring Jev at 83.2% (95% CI 77.988.0%). This run puts the full split at 80.3% (95% CI 78.881.6%, Wilson).

The two agree: our full-split figure falls inside the smaller run’s confidence interval. That is the useful result here — not that Jev scored 80.3%, but that an independent 208-row estimate held up when someone ran all 3,080 rows.
This run has no baseline, and that is its biggest weakness. On the same dataset, the third-party run scored BGE-small-en-v1.5 + logistic regression at 93.3% — well above Jev. A number measured against the dataset’s labels says nothing about whether Jev is the right choice for the task.
07 / Replay

The replay is generated from the published result manifest. It is accelerated for review; it does not replay wall-clock timing.

[ WATCH FULL RUN ]
08 / Limitations
  • No baseline was run. This measures Jev against the dataset's labels, not against any alternative model, so it says nothing about whether Jev is the right choice for this task.
  • Latency is client-observed round-trip time with no network floor measured, so it is not model compute time and cannot be compared against a latency figure from any other run.
  • The runner issues requests concurrently and the artifacts do not record the concurrency used for this run, which further rules out comparing these latencies with a serially measured figure.
  • Cost is derived from recorded input tokens at a list price, not from an invoice. The provider returned no per-request cost on any of the 3,080 calls.
  • The 77 label descriptions are produced mechanically from the dataset's label names. Different criteria, worked examples or task decomposition would likely change the result.
  • The provider returned no request identifier on any call, so an individual published row cannot be traced back to a provider-side record.
  • Three rows record a prediction that is not the argmax of their own published probability vector; in all three the top two labels are exactly tied at two-decimal rounding (0.35/0.35, 0.24/0.24, 0.50/0.50).
09 / What this run does not prove
  • It does not show that Jev is accurate in general. This is one 77-way banking intent task, and JevBench's own collected evidence puts Jev's accuracy anywhere from 62.6% to 98.6% depending on the task and the prompt.
  • It does not show that Jev beats any alternative on this task, because no alternative was run beside it. A third-party run on a 208-row sample of this same split scored a supervised encoder baseline well above Jev.
  • It does not establish that Jev's confidence scores are safe to route on outside this dataset. The threshold table is fitted on the same rows it is measured on.
  • It is not an independent result. JevBench chose the dataset, wrote the prompt and published the write-up, and a run we executed ourselves cannot corroborate our own reporting.
  • It does not measure the latency you would see. No network floor was measured and the client's region is not pinned in the artifacts.

Every failure is preserved. The evidence is the record, not the headline number.

READ TECHNICAL NOTE 001 →