JEVBENCH / INDEPENDENT MODEL EVALUATION
ARCHIVE 001 — BANKING77

Is Jev actually good enough?

Judge models by evidence, not vibes.

Real runs. Real latency. Real cost. Real failures.

test samples
3,080
accuracy
80.3%
median latency
310 ms
estimated API cost
$0.222
Recorded real run
jev-1.13.0 · Banking77 · 3,080 cases

Replay generated from the actual Banking77 benchmark manifest. Accelerated for review; muted; no synthetic cases.

01 / The claim

FAST.
CHEAP.
TYPED.

TypeSafe describes Jev as “Decisions, not strings” and “more like code”: typed decisions and probabilities for software to combine. JEVBench keeps that product language as the claim under test, not as the site’s visual identity.

OFFICIAL PRODUCT LANGUAGE ↗
02 / The test
REQUESTS
3,080
ACCURACY
80.3%
FAILURES
608
CLIENT P50
310 ms
EST. API COST
$0.222

We preserve the failures too.

03 / Where Jev gets it wrong
CASE 1039 / FAILURE RECORD

“On my last transaction it seem that my top-up was not successful.”

EXPECTEDtop_up_reverted
PREDICTEDtop_up_failed
CONFIDENCE1.00
LATENCY326 ms
CASE 2545 / FAILURE RECORD

“My card is just not working at this time.”

EXPECTEDvirtual_card_not_working
PREDICTEDcard_not_working
CONFIDENCE1.00
LATENCY290 ms
CASE 1935 / FAILURE RECORD

“My ATM got stuck and I'm not sure what to do.”

EXPECTEDcard_swallowed
PREDICTEDatm_support
CONFIDENCE1.00
LATENCY262 ms

RUN 002 audits every one of these decisions for whether Jev’s own confidence flags them. 29 came back wrong at confidence 1.00 — no threshold reaches those at any setting.

04 / Benchmark archive
RUNDATASETMODELSAMPLESACCURACYP50COSTSTATUS
001Banking77jev-1.13.03,08080.3%310ms$0.222*[ FIRST-PARTY ]
003Ad detection · SMS + YouTubejev-1.13.020094.0%480ms$0.005*[ FIRST-PARTY ]
005Chinese vs English · MASSIVEjev-1.13.02,97479.3%113ms$0.162*[ FIRST-PARTY ]
006Browser next step · Mind2Webjev-1.13.09,37838.4%151ms$0.660*[ FIRST-PARTY ]
007Page context · Mind2Webjev-1.13.09,37842.3%153ms$1.076*[ FIRST-PARTY ]
008Jev vs free open model · MASSIVEjev-1.13.02,97481.7%113ms$0.162*[ FIRST-PARTY ]
009Jev vs open decision models · MASSIVEjev-1.13.02,97481.7%113ms$0.162*[ FIRST-PARTY ]

* DERIVED ESTIMATE. FIRST-PARTY = RUN BY JEVBENCH, SO IT CARRIES NO EVIDENCE TIER AND IS NOT INDEPENDENT EVIDENCE. ARTIFACTS PUBLISHED IN FULL.

OPEN FULL ARCHIVE →
05 / Trust the protocol, inspect the limits

The dataset commit, prompt protocol, result manifest, scoring rule, latency definition, cost basis and limitations travel with the run. No universal Jev score is inferred from one classification task.

This is not a formality. On one collected phishing set, Jev’s recall moves from 85.7% to 98.4% on nothing but a change of prompt — same model, same rows. A Jev number without its protocol is not a weak claim; it is not a claim at all.

[ METHODOLOGY ]
06 / Private evaluation

Have a production workload?

Register interest without uploading customer data. No payment, no public result.

[ TEST YOUR WORKLOAD ]