“On my last transaction it seem that my top-up was not successful.”
Is Jev actually good enough?
Judge models by evidence, not vibes.
Real runs. Real latency. Real cost. Real failures.
- test samples
- 3,080
- accuracy
- 80.3%
- median latency
- 310 ms
- estimated API cost
- $0.222
Replay generated from the actual Banking77 benchmark manifest. Accelerated for review; muted; no synthetic cases.
FAST.
CHEAP.
TYPED.
TypeSafe describes Jev as “Decisions, not strings” and “more like code”: typed decisions and probabilities for software to combine. JEVBench keeps that product language as the claim under test, not as the site’s visual identity.
OFFICIAL PRODUCT LANGUAGE ↗- REQUESTS
- 3,080
- ACCURACY
- 80.3%
- FAILURES
- 608
- CLIENT P50
- 310 ms
- EST. API COST
- $0.222
We preserve the failures too.
“My card is just not working at this time.”
“My ATM got stuck and I'm not sure what to do.”
RUN 002 audits every one of these decisions for whether Jev’s own confidence flags them. 29 came back wrong at confidence 1.00 — no threshold reaches those at any setting.
| RUN | DATASET | MODEL | SAMPLES | ACCURACY | P50 | COST | STATUS |
|---|---|---|---|---|---|---|---|
| 001 | Banking77 | jev-1.13.0 | 3,080 | 80.3% | 310ms | $0.222* | [ FIRST-PARTY ] |
| 003 | Ad detection · SMS + YouTube | jev-1.13.0 | 200 | 94.0% | 480ms | $0.005* | [ FIRST-PARTY ] |
| 005 | Chinese vs English · MASSIVE | jev-1.13.0 | 2,974 | 79.3% | 113ms | $0.162* | [ FIRST-PARTY ] |
| 006 | Browser next step · Mind2Web | jev-1.13.0 | 9,378 | 38.4% | 151ms | $0.660* | [ FIRST-PARTY ] |
| 007 | Page context · Mind2Web | jev-1.13.0 | 9,378 | 42.3% | 153ms | $1.076* | [ FIRST-PARTY ] |
| 008 | Jev vs free open model · MASSIVE | jev-1.13.0 | 2,974 | 81.7% | 113ms | $0.162* | [ FIRST-PARTY ] |
| 009 | Jev vs open decision models · MASSIVE | jev-1.13.0 | 2,974 | 81.7% | 113ms | $0.162* | [ FIRST-PARTY ] |
* DERIVED ESTIMATE. FIRST-PARTY = RUN BY JEVBENCH, SO IT CARRIES NO EVIDENCE TIER AND IS NOT INDEPENDENT EVIDENCE. ARTIFACTS PUBLISHED IN FULL.
OPEN FULL ARCHIVE →The dataset commit, prompt protocol, result manifest, scoring rule, latency definition, cost basis and limitations travel with the run. No universal Jev score is inferred from one classification task.
This is not a formality. On one collected phishing set, Jev’s recall moves from 85.7% to 98.4% on nothing but a change of prompt — same model, same rows. A Jev number without its protocol is not a weak claim; it is not a claim at all.
Have a production workload?
Register interest without uploading customer data. No payment, no public result.