Experimental records
Everything JevBench has run or analysed, in order. A run collects decisions from a model. An analysis re-reads decisions already collected and asks something else of them, so it never increases the evidence base.
These are first-party records
JevBench executed these itself, so they carry no evidence tier: the A–D scale grades independence from the party with an interest in the result, and here that party is us. Evidence collected from other people lives in the benchmark archive and is kept separate.
2 records
RUN 001COLLECTED DECISIONS→
Banking77 × Jev — complete test split
The complete BANKING77 test split sent to Jev 1.13.0 as typed Choice decisions, published with its pinned dataset commit, its protocol, the runner, and every per-row result.
- Decisions
- 3,080
- Accuracy
- 80.3%
- Failures kept
- 608
- Run date
- 2026-09-20
ANALYSIS 002NO NEW EVIDENCE→
Can Jev know when it is wrong?
An offline audit of those same decisions: does Jev's confidence flag its own errors? It makes no new API calls and collects no new evidence.
- Decisions audited
- 3,080
- Error AUROC
- 0.853
- Coverage @ 95%
- 59.1%
- Wrong at conf 1.00
- 29
Supporting pages: the 608-case failure archive · the RUN 001 replay · how evidence is handled