Experimental records

Everything JevBench has run or analysed, in order. A run collects decisions from a model. An analysis re-reads decisions already collected and asks something else of them, so it never increases the evidence base.

These are first-party records
JevBench executed these itself, so they carry no evidence tier: the A–D scale grades independence from the party with an interest in the result, and here that party is us. Evidence collected from other people lives in the benchmark archive and is kept separate.
2 records

Supporting pages: the 608-case failure archive · the RUN 001 replay · how evidence is handled