AI decision models, measured
A decision model does not write answers. It reads a piece of text and a typed question (pick one option, score on a scale, yes or no) and returns a probability for every option. That is the shape of many decisions inside an agent: which tool to call, which model to route to, which category a request belongs in, whether to refuse. Jev was the first one: TypeSafe, which makes it, calls the class System One. Since mid-September a run of open models with the same shape has followed. This page lists them, and what JevBench has and has not measured, under protocols committed before each run.
MASSIVE’s 2,974 English test commands (CC BY 4.0, human labels), one 60-way intent question, the option descriptions Jev was given. Every figure links to the record that measured it.
| MODEL | KIND | LABELLED EXAMPLES USED | ACCURACY | 95% (WILSON) | RECORD |
|---|---|---|---|---|---|
| BGE-M3 + logistic regression | Trained classifier | 11,514 labelled examples | 88.1% | 86.9%–89.2% | RUN 008 |
| Jev | Decision model (API) | None | 81.7% | 80.3%–83.1% | RUN 005 |
| OpenDecider-nano | Open decision model | None | 54.4% | 52.6%–56.2% | RUN 009 |
| open-jev (DeBERTa-v3-large) | Open decision model | None | 52.9% | 51.1%–54.7% | RUN 009 |
| BGE-M3, zero-shot | Open embedding model | None | 48.3% | 46.5%–50.1% | RUN 008 |
| Laya-multilingual | Open decision model | None | 43.9% | 42.2%–45.7% | RUN 009 |
| Laya (English checkpoint) | Open decision model | None | 37.3% | 35.5%–39.0% | RUN 009 |
The trained classifier is not a like-for-like comparison: it learned from labelled examples nobody else saw. The open decision models were run as released, on CPU, without examples; their cards describe other tasks, and some describe the models as bases to fine-tune. RUN 009 sets out how to read the gap.
| MODEL | MAKER | LICENCE | SIZE | LANGUAGES | RUNS | HOW YOU ASK IT | JEVBENCH |
|---|---|---|---|---|---|---|---|
| Jev | TypeSafe | Closed weights; API only | Not published | English primary; other languages accepted at lower accuracy, per its docs | TypeSafe API ($0.042 per million input tokens, output free, per its docs) | State plus typed questions: Choice, Score, Noul (yes/no) | Measured: 001, 003, 005, 006, 007, 008, 009 |
| OpenDecider-nano | manjunathshiva | Apache-2.0 | About 400M parameters, per its card | English | Locally, CPU or GPU | State plus typed questions (choice, score, noul) | Measured: 009 |
| Laya | convaiinnovations | Apache-2.0 | 421M (English) and 322M (multilingual), per its card | English checkpoint; a multilingual checkpoint claims 100+ languages | Locally, CPU or GPU | State plus typed questions (choice, score, noul) | Measured: 009 |
| open-jev (DeBERTa-v3-large) | com-kotobalabs | Apache-2.0 | DeBERTa-v3-large backbone | English | Locally, CPU or GPU | State plus typed questions (choice, score, noul) | Measured: 009 |
| decider-2b | Mapika | Apache-2.0 | 1.9B parameters (Qwen3.5-2B base); 4B, 12B and 35B-A3B siblings, per its card | English | Locally (open weights) | State plus typed questions with explicit option lists | Not yet tested |
| GLiNER2.5-Decide | fastino | Apache-2.0 | 340M, per its card; a 1B sibling and a multilingual variant exist | English | Locally (open weights) | Text plus a label set given at call time; single- or multi-label | Not yet tested |
| strands-decider-2B | StrandsAgents | Apache-2.0 | LoRA adapter and readout head on Qwen3.5-2B-Base | Not stated on its card | Locally, CPU, CUDA or Apple MPS | State plus typed questions (noul, choice, score) | Not yet tested |
| pplx-decider-v1-27b | Perplexity | Apache-2.0 | 27B parameters (Qwen3.8-27B base) | Not stated on its card | A CUDA GPU with room for about 49 GiB of weights, per its card | Decision questions through its own inference script | Not yet tested |
- Facts are from each model’s own card or documentation, read on 2026-10-05; follow the name for the source.
- Laya: Its card describes the released checkpoints as a fast base to specialise rather than a zero-shot decision engine.
- open-jev (DeBERTa-v3-large): Not affiliated with TypeSafe, per its card.
- decider-2b: Its card calls it an open reproduction of the System One model class.
- strands-decider-2B: Its card names model routing, tool selection, argument checking, triage and guardrails as intended uses.
- No accuracy is shown for a model JevBench has not run. Figures the models’ authors report stay on their cards.
| RECORD | QUESTION | DATA | ITEMS | JEV ACCURACY |
|---|---|---|---|---|
| RUN 001 | Banking77 × Jev — complete test split | BANKING77 (complete official test split) | 3,080 | 80.3% |
| RUN 003 | Can Jev tell an ad from everything else? | VE-001 AI-adjudicated subset of two UCI spam corpora | 200 | 94.0% |
| RUN 005 | Does Jev understand Chinese? | MASSIVE 1.1 test split, paired en-US and zh-CN | 2,974 | 79.3% |
| RUN 006 | Can Jev pick the next click? | Mind2Web test splits (cross-task, cross-website, cross-domain) | 9,378 | 38.4% |
| RUN 007 | Does more page context help Jev? | Mind2Web test splits (cross-task, cross-website, cross-domain) | 9,378 | 42.3% |
Different tasks with different scoring; the figures are not a single Jev score and should not be averaged.
- Agent decisions beyond intent: tool selection, model routing and guardrails. Tool selection is the next experiment, with two of the most-downloaded open decision models on Hugging Face (decider-2b, GLiNER2.5-Decide) beside Jev.
- A large LLM on the same items. The question this page exists to answer, which decisions do not need a large model, needs one as the baseline; it will be run only through an official API, so the model behind it is known.
- The larger open models, such as pplx-decider-v1-27b, which need a GPU this project does not have.
Some model cards report results on a “JevBench public hard tier”. That is a different project with a similar name. jevbench.xyz publishes no such tier and is not affiliated with it; the records here are the RUN pages, each with its protocol and per-item results.