Decision models / what has been measured / updated 2026-10-05

AI decision models, measured

A decision model does not write answers. It reads a piece of text and a typed question (pick one option, score on a scale, yes or no) and returns a probability for every option. That is the shape of many decisions inside an agent: which tool to call, which model to route to, which category a request belongs in, whether to refuse. Jev was the first one: TypeSafe, which makes it, calls the class System One. Since mid-September a run of open models with the same shape has followed. This page lists them, and what JevBench has and has not measured, under protocols committed before each run.

What the data says so far
On 2,974 public voice-assistant commands, given the same question and no examples, Jev chose the right one of 60 intents 81.7% of the time; the best open decision model, OpenDecider-nano, 54.4%. A small free classifier trained on 11,514 labelled examples reached 88.1%. That is one task. Tool selection, routing and guardrails have not been measured yet, and neither has any large LLM on the same items.
01 / Same items, same question

MASSIVE’s 2,974 English test commands (CC BY 4.0, human labels), one 60-way intent question, the option descriptions Jev was given. Every figure links to the record that measured it.

MODELKINDLABELLED EXAMPLES USEDACCURACY95% (WILSON)RECORD
BGE-M3 + logistic regressionTrained classifier11,514 labelled examples88.1%86.9%–89.2%RUN 008
JevDecision model (API)None81.7%80.3%–83.1%RUN 005
OpenDecider-nanoOpen decision modelNone54.4%52.6%–56.2%RUN 009
open-jev (DeBERTa-v3-large)Open decision modelNone52.9%51.1%–54.7%RUN 009
BGE-M3, zero-shotOpen embedding modelNone48.3%46.5%–50.1%RUN 008
Laya-multilingualOpen decision modelNone43.9%42.2%–45.7%RUN 009
Laya (English checkpoint)Open decision modelNone37.3%35.5%–39.0%RUN 009

The trained classifier is not a like-for-like comparison: it learned from labelled examples nobody else saw. The open decision models were run as released, on CPU, without examples; their cards describe other tasks, and some describe the models as bases to fine-tune. RUN 009 sets out how to read the gap.

02 / The models
MODELMAKERLICENCESIZELANGUAGESRUNSHOW YOU ASK ITJEVBENCH
JevTypeSafeClosed weights; API onlyNot publishedEnglish primary; other languages accepted at lower accuracy, per its docsTypeSafe API ($0.042 per million input tokens, output free, per its docs)State plus typed questions: Choice, Score, Noul (yes/no)Measured: 001, 003, 005, 006, 007, 008, 009
OpenDecider-nanomanjunathshivaApache-2.0About 400M parameters, per its cardEnglishLocally, CPU or GPUState plus typed questions (choice, score, noul)Measured: 009
LayaconvaiinnovationsApache-2.0421M (English) and 322M (multilingual), per its cardEnglish checkpoint; a multilingual checkpoint claims 100+ languagesLocally, CPU or GPUState plus typed questions (choice, score, noul)Measured: 009
open-jev (DeBERTa-v3-large)com-kotobalabsApache-2.0DeBERTa-v3-large backboneEnglishLocally, CPU or GPUState plus typed questions (choice, score, noul)Measured: 009
decider-2bMapikaApache-2.01.9B parameters (Qwen3.5-2B base); 4B, 12B and 35B-A3B siblings, per its cardEnglishLocally (open weights)State plus typed questions with explicit option listsNot yet tested
GLiNER2.5-DecidefastinoApache-2.0340M, per its card; a 1B sibling and a multilingual variant existEnglishLocally (open weights)Text plus a label set given at call time; single- or multi-labelNot yet tested
strands-decider-2BStrandsAgentsApache-2.0LoRA adapter and readout head on Qwen3.5-2B-BaseNot stated on its cardLocally, CPU, CUDA or Apple MPSState plus typed questions (noul, choice, score)Not yet tested
pplx-decider-v1-27bPerplexityApache-2.027B parameters (Qwen3.8-27B base)Not stated on its cardA CUDA GPU with room for about 49 GiB of weights, per its cardDecision questions through its own inference scriptNot yet tested
  • Facts are from each model’s own card or documentation, read on 2026-10-05; follow the name for the source.
  • Laya: Its card describes the released checkpoints as a fast base to specialise rather than a zero-shot decision engine.
  • open-jev (DeBERTa-v3-large): Not affiliated with TypeSafe, per its card.
  • decider-2b: Its card calls it an open reproduction of the System One model class.
  • strands-decider-2B: Its card names model routing, tool selection, argument checking, triage and guardrails as intended uses.
  • No accuracy is shown for a model JevBench has not run. Figures the models’ authors report stay on their cards.
03 / Jev on its own
RECORDQUESTIONDATAITEMSJEV ACCURACY
RUN 001Banking77 × Jev — complete test splitBANKING77 (complete official test split)3,08080.3%
RUN 003Can Jev tell an ad from everything else?VE-001 AI-adjudicated subset of two UCI spam corpora20094.0%
RUN 005Does Jev understand Chinese?MASSIVE 1.1 test split, paired en-US and zh-CN2,97479.3%
RUN 006Can Jev pick the next click?Mind2Web test splits (cross-task, cross-website, cross-domain)9,37838.4%
RUN 007Does more page context help Jev?Mind2Web test splits (cross-task, cross-website, cross-domain)9,37842.3%

Different tasks with different scoring; the figures are not a single Jev score and should not be averaged.

04 / Not measured yet
  • Agent decisions beyond intent: tool selection, model routing and guardrails. Tool selection is the next experiment, with two of the most-downloaded open decision models on Hugging Face (decider-2b, GLiNER2.5-Decide) beside Jev.
  • A large LLM on the same items. The question this page exists to answer, which decisions do not need a large model, needs one as the baseline; it will be run only through an official API, so the model behind it is known.
  • The larger open models, such as pplx-decider-v1-27b, which need a GPU this project does not have.
05 / A note on the name

Some model cards report results on a “JevBench public hard tier”. That is a different project with a similar name. jevbench.xyz publishes no such tier and is not affiliated with it; the records here are the RUN pages, each with its protocol and per-item results.