Jev is cheap.Is it good enough?
Independent evidence for Jev.
Benchmarks · Cost · Accuracy · Fallback
Cheaper
Structured
enough?
Different tasks.
Different outcomes.
Across the runs recorded here Jev ranges from 62.6% to 95.4%. There is no universal Jev accuracy, and our schema has nowhere to store one.
PhishNChips v5.2 · Label an email as phishing or legitimate. · n=2,000 · evidence
Same model, same task, different prompt.
All three are Jev detecting phishing. The gap between the first two is nothing but the prompt, on the same 5,733-email set. A recall figure quoted without its protocol is not a fact about the model — which is the whole reason every number here travels with the run that produced it. See the run.
The author is explicit that the enriched criteria were informed by labelled errors, so part of that gain is supervision a cold-start deployment would not have.
The trade-off, on one dataset where both were measured.
Scoped to a single run deliberately. A cost–accuracy plane implies the points are alternatives to one another, and that is only true when they were measured on the same data under the same protocol.
Hover or tap a point for its full record.
Verified runs. No marketing claims.
| Jev jev-1.13.0 | 62.6% |
| Claude Haiku 4.5 | 81.3% |
| Jev | 83.2% |
| GPT-5.4 nano gpt-5.4-nano-2026-03-17 | 79.3% |
| GPT-5.6 Terra | 87.5% |
| BGE-small-en-v1.5 + logistic regression | 93.3% |
| Jev | 87.0% |
| GPT-5.4 nano gpt-5.4-nano-2026-03-17 | 79.5% |
| GPT-5.6 Terra | 91.5% |
| Jev jev-latest | 91.7% |
| Jev jev-preview | 91.7% |
| Jev jev-1.13.0, text-only | 93.6% |
| Jev jev-1.13.0, enriched evidence | 98.0% |
| Jev jev-1.13.0, enriched + evidence wording | 98.6% |
| TF-IDF + logistic regression enriched | 98.9% |
| Jev + TF-IDF ensemble | 99.3% |
| Jev jev-1.13.0, enriched evidence | 95.4% |
| TF-IDF + logistic regression enriched | 96.3% |
| Jev → DeepSeek v4 Pro cascade | 80.2% |
- 09.19Jev pricing verified at OpenRouter
- 09.19Claude Haiku 4.5 pricing verified at Anthropic (first-party API)
- 09.19Source verified · OpenRouter
- 09.19Source verified · Anthropic
- 09.19Source verified · Latent Space (AINews)
- 09.19Source verified · anisselbd
- 09.19Source verified · ickma2311
- 09.19Source verified · themsquared
How much quality can you recover before the saving is gone?
A Jev call on every request, escalating to Claude Haiku 4.5 on the ones it is not confident about. The break-even fallback rate is where that cascade stops being cheaper than running the baseline on everything.
Model your own workload →Jev can fall back on up to 99.8% of requests before this hybrid path costs more than baseline-only.
1M requests/month · Jev 300 input tokens · Claude Haiku 4.5 500 input + 910 output · verified list pricing both sides