← All evidence
Agent tool-call risk classification
AIndependent · Reproduction artifacts availableAgent tool-call riskHand-labelled tool-call set (author's own) · n = 602026-09-17
Classify an agent tool call as readonly, destructive, privileged or exfiltration.
Accuracy
91.7%
Jev jev-latest
no interval reported
91.7%
Jev jev-preview
no interval reported
Every recorded metric
| Model | Accuracy | Recall | FPR | AUROC | ECE | Latency | Cost |
|---|---|---|---|---|---|---|---|
Jev | 91.7% | — | — | — | — | 422 ms / 542 ms location not stated | $0.000017 / call computed at $0.042/MTok list price |
Jev | 91.7% | — | — | — | — | 379 ms / 484 ms location not stated | not recorded |
Protocol
Four-way classification over a hand-labelled set, comparing two Jev variants against each other. No non-Jev baseline was run.
Dataset
Hand-labelled tool-call set (author's own)
n = 60
34 clear, 14 ambiguous, 12 adversarial
Limitations (5)
- n=60. At this size a single item moves accuracy by 1.7 points and the confidence interval is far wider than the differences being discussed.
- No non-Jev baseline: the author states no provider key was available at run time. This run shows how Jev behaves, not whether it beats an alternative.
- The dataset is the author's own hand-labelled set, not a public benchmark, so label quality cannot be independently checked.
- Accuracy is identical (55/60) for both variants, so the latency difference is the only separation between them.
- The source does not state where latency was measured from, so the network floor inside these figures is unknown and they cannot be compared against latency from any other run.
Note on this run
Calibration finding relevant to confidence routing: every incorrect answer came with hedged confidence, the model never returned 1.000 and was wrong, and confidence on the five misses ranged from 0.130 to 0.785. On this task set the confidence score looks usable as an escalation trigger. Note that the CLINC150 run points the other way on calibration quality, so this should not be generalised.
What this benchmark does not prove
- 01That Jev is better than anything at tool-call risk classification. No non-Jev baseline was run, so this measures Jev's behaviour, not its standing.
- 02That 91.7% is a reliable estimate. At n=60 a single item moves accuracy by 1.7 points and the interval is far wider than any difference discussed.
- 03That Jev's confidence score is trustworthy in general. The calibration finding holds on this hand-labelled set; CLINC150 points the other way.
- 04That jev-latest and jev-preview differ in quality. They scored identically at 55/60; only latency separated them.
Source and artifacts
jev-benchmark — Jev on agent tool-call risk classification — themsquared, verified 2026-09-19
Reproduction artifacts: github.com/themsquared/jev-benchmark
Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet
Compares two Jev variants against each other. Contains no non-Jev baseline; author states no provider key was available at run time.