Confidence routing on Web of Science: do not route
The same confidence-routing pipeline, applied to a second labelled dataset.
| Model | Accuracy | Recall | FPR | AUROC | ECE | Latency | Cost |
|---|---|---|---|---|---|---|---|
Jev → DeepSeek v4 Pro cascade | — | — | — | — | — | — | not recorded |
The same routing pipeline as the Banking77 sample, with no routing parameter carried over from it. The author's stated outcome is that no threshold beat the better single model on this dataset.
- The source does not state a run date.
- No numbers at all are published for this dataset — only the verdict. We record the verdict and deliberately leave the metric fields empty rather than reconstructing anything.
- Without the underlying figures there is no way to tell whether routing was close to worthwhile or far from it.
- No confidence intervals, no calibration metrics, no baseline figures.
- The author's affiliation and who paid API costs are not stated.
What this benchmark does not prove
- 01That confidence routing never works on this kind of data. It failed on one dataset with one fallback model and one parameter search.
- 02That Jev's confidence is uninformative here. No calibration metric was published for this dataset at all.
- 03Anything quantitative. This run contributes a verdict and nothing else, which is why its metric fields are empty.
Janus — confidence routing between Jev and a larger fallback model — FirasSX914, verified 2026-09-19
Reproduction artifacts: github.com/FirasSX914/Janus
Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet
A routing library rather than a benchmark. Its measured results are the tool's own demo output on two public datasets, so the tool has an interest in routing working — it reported that it does not on one of the two, which cuts against that interest. No run date, calibration metrics or intervals are stated.