← All evidence

Confidence routing on Web of Science: do not route

AIndependent · Reproduction artifacts availableIntent & routingrouting evaluationWeb of Science (routing sample) · n = 500date not stated by sourceverdict: DO NOT ROUTE

The same confidence-routing pipeline, applied to a second labelled dataset.

Accuracy
Every recorded metric
ModelAccuracyRecallFPRAUROCECELatencyCost
Jev → DeepSeek v4 Pro cascade
not recorded
Protocol

The same routing pipeline as the Banking77 sample, with no routing parameter carried over from it. The author's stated outcome is that no threshold beat the better single model on this dataset.

Dataset
Web of Science (routing sample)
n = 500
500 labelled rows
Limitations (5)
  • The source does not state a run date.
  • No numbers at all are published for this dataset — only the verdict. We record the verdict and deliberately leave the metric fields empty rather than reconstructing anything.
  • Without the underlying figures there is no way to tell whether routing was close to worthwhile or far from it.
  • No confidence intervals, no calibration metrics, no baseline figures.
  • The author's affiliation and who paid API costs are not stated.
Note on this run
A negative result, recorded because negative results are evidence. The same pipeline that found a workable threshold on Banking77 found none here.

What this benchmark does not prove

  • 01That confidence routing never works on this kind of data. It failed on one dataset with one fallback model and one parameter search.
  • 02That Jev's confidence is uninformative here. No calibration metric was published for this dataset at all.
  • 03Anything quantitative. This run contributes a verdict and nothing else, which is why its metric fields are empty.
Source and artifacts

Janus — confidence routing between Jev and a larger fallback model FirasSX914, verified 2026-09-19

Reproduction artifacts: github.com/FirasSX914/Janus

Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet

A routing library rather than a benchmark. Its measured results are the tool's own demo output on two public datasets, so the tool has an interest in routing working — it reported that it does not on one of the two, which cuts against that interest. No run date, calibration metrics or intervals are stated.