Recent phishing messages: recall under distribution shift
Detect phishing in messages drawn from 2024–25, newer than the main training and evaluation material.
| Model | Accuracy | Recall | FPR | AUROC | ECE | Latency | Cost |
|---|---|---|---|---|---|---|---|
Jev | — | 95.3% | — | — | — | — | not recorded |
Recall-only evaluation on recent phishing messages, using the enriched configuration from the main run. Recorded separately because it is a different dataset assembled to test behaviour on newer material.
- The source does not state a run date.
- No baseline figure is recorded here. Our reading of the source returned a TF-IDF recall on this set identical to its main-set phishing recall, which we could not disambiguate, so the comparison figure is queued for re-verification rather than published.
- Recall only: precision on this set is not reported, so a model that flagged everything as phishing would score well here.
- n=853 with no confidence interval.
What this benchmark does not prove
- 01That Jev holds up under distribution shift. Recall without precision cannot establish that, and there is no baseline recorded here to compare against.
- 02That a 95.3% recall on recent phishing generalises to your mail stream. This is one author's assembly of recent messages, not a representative sample.
- 03Anything about how the classical baseline behaves on newer material, which is the comparison this run was assembled to make.
jev-spam-eval — three-way email classification: Jev variants vs TF-IDF logistic regression — bitnovus, verified 2026-09-19
Reproduction artifacts: github.com/bitnovus/jev-spam-eval
Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet
Public code and stated protocol. The repository does not state a run date, the author's affiliation, or who funded API costs; it does report measured spend of about $1.35 for the full paired run.