← All evidence

Recent phishing messages: recall under distribution shift

AIndependent · Reproduction artifacts availableSecurity detectionRecent phishing set (author's own assembly) · n = 853date not stated by source

Detect phishing in messages drawn from 2024–25, newer than the main training and evaluation material.

Accuracy
Every recorded metric
ModelAccuracyRecallFPRAUROCECELatencyCost
Jev
95.3%not recorded
Protocol

Recall-only evaluation on recent phishing messages, using the enriched configuration from the main run. Recorded separately because it is a different dataset assembled to test behaviour on newer material.

Dataset
Recent phishing set (author's own assembly)
n = 853
Phishing messages from 2024–25
Limitations (4)
  • The source does not state a run date.
  • No baseline figure is recorded here. Our reading of the source returned a TF-IDF recall on this set identical to its main-set phishing recall, which we could not disambiguate, so the comparison figure is queued for re-verification rather than published.
  • Recall only: precision on this set is not reported, so a model that flagged everything as phishing would score well here.
  • n=853 with no confidence interval.
Note on this run
Queued for completion: the baseline comparison on this set is the interesting part and is deliberately absent until it is confirmed against the source.

What this benchmark does not prove

  • 01That Jev holds up under distribution shift. Recall without precision cannot establish that, and there is no baseline recorded here to compare against.
  • 02That a 95.3% recall on recent phishing generalises to your mail stream. This is one author's assembly of recent messages, not a representative sample.
  • 03Anything about how the classical baseline behaves on newer material, which is the comparison this run was assembled to make.
Source and artifacts

jev-spam-eval — three-way email classification: Jev variants vs TF-IDF logistic regression bitnovus, verified 2026-09-19

Reproduction artifacts: github.com/bitnovus/jev-spam-eval

Independent: yes · Vendor-affiliated: no · Independently reproduced: not yet

Public code and stated protocol. The repository does not state a run date, the author's affiliation, or who funded API costs; it does report measured spend of about $1.35 for the full paired run.