Can Jev tell an ad from everything else?
200 short English texts — half SMS messages, half YouTube comments, all from two public spam corpora — sent to jev-1.13.0 as one binary question: is this primarily advertising? The labels it was scored against were written by AI annotators, not people, and the original human labels from the source corpora are published beside them so you can check the difference.
- TEXTS
- 200
- ACCURACY (AI LABELS)
- 94.0%
- ERRORS SIDING WITH HUMANS
- 5 / 12
- PRECISION IF ADS WERE 5%
- ~38.1%
| METRIC | VALUE | 95% INTERVAL (WILSON) | COUNT |
|---|---|---|---|
| Accuracy | 94.0% | 89.8%–96.5% | 188 / 200 |
| AD precision | 90.7% | 83.3%–95.0% | 88 / 97 |
| AD recall | 96.7% | 90.8%–98.9% | 88 / 91 |
| False-positive rate | 8.3% | 4.4%–15.0% | 9 / 109 |
| Client p50 / p95 latency | 479.5 / 1084 ms | — | region not recorded |
| Total cost, all texts | $0.004929 | — | list price, not an invoice |
No baseline model was run on these texts, so nothing here says whether Jev is the right choice for this task — only how it did against these labels.
For a two-option Choice, Jev’s confidence is the gap between the two option probabilities. Every one of the 12 errors came back with a small gap. The cost shows up too: a threshold that catches all of them also sends some correct answers to the fallback.
| THRESHOLD | KEPT | COVERAGE | WRONG KEPT | ACCURACY ON KEPT |
|---|---|---|---|---|
| 0.50 | 179 | 89.5% | 3 | 98.3% |
| 0.55 | 178 | 89.0% | 2 | 98.9% |
| 0.60 | 175 | 87.5% | 1 | 99.4% |
| 0.65 | 175 | 87.5% | 1 | 99.4% |
| 0.70 | 173 | 86.5% | 0 | 100.0% |
| 0.75 | 171 | 85.5% | 0 | 100.0% |
| 0.80 | 168 | 84.0% | 0 | 100.0% |
| 0.85 | 164 | 82.0% | 0 | 100.0% |
| 0.90 | 160 | 80.0% | 0 | 100.0% |
| 0.95 | 155 | 77.5% | 0 | 100.0% |
The thresholds were fixed before the run but scored on these same 200 texts, with no holdout. No fallback model was run, so accuracy on the texts routed away is unmeasured. Zero observed errors is not zero errors: at 0.70 the 95% interval still reaches down to 97.8%.
45.5% of these texts are ads — far more than most real feeds. Precision depends on that share. Holding Jev’s measured recall and false-positive rate fixed, the same classifier on a feed that is 5% ads would be right about roughly 38.1% of the texts it flags, and the false-positive rate’s own uncertainty puts that anywhere from about 25% to 54%.
Each row shows the AI-adjudicated label Jev was scored against and the source corpus’s original human label. Rows marked sides with human are the ones where the disagreement is arguably about the label rather than the model. Texts are shown exactly as published, as inert text — some are spam and contain links; none are clickable.
| CASE | TEXT | SOURCE | AI LABEL | HUMAN LABEL | JEV | CONF | |
|---|---|---|---|---|---|---|---|
| 0023 | yeah sure thing mate haunt got all my stuff sorted but im going sound anyway promoting hex for .by the way who is this? dont know number. Joke | SMS | NOT_AD | ham | AD | 0.29 | |
| 0040 | Pathaya enketa maraikara pa' | SMS | NOT_AD | ham | AD | 0.02 | |
| 0051 | http://www.bubblews.com/news/6401116-vps-solutions | YouTube comment | NOT_AD | spam | AD | 0.13 | sides with human |
| 0097 | A Boy loved a gal. He propsd bt she didnt mind. He gv lv lttrs, Bt her frnds threw thm. Again d boy decided 2 aproach d gal , dt time a truck was speeding towards d gal. Wn it was about 2 hit d girl,d boy ran like hell n saved her. She asked 'hw cn u run so fast?' D boy replied "Boost is d secret of my energy" n instantly d girl shouted "our energy" n Thy lived happily 2gthr drinking boost evrydy Moral of d story:- I hv free msgs:D;): gud ni8 | SMS | NOT_AD | ham | AD | 0.25 | |
| 0130 | LIFE has never been this much fun and great until you came in. You made it truly special for me. I won't forget you! enjoy @ one gbp/sms | SMS | AD | spam | NOT_AD | 0.06 | |
| 0160 | More views than nikki minaj Anaconda | YouTube comment | NOT_AD | ham | AD | 0.07 | |
| 0163 | lets get it to 1 BILLION | YouTube comment | AD | ham | NOT_AD | 0.12 | sides with human |
| 0165 | 1,000,000 VIEWS NEAR | YouTube comment | NOT_AD | ham | AD | 0.69 | |
| 0182 | Hi babe its Chloe, how r u? I was smashed on saturday night, it was great! How was your weekend? U been missing me? SP visionsms.com Text stop to stop 150p/text | SMS | AD | spam | NOT_AD | 0.03 | |
| 0191 | Katy has the voice of gold. this video really brings the lyrics to life. I am a rapper with 25000 subscribers.. thumbs up if you hate when you take a shit and the water splashes your nuts | YouTube comment | NOT_AD | spam | AD | 0.57 | sides with human |
| 0194 | Thumbs up if FE-FE-FE-FE-FEGELEIN brought u here | YouTube comment | NOT_AD | spam | AD | 0.34 | sides with human |
| 0197 | this song is awesome. these guys are the best. love this video too its hilarious lol. im getting popular fast because i rap with meaning. thumbs up if you piss next to the water in the toilet so it doesnt make noise | YouTube comment | NOT_AD | spam | AD | 0.52 | sides with human |
Across all 200 texts, the AI and human labels disagree on 11. Human labels are mapped spam → AD and ham → NOT_AD, which is a proxy: a subscribe-to-my-channel comment is spam without obviously being an ad.
Texts were sampled from 8 strata chosen before labelling. These are those strata. 70 of 200 texts were later re-filed by the adjudicator; the artifact keeps that grouping too, but it ends up one label per group, which makes it uninformative.
| DESIGN SLICE | N | AD / NOT_AD | ACCURACY | AD RECALL | FALSE-POSITIVE RATE |
|---|---|---|---|---|---|
| soft_ad | 25 | 21 / 4 | 88.0% | 100.0% | 75.0% |
| product_boundary | 10 | 0 / 10 | 90.0% | n/a | 10.0% |
| social_post | 35 | 1 / 34 | 91.4% | 0.0% | 5.9% |
| product_offer | 15 | 12 / 3 | 93.3% | 91.7% | 0.0% |
| normal_content | 40 | 0 / 40 | 95.0% | n/a | 5.0% |
| weak_cta | 25 | 22 / 3 | 96.0% | 100.0% | 33.3% |
| obvious_ad | 35 | 35 / 0 | 97.1% | 97.1% | n/a |
| other_boundary | 15 | 0 / 15 | 100.0% | n/a | 0.0% |
n/a where a slice holds no texts of that class. Groups are small; a single text moves a slice’s accuracy by several points.
Deterministic 200-text subset from UCI SMS Spam Collection and YouTube Spam Collection. Two separate source-blind AI annotators assigned AD/NOT_AD under a written guideline; a third AI adjudicated disagreements. The frozen prompt asks one binary Choice per text through TypeSafe's live API. A 20-text schema smoke run preceded the full 200-text run. No baseline or independent holdout was evaluated.
- The labels were assigned and adjudicated by AI agents, not by human annotators; shared model biases may raise apparent agreement and accuracy.
- On 11 of 200 texts the AI label disagrees with the source corpus's original human spam/ham label, and in 5 of Jev's 12 counted errors Jev agrees with the human label. Accuracy ranges from 93.5% to 96.5% depending on which labels are trusted.
- The sample is curated from two existing spam corpora and has 45.5% AD prevalence, which is not a deployment prevalence estimate.
- The source corpora were published in 2011 (SMS) and 2015 (YouTube); no text reflects current advertising formats.
- The 0.50–0.95 confidence sweep uses the same 200 examples as the reported accuracy and has no independent holdout validation.
- Latency is client-observed round-trip time; the runner region and network floor were not recorded.
- Cost is calculated from recorded input tokens and published list price, not reconciled to a provider invoice.
- It does not establish accuracy against human-annotated ground truth or on text outside these two UCI source corpora. Where the original human labels exist, several of the counted errors side with them against the AI labels.
- It does not establish deployment precision, because precision depends on the actual advertising prevalence and any dataset shift.
- It does not prove that a confidence threshold chosen from these 200 examples will preserve the same coverage or accuracy on new data.
- It is not an independent result: JevBench designed the task, selected the texts, generated and adjudicated the labels with AI, ran Jev, and wrote the analysis. Every step can introduce first-party bias and cannot corroborate JevBench's own claim.
Source texts are redistributed under CC BY 4.0 from two UCI Machine Learning Repository datasets:
- Almeida, T. and Hidalgo, J. (2011). SMS Spam Collection. doi:10.24432/C5CC84
- Alberto, T. and Lochter, J. (2015). YouTube Spam Collection. doi:10.24432/C58885
JevBench normalised whitespace, selected the subset, replaced the spam/ham labels with AI-adjudicated AD/NOT_AD labels while keeping the originals, and added Jev’s outputs. Every figure on this page is read from the published artifact directory, which also carries the full per-text results, both annotators’ labels, the protocol, and SHA-256 checksums for every file.
The texts are untrusted user content from spam corpora, including URLs and phone numbers. They are shown here as plain text and should be treated as data.