Research Article

Verified-source Evidence for LLM-based Smishing Detection: A Reproducible Extension and Evaluation of SmishX

University of Toronto, Canada

Article Information

Article Type: Research Article
Submitted: August 27, 2026
Accepted: September 06, 2026
Published: September 15, 2026
Pages: 1-11
DOI: Pending
Language: English
License: CC BY 4.0

Abstract

SmishX [6] showed that pairing an LLM with link- and brand-context retrieval improves explainable SMS-phishing detection, but its released detector treats brand verification as a raw search-engine query and has no mechanism for checking what an impersonated organization actually says about the request it is being accused of making. We faithfully reproduce SmishX's detection pipeline (P0) and evaluate two extensions on a 283-message cohort (34 legitimate, 133 smishing, 116 spam): a corrected evidence packet that fixes how failed link collection is presented to the model, and an added retrieval layer for first-party and known-campaign evidence (P1). On the 267 cases scored by all three adopted systems, P1 dominates P0 on all six metrics measured (recall 96.6%→99.1%, specificity 67.6%→73.5%, MCC 0.668→0.804), and the head-to-head comparison narrowly misses conventional significance (McNemar exact, matched n=267, p=0.057). An ablation isolating the packet fix alone accounts for most of the gain and is significant (p=0.006); the added evidentiary retrieval's marginal contribution over that corrected baseline is not (p=0.69). We report this as an honest, partially-negative result rather than a validated win, disclose a deviation from a predeclared analysis protocol, flag that the headline cohort blends two differently-scoped sampling waves, and note that a predeclared blinded audit found neither system's stated reasoning fully evidence-supported. This draft also corrects a significance value mislabeled in an earlier version of this paper (Appendix D).

Keywords

Smishing Detection SMS Phishing Large Language Models AI Agents Cybersecurity Phishing Detection Spam Detection Fraud Detection Information Retrieval Source Verification SmishX Link Context Retrieval Brand Context Retrieval Evidence-based Detection

Cite

Citation: David Druker (2026) Verified-source Evidence for LLM-based Smishing Detection: A Reproducible Extension and Evaluation of SmishX. Epistora J. Artif. Intell. & Intell. Syst. 1(1), 1-11. Article EJAIS-108