Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study
Abstract
1. Introduction
2. Review of the Related Literature
2.1. Evolution of Phishing Techniques
2.2. AI-Enabled Phishing
3. Research Questions
- RQ1.
- In this pilot study, how did participants rate the convincingness of GPT-4-generated and human-authored personalized smishing messages?
- RQ2.
- What content characteristics contribute to a more convincing spear phishing message?
- RQ3.
- How accurately did participants identify whether the study messages were GPT-4-generated or human-authored?
- RQ4.
- What criteria did participants describe when judging whether the study messages were GPT-4-generated or human-authored?
- RQ5.
- Can the GPT-4-generated and human-authored messages in this study be distinguished computationally from their text?
4. Methodology
4.1. TRAPD Framework and Study Procedure
- Recruit targets who willingly share personal information with potential attackers.
- Generate personalized deceptive messages aimed at the targets, for example using humans or AI.
- Have targets rank the messages from most to least convincing and choose a threshold above which they would intend to click.
- Have targets explain why they placed messages where they did.
- Optionally, have targets label messages with a variable of interest, such as perceived AI authorship, and explain their choices.
4.2. Recruiting Targets Who Shared Personal Information
4.3. Generating Personalized Deceptive Messages
4.3.1. Human Generation
4.3.2. GPT-4 Generation
4.4. Target Interview and Sorting Activity
4.4.1. Threshold Rank Order
4.4.2. Qualitative Assessment
4.4.3. Source Labeling
4.5. Convincingness Analytic Track
4.6. Source Attribution Analytic Track
4.6.1. Computational Source-Attribution Analysis Plan
4.6.2. PCA and Exploratory Clustering
4.7. Implementation and Reproducibility Details
5. Convincingness Results
5.1. Quantitative Convincingness Results
5.1.1. GPT-4 vs. Human Ranking and Intended Click Probability
5.1.2. Topic-Based Ranking and Intended Click Probability
5.1.3. Demographics and Phishing Experience
5.2. Content Characteristics of Convincing Messages
5.2.1. Personal Relevance
5.2.2. Sender Identity
5.2.3. URLs and Perceived Convincingness
5.2.4. Technology Communication Medium
5.2.5. Messaging Style
5.2.6. Urgency and Scarcity
5.2.7. Context Inaccuracies
5.2.8. Plausible Rewards
6. Source Attribution Results
6.1. Human Source Attribution
Identifying Message Origin
6.2. Human Criteria for Judging Message Source
6.2.1. Style
6.2.2. Personalization
6.2.3. Word Choice
6.2.4. Message Structure
6.2.5. Grammar/Spelling
6.2.6. Emojis
6.2.7. Message Length
6.2.8. URLs as Source-Attribution Cues
6.3. Computational Source Attribution
Descriptive PCA and Exploratory Clustering
7. Discussion
7.1. Convincingness of GPT-4- and Human-Authored Messages
7.2. Source Attribution by Participants and Models
7.3. The Proposed TRAPD Evaluation Framework
7.4. Implications of AI-Enabled Spear Phishing
7.5. Dual-Use Considerations
7.6. Mitigation and Countermeasure Implications
8. Limitations
8.1. Measurement and TRAPD Validity
8.2. Sample and Statistical Precision
8.3. Human Comparison Condition
8.4. GPT-4 Generation Reproducibility
8.5. Qualitative Coding
8.6. Interview Sequence and Response Effects
8.7. Computational Scope and Data Availability
9. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
DURC Statement
Abbreviations
| AI | Artificial intelligence |
| LLM | Large language model |
| SMS | Short message service |
| TRAPD | Threshold Ranking Approach for Personalized Deception |
References
- Alsharida, R.A.; Al-rimy, B.A.S.; Al-Emran, M.; Zainal, A. A systematic review of multi perspectives on human cybersecurity behavior. Technol. Soc. 2023, 73, 102258. [Google Scholar] [CrossRef] [Scilit]
- Yeboah-Boateng, E.O.; Amanor, P.M. Phishing, SMiShing & Vishing: An Assessment of Threats against Mobile Devices. J. Emerg. Trends Comput. Inf. Sci. 2014, 5, 297–307. [Google Scholar] [CrossRef] [Scilit]
- Federal Bureau of Investigation Internet Crime Complaint Center. 2025 IC3 Annual Report. 2026. Available online: https://www.ic3.gov/AnnualReport/Reports/2025_IC3Report.pdf (accessed on 30 September 2024).
- Dewan, P.; Kashyap, A.; Kumaraguru, P. Analyzing social and stylometric features to identify spear phishing emails. In Proceedings of the 2014 APWG Symposium on Electronic Crime Research (eCrime), Birmingham, AL, USA, 23–25 September 2014; pp. 1–13. [Google Scholar] [CrossRef] [Scilit]
- Mohamed, N.; Taherdoost, H.; Madanchian, M. Enhancing Spear Phishing Defense with AI: A Comprehensive Review and Future Directions. Eai Endorsed Trans. Scalable Inf. Syst. 2024, 12, 1–10. [Google Scholar] [CrossRef] [Scilit]
- Schmitt, M.; Flechais, I. Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing. Artif. Intell. Rev. 2024, 57, 324. [Google Scholar] [CrossRef] [Scilit]
- Proofpoint, Inc. 2024 State of the Phish—Today’s Cyber Threats and Phishing Protection. Available online: https://www.proofpoint.com/us/resources/threat-reports/state-of-phish (accessed on 30 September 2024).
- Benenson, Z.; Gassmann, F.; Landwirth, R. Unpacking Spear Phishing Susceptibility. In Proceedings of the Financial Cryptography and Data Security; Brenner, M., Rohloff, K., Bonneau, J., Miller, A., Ryan, P.Y., Teague, V., Bracciali, A., Sala, M., Pintore, F., Jakobsson, M., Eds.; Springer: Cham, Switzerland, 2017; pp. 610–627. [Google Scholar] [CrossRef] [Scilit]
- Rajivan, P.; Gonzalez, C. Creative Persuasion: A Study on Adversarial Behaviors and Strategies in Phishing Attacks. Front. Psychol. 2018, 9, 135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Heiding, F.; Lermen, S.; Kao, A.; Schneier, B.; Vishwanath, A. Evaluating Large Language Models’ Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects. arXiv 2024, arXiv:2412.00586. [Google Scholar] [CrossRef] [Scilit]
- Roy, S.S.; Thota, P.; Naragam, K.V.; Nilizadeh, S. From Chatbots to PhishBots? Preventing Phishing Scams Created Using ChatGPT, Google Bard and Claude. arXiv 2024, arXiv:2310.19181. [Google Scholar] [CrossRef] [Scilit]
- Oliveira, D.S.; Lin, T.; Rocha, H.; Ellis, D.; Dommaraju, S.; Yang, H.; Weir, D.; Marin, S.; Ebner, N.C. Empirical analysis of weapons of influence, life domains, and demographic-targeting in modern spam: An age-comparative perspective. Crime Sci. 2019, 8, 3. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, T.; Capecci, D.E.; Ellis, D.M.; Rocha, H.A.; Dommaraju, S.; Oliveira, D.S.; Ebner, N.C. Susceptibility to Spear-Phishing Emails: Effects of Internet User Demographics and Email Content. ACM Trans. Comput.-Hum. Interact. 2019, 26, 32:1–32:28. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hazell, J. Large Language Models Can Be Used To Effectively Scale Spear Phishing Campaigns. arXiv 2023, arXiv:2305.06972. [Google Scholar]
- Seymour, J.; Tully, P. Generative Models for Spear Phishing Posts on Social Media. arXiv 2018, arXiv:1802.05196. [Google Scholar] [CrossRef] [Scilit]
- Althobaiti, K.; Alsufyani, N. A Review of Organization-Oriented Phishing Research. PeerJ Comput. Sci. 2024, 10, e2487. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Barrera, D.; Naranjo, V.; Fuertes, W.; Macas, M. Literature Review of SMS Phishing Attacks: Lessons, Addresses, and Future Challenges. In Proceedings of the Advanced Research in Technologies, Information, Innovation and Sustainability; Communications in Computer and Information Science; Springer: Cham, Switzerland, 2024; Volume 1936, pp. 191–204. [Google Scholar] [CrossRef] [Scilit]
- Anti-Phishing Working Group. Phishing Activity Trends Report: 4th Quarter 2025; Anti-Phishing Working Group: Lexington, MA, USA, 2026. [Google Scholar]
- Karamagi, R. A Review of Factors Affecting the Effectiveness of Phishing. Comput. Inf. Sci. 2022, 15, 20–31. [Google Scholar] [CrossRef] [Scilit]
- Bethany, M.; Galiopoulos, A.; Bethany, E.; Karkevandi, M.B.; Beebe, N.; Vishwamitra, N.; Najafirad, P. Lateral Phishing with Large Language Models: A Large Organization Comparative Study. IEEE Access 2025, 13, 60684–60701. [Google Scholar] [CrossRef] [Scilit]
- Verizon. 2025 Data Breach Investigations Report; Verizon: New York, NY, USA, 2025. [Google Scholar]
- Mishra, S.; Soni, D. SMS Phishing and Mitigation Approaches. In Proceedings of the 2019 Twelfth International Conference on Contemporary Computing (IC3), Noida, India, 8–10 August 2019; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
- Bossetta, M. The Weaponization of Social Media: Spear Phishing and Cyberattacks on Democracy. J. Int. Aff. 2018, 71, 97–106. [Google Scholar]
- Blancaflor, E.B.; Cruz, K.J.R.; Monta, F.J.C.; Flores, P.E. Unmasking the Threat: Analyzing and Mitigating SMS Smishing Attacks. In Proceedings of the Ninth International Congress on Information and Communication Technology; Lecture Notes in Networks and Systems; Springer: Singapore, 2024; Volume 1055, pp. 69–78. [Google Scholar] [CrossRef] [Scilit]
- Parker, H.J.; Flowerday, S.V. Contributing factors to increased susceptibility to social media phishing attacks. SA J. Inf. Manag. 2020, 22, a1176. [Google Scholar] [CrossRef] [Scilit]
- Zhai, X.; Nyaaba, M.; Ma, W. Can Generative AI and ChatGPT Outperform Humans on Cognitive-demanding Problem-Solving Tasks in Science? Sci. Educ. 2025, 34, 649–670. [Google Scholar] [CrossRef] [Scilit]
- Khan, H.; Alam, M.; Al-Kuwari, S.; Faheem, Y. Offensive AI: Unification of Email Generation Through GPT-2 Model with a Game-Theoretic Approach for Spear-Phishing Attacks. In Proceedings of the Competitive Advantage in the Digital Economy (CADE 2021), Online, 2–3 June 2021; pp. 178–184. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Y.; Sadasivan, V.S.; Saberi, M.; Saha, S.; Feizi, S. Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text. In Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego, CA, USA, 2–7 December 2025. [Google Scholar]
- Wilczyński, P.; Mieleszczenko-Kowszewicz, W.; Biecek, P. Resistance Against Manipulative AI: Key factors and possible actions. arXiv 2024, arXiv:2404.14230. [Google Scholar] [CrossRef] [Scilit]
- Hoeken, H.; Fikkers, K.; Eerland, A.; Holleman, B.; van Berkum, J.; Pander Maat, H. The Perceived Convincingness Model: Why and under what conditions processing fluency and emotions are valid indicators of a message’s perceived convincingness. Commun. Theory 2022, 32, 488–496. [Google Scholar] [CrossRef] [Scilit]
- Moody, G.D.; Galletta, D.F.; Dunn, B.K. Which Phish Get Caught? An Exploratory Study of Individuals’ Susceptibility to Phishing. Eur. J. Inf. Syst. 2017, 26, 564–584. [Google Scholar] [CrossRef] [Scilit]
- Zhuo, S.; Biddle, R.; Koh, Y.S.; Lottridge, D.; Russello, G. SoK: Human-centered Phishing Susceptibility. ACM Trans. Priv. Secur. 2023, 26, 1–27. [Google Scholar] [CrossRef] [Scilit]
- Hakim, Z.M.; Ebner, N.C.; Oliveira, D.S.; Getz, S.J.; Levin, B.E.; Lin, T.; Lloyd, K.; Lai, V.T.; Grilli, M.D.; Wilson, R.C. The Phishing Email Suspicion Test (PEST) a lab-based task for evaluating the cognitive mechanisms of phishing detection. Behav. Res. Methods 2021, 53, 1342–1352. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, T.; Singh, K.; Rajivan, P. Personalized persuasion: Quantifying susceptibility to information exploitation in spear-phishing attacks. Appl. Ergon. 2023, 108, 103908. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hanus, B.; Wu, Y.A.; Parrish, J. Phish Me, Phish Me Not. J. Comput. Inf. Syst. 2022, 62, 516–526. [Google Scholar] [CrossRef] [Scilit]
- Burns, A.J.; Johnson, M.E.; Caputo, D.D. Spear Phishing in a Barrel: Insights from a Targeted Phishing Campaign. J. Organ. Comput. Electron. Commer. 2019, 29, 24–39. [Google Scholar] [CrossRef] [Scilit]
- Williams, E.J.; Hinds, J.; Joinson, A.N. Exploring susceptibility to phishing in the workplace. Int. J. Hum.-Comput. Stud. 2018, 120, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Kahlke, R.; Maggio, L.A.; Lee, M.C.; Cristancho, S.; LaDonna, K.A.; Abdallah, Z.; Khehra, A.; Kshatri, K.; Horsley, T.; Varpio, L. When Words Fail Us: An Integrative Review of Innovative Elicitation Techniques for Qualitative Interviews. Med. Educ. 2025, 59, 382–394. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Y.; Gosline, R. Human favoritism, not AI aversion: People’s perceptions (and bias) toward generative AI, human experts, and human–GAI collaboration in persuasive content generation. Judgm. Decis. Mak. 2023, 18, e41. [Google Scholar] [CrossRef] [Scilit]
- Nisbett, N.; Spaiser, V. How convincing are AI-generated moral arguments for climate action? Front. Clim. 2023, 5, 1193350. [Google Scholar] [CrossRef] [Scilit]
- Palmer, A.; Spirling, A. Large Language Models Can Argue in Convincing Ways About Politics, But Humans Dislike AI Authors: Implications for Governance. Political Sci. 2023, 75, 281–291. [Google Scholar] [CrossRef] [Scilit]
- Heiding, F.; Schneier, B.; Vishwanath, A.; Bernstein, J.; Park, P.S. Devising and Detecting Phishing: Large Language Models vs. Smaller Human Models. arXiv 2023, arXiv:2308.12287. [Google Scholar] [CrossRef] [Scilit]
- Butavicius, M.; Parsons, K.; Pattinson, M.; McCormac, A. Breaching the Human Firewall: Social engineering in Phishing and Spear-Phishing Emails. arXiv 2016, arXiv:1606.00887. [Google Scholar] [CrossRef] [Scilit]
- Stembert, N.; Padmos, A.; Bargh, M.S.; Choenni, S.; Jansen, F. A Study of Preventing Email (Spear) Phishing by Enabling Human Intelligence. In Proceedings of the 2015 European Intelligence and Security Informatics Conference, Manchester, UK, 7–9 September 2015; pp. 113–120. [Google Scholar] [CrossRef] [Scilit]
- Köbis, N.; Mossink, L. Artificial Intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry. arXiv 2020, arXiv:2005.09980. [Google Scholar] [CrossRef] [Scilit]
- Jakesch, M.; Hancock, J.; Naaman, M. Human heuristics for AI-generated language are flawed. Proc. Natl. Acad. Sci. USA 2023, 120, e2208839120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thomopoulos, G.; Lyras, D.; Fidas, C. Methodologies and Ethical Considerations in Phishing Research: A Comprehensive Review. In Proceedings of the CHIGREECE ’23: 2nd International Conference of the ACM Greek SIGCHI Chapter, Athens, Greece, 27–28 September 2023; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Distler, V. The Influence of Context on Response to Spear-Phishing Attacks: An In-Situ Deception Study. In Proceedings of the CHI ’23: 2023 CHI Conference on Human Factors in Computing Systems, Hamburg, Germany, 23–28 April 2023; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Pan, Y.; Su, Z.; Deng, Y.; Zhao, Q.; Du, L.; Luan, T.H.; Kang, J.; Niyato, D. Large Model-Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends. IEEE Commun. Surv. Tutor. 2026, 28, 1906–1949. [Google Scholar] [CrossRef] [Scilit]
- McKinsey & Company. The State of AI in Early 2024: Gen AI Adoption Spikes and Starts to Generate Value. Available online: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (accessed on 30 September 2024).
- Raja, A.K.; Zhou, J. AI Accountability: Approaches, Affecting Factors, and Challenges. Computer 2023, 56, 46–56. [Google Scholar] [CrossRef] [Scilit]
- Basit, A.; Zafar, M.; Liu, X.; Javed, A.R.; Jalil, Z.; Kifayat, K. A Comprehensive Survey of AI-Enabled Phishing Attacks Detection Techniques. Telecommun. Syst. 2021, 76, 139–154. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zieni, R.; Massari, L.; Calzarossa, M.C. Phishing or Not Phishing? A Survey on the Detection of Phishing Websites. IEEE Access 2023, 11, 18499–18519. [Google Scholar] [CrossRef] [Scilit]
- Dou, Z.; Khalil, I.; Khreishah, A.; Al-Fuqaha, A.; Guizani, M. Systematization of Knowledge (SoK): A Systematic Review of Software-Based Web Phishing Detection. IEEE Commun. Surv. Tutor. 2017, 19, 2797–2819. [Google Scholar] [CrossRef] [Scilit]








| Category | Avg. Rank | Intended Click Rate |
|---|---|---|
| Topic | ||
| Job | 5.71 | 38% |
| Hobby | 6.66 | 19% |
| Social | 7.13 | 17% |
| Source | ||
| Human-authored | 6.59 | 21.3% |
| GPT-4-generated | 6.41 | 28.0% |
| Theme | % of Participants |
|---|---|
| Relevance | 76% |
| Sender | 68% |
| URL | 64% |
| Medium | 40% |
| Style | 40% |
| Urgency and Scarcity | 32% |
| Inaccuracies | 28% |
| Rewards | 28% |
| Guessed GPT-4 | Guessed Human | Row Total | |
|---|---|---|---|
| True GPT-4 | 78 | 72 | 150 |
| True Human | 72 | 78 | 150 |
| Column Total | 150 | 150 | 300 |
| Correct Guesses | 156 of 300 (52%) | ||
| Theme | % of Participants |
|---|---|
| Uncertainty | 48% |
| Style | 40% |
| Personalization | 32% |
| Word Choice | 24% |
| Message Structure | 24% |
| Grammar | 24% |
| Emojis | 20% |
| Message Length | 16% |
| URL | 8% |
| Condition | ROC AUC | Avg. Precision | Bal. Acc. | Sensitivity | Specificity |
|---|---|---|---|---|---|
| No-link baseline | 0.977 | 0.976 | 0.920 | 0.933 | 0.907 |
| Jointly normalized | 0.961 | 0.960 | 0.910 | 0.913 | 0.907 |
| Jointly normalized and length-equalized | 0.954 | 0.948 | 0.887 | 0.900 | 0.873 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Francia, J.; Hansen, D.; Schooley, B.; Taylor, M.; Murray, S.V.; Cornelius, R.; Snow, G. Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study. J. Cybersecur. Priv. 2026, 6, 129. https://doi.org/10.3390/jcp6040129
Francia J, Hansen D, Schooley B, Taylor M, Murray SV, Cornelius R, Snow G. Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study. Journal of Cybersecurity and Privacy. 2026; 6(4):129. https://doi.org/10.3390/jcp6040129
Chicago/Turabian StyleFrancia, Jerson, Derek Hansen, Benjamin Schooley, Matthew Taylor, Shydra Valynn Murray, Rebekah Cornelius, and Greg Snow. 2026. "Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study" Journal of Cybersecurity and Privacy 6, no. 4: 129. https://doi.org/10.3390/jcp6040129
APA StyleFrancia, J., Hansen, D., Schooley, B., Taylor, M., Murray, S. V., Cornelius, R., & Snow, G. (2026). Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study. Journal of Cybersecurity and Privacy, 6(4), 129. https://doi.org/10.3390/jcp6040129

