Next Article in Journal
Laboratory Test and Constitutive Model for Quantifying the Anisotropic Swelling Behavior of Expansive Soils
Previous Article in Journal
Improved Adversarial Transfer Network for Bearing Fault Diagnosis under Variable Working Conditions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Beyond Word-Based Model Embeddings: Contextualized Representations for Enhanced Social Media Spam Detection

by
Sawsan Alshattnawi
1,†,
Amani Shatnawi
1,†,
Anas M.R. AlSobeh
1,2,† and
Aws A. Magableh
1,3,*,†
1
Faculty of Computer Science and Information Technology, Yarmouk University, Irbid 21163, Jordan
2
Information Technology, School of Computing, Southern Illinois University Carbondale, 1365 Douglas Drive, Carbondale, IL 62901, USA
3
Software Engineering, Computer and Information Sciences, Prince Sultan University, Riyadh 11586, Saudi Arabia
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Appl. Sci. 2024, 14(6), 2254; https://doi.org/10.3390/app14062254
Submission received: 31 January 2024 / Revised: 25 February 2024 / Accepted: 5 March 2024 / Published: 7 March 2024

Abstract

As social media platforms continue their exponential growth, so do the threats targeting their security. Detecting disguised spam messages poses an immense challenge owing to the constant evolution of tactics. This research investigates advanced artificial intelligence techniques to significantly enhance multiplatform spam classification on Twitter and YouTube. The deep neural networks we use are state-of-the-art. They are recurrent neural network architectures with long- and short-term memory cells that are powered by both static and contextualized word embeddings. Extensive comparative experiments precede rigorous hyperparameter tuning on the datasets. Results reveal a profound impact of tailored, platform-specific AI techniques in combating sophisticated and perpetually evolving threats. The key innovation lies in tailoring deep learning (DL) architectures to leverage both intrinsic platform contexts and extrinsic contextual embeddings for strengthened generalization. The results include consistent accuracy improvements of more than 10–15% in multisource datasets, unlocking actionable guidelines on optimal components of neural models, and embedding strategies for cross-platform defense systems. Contextualized embeddings like BERT and ELMo consistently outperform their noncontextualized counterparts. The standalone ELMo model with logistic regression emerges as the top performer, attaining exceptional accuracy scores of 90% on Twitter and 94% on YouTube data. This signifies the immense potential of contextualized language representations in capturing subtle semantic signals vital for identifying disguised spam. As emerging adversarial attacks exploit human vulnerabilities, advancing defense strategies through enhanced neural language understanding is imperative. We recommend that social media companies and academic researchers build on contextualized language models to strengthen social media security. This research approach demonstrates the immense potential of personalized, platform-specific DL techniques to combat the continuously evolving threats that threaten social media security.
Keywords: online social network (OSNs); social network analysis; cybersecurity; spam detection; neural word embeddings; RNN; contextual word embeddings; LSTM; BERT; EMLO online social network (OSNs); social network analysis; cybersecurity; spam detection; neural word embeddings; RNN; contextual word embeddings; LSTM; BERT; EMLO

Share and Cite

MDPI and ACS Style

Alshattnawi, S.; Shatnawi, A.; AlSobeh, A.M.R.; Magableh, A.A. Beyond Word-Based Model Embeddings: Contextualized Representations for Enhanced Social Media Spam Detection. Appl. Sci. 2024, 14, 2254. https://doi.org/10.3390/app14062254

AMA Style

Alshattnawi S, Shatnawi A, AlSobeh AMR, Magableh AA. Beyond Word-Based Model Embeddings: Contextualized Representations for Enhanced Social Media Spam Detection. Applied Sciences. 2024; 14(6):2254. https://doi.org/10.3390/app14062254

Chicago/Turabian Style

Alshattnawi, Sawsan, Amani Shatnawi, Anas M.R. AlSobeh, and Aws A. Magableh. 2024. "Beyond Word-Based Model Embeddings: Contextualized Representations for Enhanced Social Media Spam Detection" Applied Sciences 14, no. 6: 2254. https://doi.org/10.3390/app14062254

APA Style

Alshattnawi, S., Shatnawi, A., AlSobeh, A. M. R., & Magableh, A. A. (2024). Beyond Word-Based Model Embeddings: Contextualized Representations for Enhanced Social Media Spam Detection. Applied Sciences, 14(6), 2254. https://doi.org/10.3390/app14062254

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop