A Deep Prompt-Based Chain-of-Thought Approach to Harmful Euphemism Detection in Social Networks
Abstract
1. Introduction
- 1.
- Data Scarcity: There is a lack of large-scale, fine-grained datasets with contextual semantic annotations, which limits models’ semantic understanding and generalization capabilities.
- 2.
- Cognitive Reasoning Deficit in Lightweight Models: Existing methods largely rely on keywords or explicit features extracted by conventional pre-trained language models (PLMs). However, euphemistic expressions often depend heavily on contextual semantics and sociocultural backgrounds. Current lightweight models possess limited capacity for external knowledge integration and explicit logical deduction. This limitation significantly hinders detection accuracy and robustness when countering complex rhetorical disguises.
- 3.
- Latency Constraints of Large Language Models (LLMs): While modern LLMs possess strong zero-shot reasoning abilities capable of deciphering implicit intents, their large parameter sizes incur prohibitive computational overhead. This renders direct LLM deployment infeasible for the microsecond-level, high-concurrency real-time moderation demanded by modern social networks.
- •
- We construct a large-scale Chinese harmful euphemism dataset with fine-grained semantic annotations. The dataset is built from the Bilibili platform and covers multiple dimensions, including topic categories, attack types, euphemistic expressions, and their corresponding explanations. To the best of our knowledge, this is the first dataset to incorporate explicit semantic explanations of harmful euphemisms.
- •
- We propose a novel representation learning framework that integrates chain-of-thought (CoT) prompting with multi-head contrastive learning. By leveraging external knowledge from LLMs in an offline manner, the framework enhances the diversity and precision of semantic representations while avoiding the computational overhead of online inference.
- •
- We develop a multi-dimensional semantic-aware detection model with dynamic fusion. The proposed architecture captures implicit semantics and contextual dependencies more effectively. It achieves strong performance while maintaining low inference latency.
2. Related Work
2.1. Semantics-Driven Representation of Harmful Euphemisms
2.2. Contrastive Learning for Implicit Semantics
2.3. Enhancement Methods Fusing External Knowledge
2.4. Multi-Task Learning and Debiasing Research
- 1.
- Single-Dimensional Perception and Shallow Feature Fusion: Existing traditional and lightweight frameworks predominantly rely on single-dimensional text sequences. Even when external features or multi-task sharing mechanisms are introduced, they are typically integrated through rigid and shallow concatenation. When confronted with highly obscure rhetorical disguises, this structural limitation inherently causes the aforementioned cognitive reasoning deficit, inevitably leading to semantic collapse. Consequently, traditional models frequently generate spurious correlations and overfit to high-frequency surface words, rather than structurally decouple underlying syntactic, contextual, and rhetorical anomalies.
- 2.
- Deployment Dilemmas and Shallow Integration of LLMs: Recently, LLMs have shown strong zero-shot reasoning capabilities via CoT prompting. However, current LLM-augmented methods typically inject external knowledge by appending generated reasoning rationales to prompts during online inference. This shallow, text-level fusion fails to organically reshape the model’s latent cognitive space. Furthermore, relying on massive LLMs for direct online inference is highly susceptible to pre-training biases and noisy rationales, while also incurring the prohibitive computational overhead and latency mentioned earlier. This makes them impractical for real-time moderation in dynamic, high-throughput social platforms.
3. Methodology
3.1. Dataset Construction
3.2. Representation Learning
3.2.1. Seed Lexicon Matching Stage
3.2.2. Chain-of-Thought Template Generation Stage
3.2.3. Three-Fold Data Augmentation Stage
- (1)
- Mapping-based Augmentation: The original harmful euphemisms in the samples are replaced with their corresponding harmful meaning mapped words from the lexicon, generating similar augmented samples of the original harmful euphemisms.
- (2)
- Substitution-based Augmentation: Utilizing the morphological variants of harmful euphemisms generated by the deep prompt CoT, the original harmful euphemisms in the sentence samples are replaced with these variants to generate similar augmented samples.
- (3)
- Randomized Augmentation: Three randomization strategies commonly used in traditional contrastive learning frameworks are comprehensively applied, namely dropout, cutoff (dimensional clipping), and character shuffling.
3.2.4. Multi-Head Contrastive Learning Stage
3.2.5. Representation Learning Training Workflow
- (1)
- Input Preprocessing and Knowledge Matching: We first construct training samples of harmful euphemisms. These samples are matched against a local harmful euphemism dictionary to extract local sensitive vocabulary and related knowledge. This process constructs a structured corpus containing the original sentence, euphemism, and its semantic interpretation.
- (2)
- Deep Prompt Chain-of-Thought Generation: Based on a cognition-driven prompt design, we construct a CoT template. An LLM is then used to complete the CoT template and leverage external knowledge for the harmful euphemisms based on this final template. Structured auxiliary information is extracted according to this chain, representing the original sentence, euphemism, semantic interpretation, transformed variants, CoT reasoning, and attack type. This stage provides guidance for external knowledge augmentation for subsequent representation learning.
- (3)
- Contrastive Learning Sample Construction: A pre-trained language model encodes the original sentences to obtain base semantic representations. Subsequently, the acquired auxiliary knowledge of harmful euphemisms is used to construct contrastive pairs.
- (4)
- Multi-Head Contrastive Learning Training: The enhanced inputs are passed through 12 projection heads, mapping them into sub-spaces of different dimensions. After constructing positive and negative pairs, the multi-head contrastive learning module calculates the contrastive loss within each sub-space. A dynamic weighting mechanism is applied to adjust the various distances, yielding a comprehensive contrastive loss. Next, the Adam optimizer is utilized to optimize this final loss, thereby updating the parameters across all layers of the model. This enables the semantic encoder to achieve optimized representations for euphemism detection.
- (5)
- Enhanced Semantic Vector Output: After freezing the encoder parameters, the text of harmful euphemisms is fed into the model, and the encoder outputs enhanced semantic vectors through forward propagation. The harmful euphemism feature vectors, optimized through this representation learning process, can be directly used for downstream tasks, such as harmful euphemism detection and harmful word recognition.
3.3. Detection Model
3.3.1. Semantic Perception Modules
- (1)
- Syntactic Semantic Perception Module
- (2)
- Contextual Semantic Perception Module
- (3)
- Rhetorical Semantic Perception Module
3.3.2. Cross-Channel Dynamic Adaptive Fusion Method
- (1)
- Input Layer
- (2)
- Adaptive Multi-channel Attention Mechanism Layer
- (3)
- Dynamic Parameter Fusion Layer
- (4)
- Output Layer
4. Experiments
4.1. Experimental Setup
4.1.1. Dataset
4.1.2. Baselines
4.1.3. Evaluation Metrics
4.1.4. Evaluation Settings for LLMs and APIs
4.1.5. Experimental Setup
5. Results and Analysis
5.1. Representation Learning Comparative Experiments
5.2. Ablation Study of the Representation Learning Module
- HER-MCL (-DPTC): Removes the deep prompt chain-of-thought module, employing a vanilla contrastive learning mechanism.
- HER-MCL (-mapping): Removes the mapping-based augmented samples from the three-fold data augmentation module, utilizing only the other two types to construct contrastive learning samples.
- HER-MCL (-replacing): Removes the replacement-based augmented samples from the three-fold data augmentation module, utilizing only the other two types to construct contrastive learning samples.
- HER-MCL (-MCL): Removes the multi-head contrastive learning mechanism, relying solely on the SimCSE mechanism.
5.3. Model Comparison Experiment
5.4. Real-Time Performance Experiments
5.5. Ablation Study
- HED-MSP (-SSPM): Removes the syntactic lexical semantic perception module, retaining the rest.
- HED-MSP (-CSPM): Removes the contextual background semantic perception module, retaining the rest.
- HED-MSP (-RSPM): Removes the rhetorical behavior semantic perception module, retaining the rest.
- HED-MSP (-CDAF): Removes the cross-channel fusion module, retaining the rest.
5.6. Model Generalization Experiments
5.7. Cross-Dataset Experiments
5.8. Case Study
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Liu, Z.; Shao, Z.; Wang, H.; Li, B. DDML: Multi-Student Knowledge Distillation for Hate Speech. Entropy 2025, 27, 417. [Google Scholar] [CrossRef] [PubMed]
- Founta, A.; Djouvas, C.; Chatzakou, D.; Leontiadis, I.; Blackburn, J.; Stringhini, G.; Vakali, A.; Sirivianos, C.; Kourtellis, N. Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior. In Proceedings of the 12th International Conference on Web and Social Media, Stanford, CA, USA, 25–28 June 2018; pp. 491–500. [Google Scholar]
- Fortuna, P.; Nunes, S. A Survey on Automatic Detection of Hate Speech in Text. ACM Comput. Surv. 2018, 51, 1–30. [Google Scholar] [CrossRef]
- ElSherief, M.; Ziems, C.; Muchlinski, D.; Rezvan, M.; Paasch, C.; Stork, J.; Glass, J.; Walentowska, M.; Zhuravskaya, M.; Yang, D. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. In Proceedings of the 18th Conference on Empirical Methods in Natural Language Processing (EMNLP), Punta Cana, Dominican Republic, 7–11 November 2021; pp. 345–363. [Google Scholar]
- Davidson, T.; Warmsley, D.; Macy, M.; Weber, I. Automated Hate Speech Detection and the Problem of Offensive Language. In Proceedings of the 11th International Conference on Web and Social Media (ICWSM), Montréal, QC, Canada, 15–18 May 2017; pp. 512–515. [Google Scholar]
- Waseem, Z.; Hovy, D. Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. In Proceedings of the 14th NAACL Student Research Workshop, San Diego, CA, USA, 12–17 June 2016; pp. 88–96. [Google Scholar]
- Yuan, K.; Lu, H.; Liao, X.; Wang, X. Reading Thieves’ Cant: Automatically Identifying and Understanding Dark Jargons from Cybercrime Marketplaces. In Proceedings of the 27th USENIX Security Symposium, Baltimore, MD, USA, 15–17 August 2018; pp. 1027–1041. [Google Scholar]
- Taylor, J.; Peignon, M.; Chen, Y.S. Surfacing Contextual Hate Speech Words within Social Media. arXiv 2017, arXiv:1711.10093. [Google Scholar] [CrossRef]
- Zhu, W.; Gong, H.; Bansal, R.; Wu, J.; Li, S.; Liu, J. Self-Supervised Euphemism Detection and Identification for Content Moderation. In Proceedings of the 42nd IEEE Symposium on Security and Privacy (SP), San Francisco, CA, USA, 24–27 May 2021; pp. 229–246. [Google Scholar]
- Huang, L.; Wang, S.; Liu, C.; Zhang, X. Low-Frequency Aware Unsupervised Detection of Dark Jargon Phrases on Social Platforms. In Proceedings of the 20th Pacific Rim International Conference on Artificial Intelligence (PRICAI), Jakarta, Indonesia, 15–19 November 2023; pp. 198–209. [Google Scholar]
- Guan, Z.; Zhang, P.; Gu, H.; Wang, X. JargonFM: A Framework With Multiple Interpretation Modes for Jargon Understanding in Online Communities. IEEE Trans. Comput. Soc. Syst. 2024, 11, 1853–1864. [Google Scholar] [CrossRef]
- Kim, Y.; Park, S.; Han, Y.S. Generalizable Implicit Hate Speech Detection using Contrastive Learning. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), Gyeongju, Republic of Korea, 12–17 October 2022; pp. 6667–6679. [Google Scholar]
- Lu, J.; Lin, H.; Zhang, X.; Xu, B.; Yang, L.; Luo, J. Hate Speech Detection via Dual Contrastive Learning. IEEE/ACM Trans. Audio Speech Lang. Process. 2023, 31, 2787–2795. [Google Scholar] [CrossRef]
- Deng, J.; Zhou, J.; Sun, H.; Zhang, C.; Zheng, F.; Min, L.; Zheng, H. COLD: A Benchmark for Chinese Offensive Language Detection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 11580–11599. [Google Scholar]
- Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, V.; Kamar, E. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland, 22–27 May 2022; pp. 3309–3326. [Google Scholar]
- Pavlopoulos, J.; Laugier, L.; Xenos, A.; Kougia, V.; Androutsopoulos, I. From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland, 22–27 May 2022; pp. 3721–3734. [Google Scholar]
- Wang, Y.; Su, H.; Wu, Y.; Zhang, X. SICM: A Supervised-Based Identification and Classification Model for Chinese Jargons Using Feature Adapter Enhanced BERT. In Proceedings of the 20th Pacific Rim International Conference on Artificial Intelligence (PRICAI), Shanghai, China, 10–13 November 2022; pp. 297–308. [Google Scholar]
- Huang, X.; Zhao, J.; Shi, J.; Wang, Z. Syntactic Enhanced Euphemisms Identification Based on Graph Convolution Networks and Dependency Parsing. In Proceedings of the 8th International Conference on Data Science in Cyberspace (DSC), Hefei, China, 20–22 October 2023; pp. 172–180. [Google Scholar]
- Ghosh, S.; Suri, M.; Chiniya, P.; Singhania, S.; Singh, R.K. CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 6159–6173. [Google Scholar]
- Zhang, L.; Jin, L.; Sun, X.; Zhao, Y. TOT: Topology-Aware Optimal Transport for Multimodal Hate Detection. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; pp. 4884–4892. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Maeda, M.; Zhao, F.; Xia, Y.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar]
- Kojima, T.; Gu, S.S.; Reid, M.; Matsuo, Y.; Iwasawa, Y. Large Language Models are Zero-Shot Reasoners. Adv. Neural Inf. Process. Syst. 2022, 35, 22199–22213. [Google Scholar]
- Zhao, R.; Zhao, F.; Wang, L.; Zhang, X. KG-CoT: Chain-of-Thought Prompting of Large Language Models over Knowledge Graphs for Knowledge-Aware Question Answering. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI), Jeju, Republic of Korea, 3–9 August 2024; pp. 6642–6650. [Google Scholar]
- Mondal, D.; Modi, S.; Panda, S.; Singh, A. KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; pp. 18798–18806. [Google Scholar]
- Magu, R.; Luo, J. Determining Code Words in Euphemistic Hate Speech Using Word Embedding Networks. In Proceedings of the 2nd Workshop on Abusive Language Online (ALW2), Brussels, Belgium, 31 October 2018; pp. 93–100. [Google Scholar]
- Wang, C.; Shen, Y.; Li, Y.; Liu, J.; Zhang, Y. A Systematic Empirical Study on Word Embedding Based Methods in Discovering Chinese Black Keywords. Eng. Appl. Artif. Intell. 2023, 125, 106775. [Google Scholar] [CrossRef]
- Wiegand, M.; Ruppenhofer, J.; Eder, E. Implicitly Abusive Language—What does it actually look like and why are we not getting there? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, 6–11 June 2021; pp. 576–587. [Google Scholar]
- Wiegand, M.; Ruppenhofer, J.; Kleinbauer, T. Detection of Abusive Language: The Problem of Biased Datasets. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 602–608. [Google Scholar]
- Breitfeller, L.; Ahn, E.; Jurgens, D.; Tsvetkov, Y. Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 1664–1674. [Google Scholar]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020; pp. 9726–9735. [Google Scholar]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning (ICML), Online, 13–18 July 2020; pp. 1575–1585. [Google Scholar]
- Gao, T.; Yao, X.; Chen, D. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online, 7–11 November 2021; pp. 6894–6910. [Google Scholar]
- Gunel, B.; Du, J.; Conneau, A.; Stoyanov, V. Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning. In Proceedings of the 9th International Conference on Learning Representations (ICLR), Online, 3–7 May 2021. [Google Scholar]
- Lin, Y.; Gou, Y.; Liu, X.; Bai, J.; Lv, J.; Peng, X. Dual Contrastive Prediction for Incomplete Multi-View Representation Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 4447–4461. [Google Scholar] [CrossRef] [PubMed]
- Nejadgholi, I.; Fraser, K.; Kiritchenko, S. Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; pp. 5517–5529. [Google Scholar]
- Zhang, Z.; Zhang, A.; Li, M.; Smola, A. Automatic Chain of Thought Prompting in Large Language Models. arXiv 2022, arXiv:2210.03493. [Google Scholar] [CrossRef]
- Zhou, Z.; Tao, R.; Zhu, J.; Liu, Y. Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? Adv. Neural Inf. Process. Syst. 2024, 37, 123846–123910. [Google Scholar]
- Wu, J.; Yu, T.; Chen, X.; Wang, S. DeCoT: Debiasing Chain-of-Thought for Knowledge-Intensive Tasks in Large Language Models via Causal Intervention. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand, 11–16 August 2024; pp. 14073–14087. [Google Scholar]
- Chen, W.; Dang, Y.; Zhang, X. A Multimodal Semantic-Enhanced Attention Network for Fake News Detection. Entropy 2025, 27, 746. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Y.; Yang, Q. A Survey on Multi-Task Learning. IEEE Trans. Knowl. Data Eng. 2021, 34, 5586–5609. [Google Scholar] [CrossRef]
- Liu, H.; Burnap, P.; Alorainy, W.; Williams, M.L. Fuzzy Multi-task Learning for Hate Speech Type Identification. In Proceedings of the World Wide Web Conference (WWW), San Francisco, CA, USA, 13–17 May 2019; pp. 3006–3012. [Google Scholar]
- Bendjoudi, I.; Vanderhaegen, F.; Hamad, D.; Dornaika, F. Multi-label, Multi-task CNN Approach for Context-based Emotion Recognition. Inf. Fusion 2021, 76, 422–428. [Google Scholar] [CrossRef]
- Zhang, Y.; Wang, J.; Liu, Y.; Zhang, X. A Multitask Learning Model for Multi-modal Sarcasm, Sentiment and Emotion Recognition in Conversations. Inf. Fusion 2023, 93, 282–301. [Google Scholar] [CrossRef]
- Liu, X.; He, P.; Chen, W.; Gao, J. Multi-Task Deep Neural Networks for Natural Language Understanding. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 4487–4496. [Google Scholar]
- Ma, J.; Zhao, Z.; Yi, X.; Chen, J.; Lichman, M.; Chi, E.H. Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK, 19–23 August 2018; pp. 1930–1939. [Google Scholar]
- Duong, L.; Cohn, T.; Bird, S.; Cook, P. Low Resource Dependency Parsing: Cross-lingual Parameter Sharing in a Neural Network Parser. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, Beijing, China, 26–31 July 2015; pp. 845–850. [Google Scholar]
- Zhou, J.; Deng, J.; Mi, F.; Li, Z.; Zheng, H.; Yao, X. Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 3576–3591. [Google Scholar]
- Manzini, T.; Yao Chong, L.; Black, A.W.; Tsvetkov, Y. Black is to Criminal as Caucasian is to Police: Detecting and Removing Multiclass Bias in Word Embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, Minneapolis, MN, USA, 2–7 June 2019; pp. 615–621. [Google Scholar]
- Wang, W.; Huang, J.; Chen, C.; Zhang, X. Validating Multimedia Content Moderation Software via Semantic Fusion. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), Vienna, Austria, 17–21 July 2023; pp. 576–588. [Google Scholar]
- Pérez, J.M.; Luque, F.M. Atalaya at SemEval-2019 Task 5: Robust Embeddings for Tweet Classification. In Proceedings of the 13th International Workshop on Semantic Evaluation, Minneapolis, MN, USA, 6–7 June 2019; pp. 64–69. [Google Scholar]
- Li, J.; Du, T.; Ji, S.; Zhang, R.; Lyu, M.R. TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine Translation. In Proceedings of the 29th USENIX Security Symposium, Online, 12–14 August 2020; pp. 1381–1398. [Google Scholar]
- Dessì, D.; Recupero, D.R.; Sack, H. An Assessment of Deep Learning Models and Word Embeddings for Toxicity Detection within Online Textual Comments. Electronics 2021, 10, 779. [Google Scholar] [CrossRef]
- Neog, M.; Baruah, N. A Deep Learning Framework for Assamese Toxic Comment Detection: Leveraging LSTM and BiLSTM Models with Attention Mechanism. In Proceedings of the 2nd International Conference on Advances in Data-Driven Computing and Intelligent Systems (ADCIS), Singapore, 21–23 September 2023; pp. 485–497. [Google Scholar]
- Huang, Z.; Xu, W.; Yu, K. Bidirectional LSTM-CRF Models for Sequence Tagging. arXiv 2015, arXiv:1508.01991. [Google Scholar] [CrossRef]
- Lai, S.; Xu, L.; Liu, K.; Zhao, J. Recurrent Convolutional Neural Networks for Text Classification. In Proceedings of the 29th AAAI Conference on Artificial Intelligence, Austin, TX, USA, 25–30 January 2015; pp. 2267–2273. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Lee, P.; Trujillo, A.C.; Plancarte, D.C. MEDs for PETs: Multilingual Euphemism Disambiguation for Potentially Euphemistic Terms. arXiv 2024, arXiv:2401.14526. [Google Scholar] [CrossRef]
- He, W.; Vieira, T.K.; Garcia, M. Investigating Idiomaticity in Word Representations. Comput. Linguist. 2025, 51, 505–555. [Google Scholar] [CrossRef]
- Hu, Y.; Li, J.; Wang, T. A Unified Generative Framework for Bilingual Euphemism Detection and Identification. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 6753–6766. [Google Scholar]




| Category | Reference | Specific Dataset | Backbone Model | External Knowledge | |
|---|---|---|---|---|---|
| Lexicon | LLM/CoT | ||||
| Rules & Static Rep. | Davidson [5] | Twitter tweets | LR, SVM, Random Forest | √ | × |
| Waseem [6] | Twitter tweets | LR, SVM | √ | × | |
| Yuan [7] | Dark Web | Word2Vec | √ | × | |
| Taylor [8] | DailyStormer, Twitter | Word2Vec + PageRank | √ | × | |
| Dynamic Rep. | Zhu [9] | Reddit, Gab | BERT | × | × |
| Huang [10] | Tieba, Sina Weibo | Transformer | × | × | |
| Guan [11] | Online Communities | RoBERTa | × | × | |
| Contrastive Learning | Kim [12] | IHC, SBIC, DynaHate | Dual-Encoder (BERT) | × | × |
| Lu [13] | OffensEval, HatEval | Dual-Encoder (RoBERTa) | × | × | |
| Deng [14] | AugCOLD | Multi-teacher Distillation | × | × | |
| Hartvigsen [15] | ToxiGen | HateBERT, RoBERTa | × | × | |
| Structure & Multimodal | Pavlopoulos [16] | TOXICSPANS | Sequence Labeling | × | × |
| Wang [17] | CBKD | Feature Adapter BERT | √ | × | |
| Huang [18] | GCN | × | × | ||
| Ghosh [19] | Hyperbolic Network | × | × | ||
| Zhang [20] | MultiOFF, Memotion | Optimal Transport Framework | × | × | |
| LLM Reasoning | Wei [21] | GSM8K, SVAMP | PaLM, GPT-3 (CoT) | × | √ |
| Kojima [22] | MultiArith | GPT-3 (Zero-shot) | × | √ | |
| Zhao [23] | CSQA, OpenBookQA | LLM + GNN | × | √ | |
| Mondal [24] | ScienceQA | LLM + Vision + KG | × | √ | |
| Topic Category | Seed Keyword List (Examples in Chinese) |
|---|---|
| Group Attacks | 猩猩 [Lit: “Gorilla”, a racist slur for people of African descent], 幕刃 [Lit: “Twilight Blade”, a homophonic derogatory term for women]. |
| Vulgar Speech | 金针菇 [Lit: “Enoki Mushroom”, denoting male genitalia], 菊花 [Lit: “Chrysanthemum”, denoting the anus]. |
| Political Content | 大毛 [Lit: “Big Fur”, implying Russians], 棒子 [Lit: “Stick”, a derogatory slur for Koreans]. |
| General Insults | 鸡 [Lit: “Chicken”, a metaphor for sex workers], 孙子 [Lit: “Grandson”, an insult to one’s lineage]. |
| Topic Type | Attack Type | Harmful | Harmless | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| HP | PS | VS | Meme | Abbrev. | MP | Irony | CP | |||
| Group Attacks | 356 | 370 | 514 | 586 | 499 | 633 | 771 | 589 | 4318 | 5594 |
| Vulgar Speech | 496 | 412 | 294 | 123 | 599 | 480 | 322 | 331 | 3057 | 3356 |
| Political Content | 146 | 235 | 128 | 164 | 217 | 589 | 288 | 186 | 1953 | 2913 |
| General Insults | 230 | 188 | 94 | 11 | 298 | 416 | 499 | 417 | 2153 | 3173 |
| Total | 1228 | 1205 | 1030 | 884 | 1613 | 2118 | 1880 | 1523 | 11,481 | 15,036 |
| Stage & Objective | Deep Prompt CoT Template |
|---|---|
| Phase I: Task Definition Obj: Understand requirements and task formulation. | At this stage, I assume the role of a linguistics and cognitive science expert with specialized knowledge in harmful euphemism analysis. The input consists of: (i) a target sample or text segment containing potential harmful euphemisms, (ii) identified euphemistic expressions, (iii) corresponding semantic interpretation labels, and (iv) auxiliary prompt seed words. The expected output includes: (i) cognitive reasoning analysis of the input sample, (ii) semantically similar samples, (iii) extracted euphemistic expressions, (iv) mapped harmful semantic interpretations, and (v) refined cognitive reasoning chains. A step-by-step reasoning process is required, integrating all relevant linguistic and contextual factors to produce a structured analytical output. |
| Phase II: Cognitive Understanding Obj: Formulate structured understanding of euphemisms. | In this stage, I establish a formal understanding of harmful euphemisms. From a semantic perspective, harmful euphemisms refer to textual expressions in which discriminatory, violent, or illegal intents are concealed through euphemistic transformations. From a structural perspective, they may appear as either short phrases or long-form discourse units, while euphemistic expressions typically exist at lexical or phrasal levels with context-dependent meanings. Furthermore, harmful euphemism interpretation is guided by five core cognitive dimensions, including eight types of attack strategies: homophony, polysemy, visual similarity, meme-based transformation, abbreviation, metaphor, irony, and compound structures. |
| Phase III: Prompt Alignment Obj: Analyze semantic alignment with examples. | Given the current input sample and provided prompt examples, I analyze their semantic relationships and describe how both jointly contribute to fulfilling the task objective. In particular, I evaluate semantic similarity between the target sample and reference examples, and identify key divergences in linguistic expression and intent encoding. |
| Phase IV: Cognitive Reasoning Obj: Derive structured reasoning chains. | I conduct a detailed analysis of how harmful euphemistic expressions are manifested within the input text and how they are interpreted as implicit harmful meanings. Based on the defined cognitive framework, I identify specific attack strategies and construct a structured reasoning chain explaining the detection process. This includes linguistic interpretation grounded in cognitive linguistics principles and logical inference over euphemistic transformations. |
| Phase V: Target Generation Obj: Generate new samples with semantic mappings. | Based on the task description and cognitive analysis, I generate semantically similar harmful euphemistic expressions. I further identify corresponding euphemistic terms and their mapped harmful meanings. A structured representation is required in the form: euphemistic term–attack strategy–semantic interpretation. Additionally, generated samples must be consistent with reference examples and cognitive reasoning constraints. |
| Phase VI: Output Evaluation Obj: Ensure quality and task compliance. | Finally, I evaluate whether the generated outputs satisfy all task constraints. This includes assessing semantic correctness, structural consistency, and alignment with prior definitions. If the outputs do not meet the requirements, the process returns to Stage 2 for iterative refinement. Otherwise, the final results are accepted and output. |
| Method | Pre-Trained Model | Pearson Correlation | Spearman’s Rank Correlation |
|---|---|---|---|
| HER-MCL (-DPTC) | BERT | 0.7924 | 0.7948 |
| DeBERTa | 0.7543 | 0.7689 | |
| DistilBERT | 0.7256 | 0.7327 | |
| RoBERTa | 0.7823 | 0.7839 | |
| HER-MCL (-mapping) | BERT | 0.8067 | 0.8135 |
| DeBERTa | 0.7629 | 0.7812 | |
| DistilBERT | 0.7359 | 0.7416 | |
| RoBERTa | 0.8114 | 0.8032 | |
| HER-MCL (-replacing) | BERT | 0.8239 | 0.8297 |
| DeBERTa | 0.7784 | 0.7928 | |
| DistilBERT | 0.7523 | 0.7584 | |
| RoBERTa | 0.8337 | 0.8219 | |
| HER-MCL (-MCL) | BERT | 0.7532 | 0.7536 |
| DeBERTa | 0.7689 | 0.7886 | |
| DistilBERT | 0.7187 | 0.7353 | |
| RoBERTa | 0.7976 | 0.7951 | |
| HER-MCL (Ours) | BERT | 0.8387 | 0.8463 |
| DeBERTa | 0.8574 | 0.8643 | |
| DistilBERT | 0.7987 | 0.8073 | |
| RoBERTa | 0.8693 | 0.8765 |
| Category | Model | Acc. (%) | Pre. (%) | Rec. (%) | F1 (%) |
|---|---|---|---|---|---|
| API-based Systems | Perspective-API | 61.33 | 85.89 | 16.75 | 28.04 |
| Tencent-API | 62.64 | 66.67 | 33.84 | 44.89 | |
| Machine Learning Methods | RF | 81.42 | 79.57 | 78.97 | 79.27 |
| SVM | 69.98 | 66.63 | 66.67 | 66.65 | |
| k-NN | 75.49 | 69.65 | 80.67 | 74.76 | |
| AdaBoost | 67.58 | 64.15 | 63.33 | 63.74 | |
| Toxic Content Classifiers | TextFNN | 84.01 | 84.07 | 81.04 | 82.53 |
| TextCNN | 73.10 | 71.23 | 67.39 | 69.25 | |
| RCNN | 83.20 | 82.50 | 80.10 | 81.25 | |
| BiLSTM | 87.96 | 87.88 | 85.67 | 86.76 | |
| BiLSTM-Att | 88.19 | 88.35 | 85.79 | 87.05 | |
| BiLSTM-CRF | 88.30 | 88.20 | 86.50 | 87.30 | |
| BERT-base | 88.81 | 86.69 | 88.99 | 87.60 | |
| BERT-large | 88.37 | 81.36 | 92.06 | 86.14 | |
| RoBERTa-base | 89.97 | 87.56 | 90.61 | 88.90 | |
| RoBERTa-large | 89.01 | 85.66 | 90.67 | 87.95 | |
| SBERT | 86.53 | 84.43 | 85.43 | 84.92 | |
| Large Language Models | GPT-4o | 72.54 | 63.06 | 93.97 | 75.48 |
| GPT-4-Turbo | 77.26 | 70.67 | 84.48 | 76.96 | |
| OpenAI-o1 | 85.35 | 85.12 | 85.48 | 85.30 | |
| Qwen2-72B | 70.17 | 61.86 | 87.82 | 72.59 | |
| LLaMA3-70B | 55.13 | 51.45 | 4.16 | 7.69 | |
| ChatGLM3-6B | 56.52 | 64.32 | 7.49 | 13.42 | |
| DeepSeek-V3 | 83.64 | 83.27 | 82.92 | 83.15 | |
| DeepSeek-R1 | 84.85 | 84.63 | 85.12 | 84.88 | |
| HED-MSP (Ours) | 93.94 | 93.09 | 93.36 | 93.23 | |
| Model | Latency | Throughput | Memory | Params | Cost |
|---|---|---|---|---|---|
| (ms/Sample) | (Samples/s) | (GB) | (B) | (Yuan/1M) | |
| BERT-large | 5–10 | 150–250 | 8–10 | 0.34 | ∼0 |
| RoBERTa-large | 5–10 | 150–250 | 8–10 | 0.25 | ∼0 |
| ChatGLM3-6B | 50–100 | 40–80 | 15–20 | 6.0 | ∼0 |
| LLaMA3-70B | 1000–2000 | 5–15 | 70–80 | 70.0 | ∼0 |
| Qwen2-72B | 1000–2000 | 5–15 | 60–80 | 72.0 | 3–5 |
| DeepSeek-V3 | 1500–3000 | 20–50 | N/A | N/A | 2–8 |
| DeepSeek-R1 | 2000–4000 | 20–50 | N/A | N/A | 4–16 |
| GPT-4o (API) | 1000–2500 | 50–100 | N/A | N/A | 35–105 |
| GPT-4-Turbo | 1500–3000 | ∼50 | N/A | N/A | 70–210 |
| HED-MSP (Ours) | 2–5 | 300–500 | 4–9 | 0.23 | ∼0 |
| Model | Accuracy (%) | Precision (%) | Recall (%) | F1 (%) |
|---|---|---|---|---|
| RoBERTa-base | 89.97 | 87.56 | 90.61 | 88.90 |
| OpenAI-o1 | 85.35 | 85.12 | 85.48 | 85.30 |
| HED-MSP (-SSPM) | 92.15 | 91.32 | 91.58 | 91.45 |
| HED-MSP (-CSPM) | 91.73 | 90.85 | 91.12 | 90.98 |
| HED-MSP (-RSPM) | 91.35 | 90.47 | 90.74 | 90.60 |
| HED-MSP (-CDAF) | 92.84 | 92.05 | 92.32 | 92.18 |
| HED-MSP (Ours) | 93.94 | 93.09 | 93.36 | 93.23 |
| Baseline Type | Baseline Name | Group Attacks → General Insults | General Insults → Group Attacks | ||||
|---|---|---|---|---|---|---|---|
| Acc. (%) | Pre. (%) | F1 (%) | Acc. (%) | Pre. (%) | F1 (%) | ||
| Harmful Content Classifier | TextFNN | 56.85 | 53.91 | 52.50 | 55.28 | 55.20 | 52.08 |
| TextCNN | 55.14 | 52.37 | 50.88 | 53.57 | 53.46 | 50.56 | |
| RCNN | 55.50 | 53.31 | 51.56 | 54.33 | 52.70 | 49.49 | |
| BiLSTM | 56.93 | 54.05 | 56.02 | 54.46 | 51.10 | 47.13 | |
| BiLSTM-Att | 55.76 | 53.57 | 53.59 | 53.97 | 50.92 | 48.85 | |
| BiLSTM-CRF | 57.83 | 55.52 | 54.27 | 56.33 | 53.54 | 52.81 | |
| BERT-base | 61.85 | 58.61 | 59.06 | 60.06 | 55.80 | 57.68 | |
| BERT-large | 62.89 | 60.18 | 57.43 | 57.21 | 53.45 | 53.64 | |
| RoBERTa-base | 60.27 | 59.29 | 52.89 | 57.11 | 53.90 | 54.91 | |
| RoBERTa-large | 64.63 | 62.64 | 59.27 | 58.43 | 54.97 | 55.10 | |
| SBERT | 50.73 | 48.37 | 40.72 | 53.76 | 51.06 | 54.75 | |
| LLMs | GPT-4o | 84.05 | 76.62 | 86.13 | 72.00 | 65.44 | 77.02 |
| GPT-4-Turbo | 83.20 | 75.04 | 84.54 | 70.51 | 63.84 | 75.67 | |
| Qwen2-72B | 80.17 | 73.55 | 82.91 | 68.95 | 63.99 | 73.80 | |
| LLaMa3-70B | 47.84 | 41.67 | 7.58 | 49.90 | 51.43 | 12.00 | |
| ChatGLM3-6B | 46.55 | 14.29 | 1.57 | 51.62 | 63.16 | 15.84 | |
| DeepSeek-V3 | 68.50 | 65.00 | 66.80 | 62.00 | 56.30 | 60.50 | |
| DeepSeek-R1 | 79.20 | 72.50 | 81.00 | 69.80 | 64.20 | 72.50 | |
| Our Model | HED-MSP | 79.96 | 75.24 | 78.43 | 66.05 | 63.16 | 65.55 |
| Baseline Type | Model | EACL | FigLang | JointEDI | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pre. | Rec. | F1 | Pre. | Rec. | F1 | Pre. | Rec. | F1 | ||
| Harmful Content Classifier | TextCNN | 80.85 | 85.31 | 82.99 | 71.95 | 77.73 | 74.69 | 86.74 | 89.90 | 88.28 |
| BiLSTM | 80.86 | 83.42 | 82.12 | 65.04 | 64.75 | 64.86 | 83.23 | 85.26 | 84.21 | |
| BiLSTM-Att | 80.33 | 80.45 | 80.20 | 72.58 | 72.87 | 72.71 | 86.40 | 86.51 | 86.45 | |
| BiLSTM-CRF | 81.30 | 82.05 | 81.64 | 73.89 | 71.78 | 72.74 | 85.54 | 87.17 | 86.31 | |
| RCNN | 80.86 | 83.42 | 82.12 | 72.58 | 72.87 | 72.71 | 86.40 | 86.51 | 86.45 | |
| BERT-base | 83.91 | 88.15 | 85.93 | 72.71 | 81.83 | 76.86 | 86.48 | 90.87 | 88.60 | |
| BERT-large | 84.12 | 86.75 | 85.39 | 78.69 | 80.38 | 79.38 | 87.54 | 87.78 | 87.63 | |
| RoBERTa-base | 81.45 | 90.99 | 85.85 | 70.14 | 87.08 | 77.58 | 85.82 | 83.69 | 84.71 | |
| RoBERTa-large | 82.18 | 93.56 | 87.47 | 72.48 | 82.40 | 76.80 | 85.11 | 88.44 | 86.66 | |
| LLMs | LLaMa3-70B | 77.42 | 73.97 | 75.66 | 66.88 | 82.24 | 73.77 | 79.92 | 50.76 | 62.09 |
| ChatGLM3-6B | 71.56 | 50.00 | 58.87 | 67.34 | 72.13 | 69.66 | 90.79 | 68.84 | 78.31 | |
| Qwen2-72B | 83.75 | 79.45 | 81.55 | 68.43 | 95.35 | 79.68 | 79.95 | 58.57 | 67.61 | |
| DeepSeek-V3 | 84.50 | 91.00 | 87.50 | 77.85 | 84.20 | 80.89 | 86.30 | 90.75 | 88.44 | |
| DeepSeek-R1 | 85.65 | 91.25 | 88.35 | 79.12 | 85.40 | 82.13 | 89.70 | 89.15 | 89.42 | |
| GPT-4o | 84.09 | 88.70 | 86.33 | 72.63 | 92.08 | 81.20 | 86.83 | 58.38 | 69.82 | |
| GPT-4-Turbo | 83.40 | 75.68 | 79.35 | 73.05 | 87.43 | 79.60 | 92.00 | 43.14 | 58.89 | |
| Our Model | HED-MSP | 85.16 | 90.41 | 87.71 | 77.47 | 83.61 | 80.42 | 89.48 | 88.79 | 89.13 |
| Euphemisms (Connotation) | Comments | True Label | Prediction |
|---|---|---|---|
| Animals (Black people) 动物(黑人) | Case 1. I can only say, Humans and animals cannot mix together. 案例1. 我只能说,人跟动物[猴子表情]就不能混在一起 | Harmful | Harmless |
| Stick (Korean & Korea) 棒子(韩国人及韩国) | Case 2. Seriously, who actually likes those sticks? Can’t you tell your “oppas” to hit the gym a bit more? 案例2. 到底谁喜欢棒子啊,让你家哥哥多练练好吗[恶心表情*3] | Harmful | Harmless |
| China aluminum (Chinese females) 国铝(中国女性) | Case 3. Sorry, we domestic aluminum, who are short in height and low in education, are simply not good enough for you. 案例3. 不好意思,我们身高矮、学历低的国铝根本配不上您。 | Harmless | Harmful |
| Woman’s Fist (Feminism) 女拳(女权) | Case 4. You cannot just say “woman’s fist” the moment you see a woman. Isn’t it better for everyone to get along peacefully? 案例4. 你们不能看到一个女的就说女拳,大家和平相处不好嘛 | Harmless | Harmful |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xie, S.; Zhou, G.; Wang, H. A Deep Prompt-Based Chain-of-Thought Approach to Harmful Euphemism Detection in Social Networks. Entropy 2026, 28, 560. https://doi.org/10.3390/e28050560
Xie S, Zhou G, Wang H. A Deep Prompt-Based Chain-of-Thought Approach to Harmful Euphemism Detection in Social Networks. Entropy. 2026; 28(5):560. https://doi.org/10.3390/e28050560
Chicago/Turabian StyleXie, Siyu, Gang Zhou, and Haizhou Wang. 2026. "A Deep Prompt-Based Chain-of-Thought Approach to Harmful Euphemism Detection in Social Networks" Entropy 28, no. 5: 560. https://doi.org/10.3390/e28050560
APA StyleXie, S., Zhou, G., & Wang, H. (2026). A Deep Prompt-Based Chain-of-Thought Approach to Harmful Euphemism Detection in Social Networks. Entropy, 28(5), 560. https://doi.org/10.3390/e28050560
