Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis
Abstract
1. Summary
State of the Art
- We developed a privacy-aware code-mixed cyberbullying dataset (PCCD) with session-based identifiers, profane word annotation, and abuser–victim identity annotations for examining the power imbalance among abuser(s) and victim(s). To the best of our knowledge, PCCD is the first code-mixed Cantonese dataset of this kind, which enables machine learning models to effectively evaluate the above three aspects for the detection of cyberbullying incidents.
- We proposed and demonstrated a privacy-aware anonymization procedure that helps improve the prediction model by simplifying names semantically. It was empirically demonstrated that the use of the anonymized data improved performance in terms of precision, recall, and F1 score. This anonymization could also help improve generalization by prohibiting the model from memorizing famous usernames.
- Our dataset contains code-mixed linguistic information and relevant toxic patterns observed in daily conversational dialogues, which should help to build a more holistic multilingual cyberbullying corpus for transfer learning.
2. Methods
2.1. Data Source
2.2. Data Cleansing
2.3. Data Anonymization
2.4. Human Annotation
2.5. Threshold-Based Cyberbullying Binary Label
3. Data Description
3.1. Session-Based Messages
3.2. Characteristics of Tweet Textual Content
3.3. Data Statistics
3.4. Validation of Proposed Dataset Construction Approach
4. User Notes
4.1. Limitations
4.2. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Dinakar, K.; Jones, B.; Havasi, C.; Lieberman, H.; Picard, R. Common sense reasoning for detection, prevention, and mitigation of cyberbullying. ACM Trans. Interact. Intell. Syst. 2012, 2, 1–30. [Google Scholar] [CrossRef] [Scilit]
- Van Hee, C.; Jacobs, G.; Emmery, C.; Desmet, B.; Lefever, E.; Verhoeven, B.; De Pauw, G.; Daelemans, W.; Hoste, V. Automatic detection of cyberbullying in social media text. PLoS ONE 2018, 13, e0203794. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kowalski, R.M.; Limber, S.P. Psychological, physical, and academic correlates of cyberbullying and traditional bullying. J. Adolesc. Health 2013, 53, S13–S20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khine, A.T.; Saw, Y.M.; Htut, Z.Y.; Khaing, C.T.; Soe, H.Z.; Swe, K.K.; Thike, T.; Htet, H.; Saw, T.N.; Cho, S.M.; et al. Assessing risk factors and impact of cyberbullying victimization among university students in Myanmar: A cross-sectional study. PLoS ONE 2020, 15, e0227051. [Google Scholar] [CrossRef] [Scilit]
- Aoyama, I.; Saxon, T.F.; Fearon, D.D. Internalizing problems among cyberbullying victims and moderator effects of friendship quality. Multicult. Educ. Technol. J. 2011, 5, 92–105. [Google Scholar] [CrossRef] [Scilit]
- Hinduja, S.; Patchin, J.W. Bullying, cyberbullying, and suicide. Arch. Suicide Res. 2010, 14, 206–221. [Google Scholar] [CrossRef] [Scilit]
- Sampasa-Kanyinga, H.; Roumeliotis, P.; Xu, H. Associations between Cyberbullying and School Bullying Victimization and Suicidal Ideation, Plans and Attempts among Canadian Schoolchildren. PLoS ONE 2014, 9, e102145. [Google Scholar] [CrossRef] [Scilit]
- Salawu, S.; He, Y.; Lumsden, J. Approaches to Automated Detection of Cyberbullying: A survey. IEEE Trans. Affect. Comput. 2017, 11, 3–24. [Google Scholar] [CrossRef] [Scilit]
- Rosa, H.; Pereira, N.; Ribeiro, R.; Ferreira, P.C.; Carvalho, J.P.; Oliveira, S.; Coheur, L.; Paulino, P.; Simão, A.M.V.; Trancoso, I. Automatic cyberbullying detection: A systematic review. Comput. Hum. Behav. 2018, 93, 333–345. [Google Scholar] [CrossRef] [Scilit]
- Cheng, L.; Silva, Y.N.; Hall, D.; Liu, H. Session-Based Cyberbullying Detection: Problems and challenges. IEEE Internet Comput. 2020, 25, 66–72. [Google Scholar] [CrossRef] [Scilit]
- Yi, P.; Zubiaga, A. Session-based cyberbullying detection in social media: A survey. Online Soc. Netw. Media 2023, 36, 100250. [Google Scholar] [CrossRef] [Scilit]
- Maity, K.; Sen, T.; Saha, S.; Bhattacharyya, P. MTBullyGNN: A Graph neural Network-Based multitask framework for cyberbullying detection. IEEE Trans. Comput. Soc. Syst. 2022, 11, 849–858. [Google Scholar] [CrossRef] [Scilit]
- Jacobs, G.; Van Hee, C.; Hoste, V. Automatic classification of participant roles in cyberbullying: Can we detect victims, bullies, and bystanders in social media text? Nat. Lang. Eng. 2020, 28, 141–166. [Google Scholar] [CrossRef] [Scilit]
- Dewani, A.; Memon, M.A.; Bhatti, S.; Sulaiman, A.; Hamdi, M.; Alshahrani, H.; Alghamdi, A.; Shaikh, A. Detection of Cyberbullying Patterns in Low Resource Colloquial Roman Urdu Microtext using Natural Language Processing, Machine Learning, and Ensemble Techniques. Appl. Sci. 2023, 13, 2062. [Google Scholar] [CrossRef] [Scilit]
- Yadav, J.; Kumar, D.; Chauhan, D. Cyberbullying Detection using Pre-Trained BERT Model. In Proceedings of the 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC); IEEE: Piscataway, NJ, USA, 2020; pp. 1096–1100. [Google Scholar] [CrossRef] [Scilit]
- López-Vizcaíno, M.F.; Nóvoa, F.J.; Carneiro, V.; Cacheda, F. Early detection of cyberbullying on social media networks. Future Gener. Comput. Syst. 2021, 118, 219–229. [Google Scholar] [CrossRef] [Scilit]
- Zampieri, M.; Malmasi, S.; Nakov, P.; Rosenthal, S.; Farra, N.; Kumar, R. Predicting the Type and Target of Offensive Posts in Social Media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 1415–1420. [Google Scholar] [CrossRef] [Scilit]
- Zampieri, M.; Morgan, S.; North, K.; Ranasinghe, T.; Simmmons, A.; Khandelwal, P.; Rosenthal, S.; Nakov, P. Target-Based Offensive Language Identification. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 762–770. [Google Scholar] [CrossRef] [Scilit]
- Pavlopoulos, J.; Sorensen, J.; Laugier, L.; Androutsopoulos, I. SEMEVAL-2021 Task 5: Toxic Spans Detection. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
- Lu, J.; Xu, B.; Zhang, X.; Min, C.; Yang, L.; Lin, H. Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 16235–16250. [Google Scholar] [CrossRef] [Scilit]
- Jiang, A.; Yang, X.; Liu, Y.; Zubiaga, A. SWSR: A Chinese dataset and lexicon for online sexism detection. Online Soc. Netw. Media 2021, 27, 100182. [Google Scholar] [CrossRef] [Scilit]
- Yang, Q.; Chen, Y.; Xu, Z.; Shang, Y.-M.; Guo, S.; Zhang, X. SCCD: A session-based dataset for Chinese cyberbullying detection. In Proceedings of the 31st International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 9533–9545. Available online: https://aclanthology.org/2025.coling-main.639 (accessed on 26 February 2026).
- Gupta, V.; Roychowdhury, S.; Das, M.; Banerjee, S.; Saha, P.; Mathew, B.; Vanchinathan, H.; Mukherjee, A. MACD: Multilingual Abusive Comment Detection at Scale for Indic Languages. Adv. Neural Inf. Process. Syst. 2022, 35, 26176–26191. [Google Scholar]
- Muhammad, S.H.; Abdulmumin, I.; Ayele, A.A.; Adelani, D.I.; Ahmad, I.S.; Aliyu, S.M.; Röttger, P.; Oppong, A.; Bukula, A.; Chukwuneke, C.I.; et al. AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 1854–1871. [Google Scholar] [CrossRef] [Scilit]
- Oriola, O.; Kotze, E. Evaluating machine learning techniques for detecting offensive and hate speech in South African tweets. IEEE Access 2020, 8, 21496–21509. [Google Scholar] [CrossRef] [Scilit]
- Chavan, T.; Gokhale, O.; Kane, A.; Patankar, S.; Joshi, R. My Boli: Code-mixed Marathi-English Corpora, Pretrained Language Models and Evaluation Benchmarks. In Findings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 242–249. [Google Scholar] [CrossRef] [Scilit]
- Alhazmi, A.; Mahmud, R.; Idris, N.; Abo, M.E.M.; Eke, C.I. Code-mixing unveiled: Enhancing the hate speech detection in Arabic dialect tweets using machine learning models. PLoS ONE 2024, 19, e0305657. [Google Scholar] [CrossRef] [Scilit]
- Kaware, P.D.; Raut, A.B. Indo-HateSpeech Analysis: A Multi-Level hate speech classification framework using BERT features and machine learning models. Adv. Nonlinear Var. Inequalities 2025, 28, 491–503. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Lin, C.J. TOCP: A Dataset for Chinese Profanity Processing. In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying; European Language Resources Association (ELRA): Paris, France, 2020; pp. 6–12. [Google Scholar]
- Wang, J.; Fu, K.; Lu, C.-T. SOSNET: A Graph Convolutional Network Approach to Fine-Grained Cyberbullying Detection. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data); IEEE: Piscataway, NJ, USA, 2020; pp. 1699–1708. [Google Scholar] [CrossRef] [Scilit]
- Llama Team, AI @ Meta. The Llama 3 Herd of Models. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Kalajdzievski, D. A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Chen, W. LORA: Low-Rank adaptation of Large Language Models. arXiv 2022, arXiv:2106.09685. [Google Scholar]
- OpenAI. GPT-4 Technical Report. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
- Gemini Team Google. Gemini: A Family of Highly Capable Multimodal Models. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]



| Dataset | Platform | Language | Size | Data Collection Period | Annotation Type | Session- Based | Profane Word Annotation | Abuser- Victim Identity |
|---|---|---|---|---|---|---|---|---|
| OLID [17] | English | 14,100 messages | Not mentioned | Manual (Crowdsourcing) | ✗ | ✗ | ✗ | |
| TBO [18] | English | 4673 messages | Not mentioned | Manual | ✗ | ✓ | ✓ | |
| Toxic Span Detection [19] | Civil Comments | English | 10,629 messages | 2016–2017 | Manual (Crowdsourcing) | ✗ | ✓ | ✗ |
| TOXICN [20] | Zhihu, Tieba | Chinese | 12,011 messages | Not mentioned | Manual | ✗ | ✗ | ✗ |
| SWSR [21] | Chinese | 1527 sessions | 2015–2020 | Manual | ✓ | ✓ | ✗ | |
| SCCD [22] | Chinese | 677 sessions | 2018–2024 | Machine- assisted | ✓ | ✗ | ✗ | |
| MACD [23] | ShareChat | Hindi, Tamil, Telugu, Malayalam, Kannada | 152,422 comments | 2021–2022 | Manual | ✓ | ✗ | ✗ |
| AfriHate [24] | 15 African languages | 90,437 messages | 2012–2023 | Manual | ✗ | ✗ | ✗ | |
| South African Tweets [25] | Code-mixed Afrikaans, IsiZulu, Sesotho, English | 21,350 messages | 2019 | Manual | ✗ | ✗ | ✗ | |
| L3Cube- MeHate [26] | Twitter, YouTube | Code-mixed Marathi–English | 2763 messages | 2008–2023 | Manual | ✗ | ✗ | ✗ |
| Large Transliterated Arabic Corpus [27] | Code-mixed Arabic–English | 23,587 messages | 2019–2023 | Manual, Machine-assisted | ✗ | ✗ | ✗ | |
| Indo- HateSpeech [28] | Code-mixed Hindi–English | 77,926 comments | 2019–2024 | Manual | ✓ | ✗ | ✗ | |
| PCCD (ours) | Code-mixed Cantonese–English | 1668 sessions | 2021–2023 | Manual | ✓ | ✓ | ✓ |
| Dialogue Length | Count | No. (%) of Non- Aggressive Tweets | No. (%) of Aggressive Tweets | ||
|---|---|---|---|---|---|
| 5 | 444 | 1551 | (69.9%) | 669 | (30.1%) |
| 6 | 294 | 1249 | (70.8%) | 515 | (29.2%) |
| 7 | 205 | 982 | (68.4%) | 453 | (31.6%) |
| 8 | 139 | 814 | (73.2%) | 298 | (26.8%) |
| 9 | 121 | 716 | (65.7%) | 373 | (34.3%) |
| 10 | 84 | 593 | (70.6%) | 247 | (29.4%) |
| 11 | 69 | 520 | (68.5%) | 239 | (31.5%) |
| 12 | 64 | 478 | (62.2%) | 290 | (37.8%) |
| 13 | 42 | 370 | (67.8%) | 176 | (32.2%) |
| 14 | 34 | 330 | (69.3%) | 146 | (30.7%) |
| 15 | 31 | 311 | (66.9%) | 154 | (33.1%) |
| 16 | 20 | 220 | (68.8%) | 100 | (31.2%) |
| 17 | 14 | 185 | (77.7%) | 53 | (22.3%) |
| 18 | 17 | 226 | (73.9%) | 80 | (26.1%) |
| 19 | 23 | 309 | (70.7%) | 128 | (29.3%) |
| 20 or above | 67 | 971 | (72.5%) | 369 | (27.5%) |
| Total | 1668 | 9825 | (69.6%) | 4290 | (30.4%) |
| Language | Anonymized Name |
|---|---|
| English | confident_morse |
| hardcorehamilton | |
| ElasticWilbur | |
| magical-kowalevski | |
| Chinese | 白美賢 |
| 袁蓉 |
| Category | Definition in This Paper | Wang’s Definition [30] (% in the Dataset) |
|---|---|---|
| 0 | No Aggression | No Aggression (69.6%) |
| 1 | Insult | Racism (1.0%) Religious Discrimination (<%) Other Insults (e.g., Ageism, Disability) (22.6%) |
| 2 | Sexual Harassment | Sexism (6.7%) |
| Column | Description |
|---|---|
| TweetID | The unique identifier of a tweet |
| ThreadID | The unique identifier of a thread |
| RowCount | The number of tweets belonging to a thread |
| TweetCreatedAt | The date and time of publishing of a tweet |
| UserScreenName | The username of the sender of a tweet |
| TweetText | The content of a tweet |
| UserFollowerCount | The number of followers that a tweet’s publisher has |
| TextProfane | The profane word(s) found in the content |
| Abuser | The perpetrator(s) of aggression |
| Victim | The target(s) of aggression, without considering the target type |
| Category | The aggression category of a tweet |
| IsBullying | Indicator of the presence of cyberbullying in a given thread |
| Sender | Text | Victims | Implicit/ Explicit |
|---|---|---|---|
| HappyMan | jumping_girl絕密sex影片 大家一齊睇 (Secret sex video featuring jumping_girl. Let’s watch together.) | jumping_girl | Explicit |
| lucky-bird | @HappyMan Rubbish 這樣也叫絕密? (@HappyMan Rubbish. This video could be called as secret?) | HappyMan | Explicit |
| lucky-bird | 白痴垃圾 去死 (Retarded rubbish. Die.) | HappyMan | Implicit |
| FastMonkey | @HappyMan Fuck! 狗up騙子 (@HappyMan F*ck. Swindler speaking like a dog.) | HappyMan | Explicit |
| FastMonkey | 去用條sh*t片草你自己媽逼啦 (To f*ck your mother’s c*nt with this sh*t video.) | HappyMan | Implicit |
| Category Label | Definition | % in Dataset |
|---|---|---|
| 0 | No Aggression | 69.6% |
| 1 | Insult | 23.7% |
| 2 | Sexual Harassment | 6.7% |
| Session Length | Average No. of Victims | % Implicit Victims | % Explicit Victims |
|---|---|---|---|
| 5 | 2.198 | 45.1% | 54.9% |
| 6 | 2.364 | 49.2% | 50.8% |
| 7 | 2.932 | 53.9% | 46.1% |
| 8 | 2.945 | 45.2% | 54.8% |
| 9 | 3.615 | 62.2% | 37.8% |
| 10 | 3.789 | 55.0% | 45.0% |
| 11 | 4.049 | 61.1% | 38.9% |
| 12 | 5.463 | 70.5% | 29.5% |
| 13 | 5.139 | 66.5% | 33.5% |
| 14 | 4.967 | 57.7% | 42.3% |
| 15 | 5.857 | 60.4% | 39.6% |
| 16 | 5.882 | 61.0% | 39.0% |
| 17 | 5.727 | 57.1% | 42.9% |
| 18 | 5.143 | 52.8% | 47.2% |
| 19 | 7.263 | 60.9% | 39.1% |
| 20 or above | 6.836 | 51.9% | 48.1% |
| Fold | No. of Sessions | No. of Tweets | % of Non-Aggressive | % of Aggressive |
|---|---|---|---|---|
| 1 | 562 | 4723 | 70.2% | 29.8% |
| 2 | 562 | 4699 | 69.4% | 30.6% |
| 3 | 561 | 4693 | 69.3% | 30.7% |
| Without Anonymization | With Anonymization | |||||
|---|---|---|---|---|---|---|
| Folds | Precision | Recall | F1 | Precision | Recall | F1 |
| Fold 1 | 0.481 | 0.471 | 0.476 | 0.514 | 0.480 | 0.496 |
| Fold 2 | 0.420 | 0.386 | 0.402 | 0.476 | 0.477 | 0.477 |
| Fold 3 | 0.493 | 0.424 | 0.456 | 0.505 | 0.506 | 0.506 |
| Mean | 0.465 | 0.427 | 0.445 | 0.499 | 0.488 | 0.493 |
| Precision | Recall | F1 | |
|---|---|---|---|
| p-value | 0.0583 | 0.0739 | 0.0458 |
| Cohen’s d | 1.5393 | 1.3494 | 1.7745 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chu, C.C.F.; Tong, C.C.H.; Chiu, C.H.; Chan, D.P.K.; Lam, S.C. Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data 2026, 11, 51. https://doi.org/10.3390/data11030051
Chu CCF, Tong CCH, Chiu CH, Chan DPK, Lam SC. Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data. 2026; 11(3):51. https://doi.org/10.3390/data11030051
Chicago/Turabian StyleChu, Carlin Chun Fai, Calvin Chun Ho Tong, Chun Hung Chiu, David Po Kin Chan, and Simon Ching Lam. 2026. "Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis" Data 11, no. 3: 51. https://doi.org/10.3390/data11030051
APA StyleChu, C. C. F., Tong, C. C. H., Chiu, C. H., Chan, D. P. K., & Lam, S. C. (2026). Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data, 11(3), 51. https://doi.org/10.3390/data11030051



