Next Article in Journal
Tracking K-12 and Higher Education Job Postings Through Web-Scraped Longitudinal Data
Previous Article in Journal
Frequency-Band Acoustic Feature Dataset for Comparative Analysis of Electric Vehicle Gearbox Housing Stiffness
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Data Descriptor

Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis

by
Carlin Chun Fai Chu
1,*,
Calvin Chun Ho Tong
1,
Chun Hung Chiu
2,
David Po Kin Chan
3 and
Simon Ching Lam
4
1
Department of Computer Science, School of Decision Sciences, The Hang Seng University of Hong Kong, Hong Kong SAR, China
2
School of Business, Sun Yat-sen University, Guangzhou 510275, China
3
Mailsaverse, Hong Kong SAR, China
4
School of Nursing, Tung Wah College, Hong Kong SAR, China
*
Author to whom correspondence should be addressed.
Data 2026, 11(3), 51; https://doi.org/10.3390/data11030051
Submission received: 4 December 2025 / Revised: 19 January 2026 / Accepted: 27 February 2026 / Published: 5 March 2026

Abstract

Cyberbullying behaviors manifest uniquely in different regions, shaped strongly by local slang, dialectal expressions, and cultural context. Code-mixed Chinese–English colloquial language (Cantonese) is commonly used in Hong Kong, Macau, and parts of southern China. Code-mixing is the use of multiple languages concurrently, and Cantonese text includes distinct phonetic, lexical, and syntactic features that are not exhibited in datasets developed for either Chinese or English applications. In this study, a privacy-aware code-mixed cyberbullying dataset (PCCD), containing 14,115 annotated tweets organized into 1668 sessions, was developed. Personally identifiable information and well-known identifiers, such as the names of famous celebrities, politicians, and organizations, were replaced with randomly generated dummy names. The anonymized data empirically demonstrated improved performance in terms of precision, recall, and F1 score, indicating a greater generalization ability when handling unseen participants. To the best of our knowledge, the PCCD is the first code-mixed Chinese–English dataset that includes abuser and victim identity annotation. Our dataset facilitates the development of robust cyberbullying detection tools that researchers and developers can use to accurately measure aggressiveness, attack frequency, and abuser–victim power imbalance in a dialogue session.
Dataset License: CC-BY

1. Summary

Cyberbullying is defined as repeated aggressive actions involving an imbalance in power between abusers and victims [1,2]. It has been reported to be detrimental to mental health [3], reducing students’ academic performance [4], causing depression [5], and increasing the risk of suicide [6,7]. To facilitate its timely detection and support, researchers have focused on detecting potential cyberbullying cases using machine learning models [8,9,10,11].
A properly annotated dataset is essential for developing an effective machine learning model for cyberbullying detection. Many models designed for natural language processing are trained with texts in major languages such as English, Chinese, and Spanish. Code-mixing is the use of multiple languages concurrently, and Cantonese text includes distinct phonetic, lexical, and syntactic features that are not exhibited in datasets developed for either Chinese or English applications. However, resources for code-mixed Chinese–English (Cantonese) text are rare, despite the dialect’s widespread use in Hong Kong, Macau, and parts of southern China. Cyberbullying behaviors manifest uniquely across different regions, and are shaped heavily by local slang, dialectal expressions, and cultural context. The lack of appropriate resources may have hampered the development of effective tools for identifying harmful online behaviors in Cantonese communities [12].
In this study, we present a dataset of cyberbullying messages collected from Twitter (recently rebranded as “X”), organized into annotated dialogue sessions. This session-based dataset can be used to fine-tune models for the detection of individual aggressive messages [13,14,15] as well as cyberbullying incidents [16]. The dataset contains code-mixed textual messages with Chinese and English characters, manually annotated abuser and victim identities, and anonymized user data. Our empirical study demonstrates that the data anonymization process helps to fine-tune more robust models that achieve better generalization for unseen participants.

State of the Art

According to Dinakar et al. and Van Hee et al. [1,2], cyberbullying is the occurrence of repetitive aggression between abusers and victims with a power imbalance. In other words, detecting a cyberbullying incident relies on three types of information: the aggressiveness of messages, attack frequency, and the power difference between the abuser(s) and victim(s).
In this data descriptor study, we adopted a systematic approach to identifying cyberbullying and hate speech datasets for comparison. Our team members searched for articles using Google Scholar with combinations of keywords including ‘cyberbullying’, ‘profane’, ‘aggressive’, ‘code-mixed’, ‘multilingual’, and/or ‘Chinese’. Articles related to the construction of cyberbullying datasets were located, and the corresponding datasets were inspected. In total, our team obtained 12 datasets which are currently accessible. However, it is possible that some cyberbullying datasets may be omitted.
Datasets have been developed for detecting cyberbullying in various languages, as shown in Table 1. The OLID dataset [17] contains 14,100 English messages retrieved from Twitter. It uses a hierarchical annotation of three levels, namely offensive language identification, offensive type categorization, and offensive target type identification. The dataset is annotated through crowdsourcing. TBO dataset [18] is an English dataset building on the framework of the OLID dataset. It consists of 4673 messages from Twitter with manual annotation. In addition to profane word annotation, it provides the identities of victims in harmful tweets. The annotation methodologies for many later works have been developed based on this dataset. The Toxic Span Detection dataset [19] is another English cyberbullying dataset annotated through crowdsourcing. It contains 10,629 messages from Civil Comments during 2016 and 2017, with explicit annotation of profane words. Importantly, none of these datasets provides session-based information, which is critical for detecting cyberbullying incidents.
There are annotated datasets for Chinese cyberbullying detection. The TOXICN dataset [20] contains 12,011 Chinese messages retrieved from the Chinese social media platforms Zhihu and Tieba. This manually annotated dataset consists of binary toxicity labels and three toxic expression categories containing explicitness, implicitness, and reporting. However, it lacks session-based information, profane word annotation, and abuser–victim identity annotations. The SWSR [21] and SCCD [22] datasets provide a session-based identifier for analysis. The SWSR dataset contains 1527 sessions retrieved from Weibo during the period from 2015 to 2020, while SCCD contains 677 sessions in the period 2018 to 2024. The SCCD dataset has multiple layers of annotations determining the presence of cyberbullying, whether the expressions are explicit or implicit, whether the expressions are sarcastic, whether the targets are individuals or groups, and the categories of the groups. The SWSR dataset has binary sexism labels and four categories of sexism. It also contains two target type labels: individual and generic. However, as these datasets do not provide abuser–victim identity annotation, the power difference between the abuser(s) and victim(s) cannot be examined.
Some datasets cover more than one language. The MACD dataset [23] is a non-English multilingual cyberbullying dataset. It consists of 152,422 comments written in Hindi, Tamil, Telugu, Malayalam, and Kannada. The comments are extracted from the social media platform ShareChat, which is very popular in India. The sampling period is from 2021 to 2022. The AfriHate dataset [24] is a composite hate speech dataset combining messages from Twitter and existing datasets in 15 African languages. It contains 90,437 messages collected from 2012 to 2023 and annotated manually. Messages were classified as hate, abusive, or normal. Furthermore, there are datasets that specifically contain code-mixed textual data. The South African Tweet dataset [25] contains 21,350 annotated Twitter messages written in code-mixed Afrikaans, IsiZulu, Sesotho, and English. The messages were collected in 2019 and categorized as hate speech, offensive speech, or free speech. The L3Cube-MeHate dataset [26] contains 2763 annotated messages written in Marathi and English. Data were collected between 2008 and 2023 from Twitter and YouTube. The messages were labeled as either hateful or non-hateful, depending on the presence of expressions of strongly negative feelings. The Large Transliterated Arabic Corpus dataset [27] is another Twitter dataset featuring Arabic–English content. Both manual and machine-assisted annotation methods were applied to the 23,587 messages collected in 2023. The messages were annotated as offensive or non-offensive. The Indo-HateSpeech dataset [28] is a manually annotated Hindi-English dataset containing 77,926 comments extracted from Instagram between 2019 and 2024. It contains labels representing three severity levels of hatred. Although Indo-HateSpeech and MACD provide session IDs for session-based analysis, none of the above multilingual datasets provides annotations of profane words or abuser–victim identity.
Notably, none of the listed datasets provide enough information for evaluating the aggressiveness, attack frequency, and power imbalance associated with a dialogue session. The contributions of this work are summarized below.
  • We developed a privacy-aware code-mixed cyberbullying dataset (PCCD) with session-based identifiers, profane word annotation, and abuser–victim identity annotations for examining the power imbalance among abuser(s) and victim(s). To the best of our knowledge, PCCD is the first code-mixed Cantonese dataset of this kind, which enables machine learning models to effectively evaluate the above three aspects for the detection of cyberbullying incidents.
  • We proposed and demonstrated a privacy-aware anonymization procedure that helps improve the prediction model by simplifying names semantically. It was empirically demonstrated that the use of the anonymized data improved performance in terms of precision, recall, and F1 score. This anonymization could also help improve generalization by prohibiting the model from memorizing famous usernames.
  • Our dataset contains code-mixed linguistic information and relevant toxic patterns observed in daily conversational dialogues, which should help to build a more holistic multilingual cyberbullying corpus for transfer learning.
PCCD can be considered a complement to existing cyberbullying datasets. It is suited to tasks requiring Chinese–English code-mixing, session-based, or abuser–victim information. However, other Chinese datasets, such as TOXICN [20], SWSR [21] and SCCD [22], are better suited to tasks requiring Chinese-only content, without Cantonese and English. Additionally, the MACD [23], AfriHate [24] and Indo-HateSpeech [28] datasets would be better suited if larger, non-Chinese datasets are required.

2. Methods

This section details the data collection, handling, and annotation procedures. The data collection subsection addresses the Twitter (now known as X) API program parameters. Text standardization is discussed in the data cleansing subsection. Detailed data anonymization and human annotation procedures are described in this section as well. We propose a threshold-based approach to determine the existence of cyberbullying incidents in dialogue sessions.

2.1. Data Source

The conversation dialogue threads in our dataset were collected from Twitter through its official API v1.1. The query parameter “q” was set to “*” (the wildcard character) to retrieve all tweets without filtering content by keywords. If profane words from the profanity list were used in the query, tweets without profane words would not be retrieved, which would complicate the retrieval of entire dialogues. The parameters “since_id” and “until” were set as the last retrieved “post_id” and the last day 00:00, respectively, to retrieve tweets daily. The parameter “geocode” was set to “22.324036, 114.168475, 200 km” to collect messages originating from locations within a 200 km radius of the Hong Kong Mass Transit Railway Prince Edward Station. The geo-filtering process facilitated downloading of tweets sent from regions near Hong Kong (i.e., the Pearl River Delta region of China), where a large portion of people communicate in Cantonese, a Chinese–English code-mixed language. Our research team collected 1,423,144 tweets posted between 19 December 2021 and 13 June 2023. The data were stored in MongoDB for further processing.
The first step was to select dialogues with code-mixed messages; therefore, dialogues that did not contain both English and Chinese characters were filtered out. Second, dialogues with a higher likelihood of containing aggressive content were selected using a predefined list of English and Chinese profanities. The profanity list contained 245 profane terms suggested by the annotators, of which 105 terms were in English and 140 terms were in Chinese (Cantonese). Each selected dialogue contained at least one English or Chinese profane word. Third, dialogues without the first tweet were considered incomplete and discarded. This is because the first post generally contains information crucial for understanding the conversation. Finally, dialogues with fewer than 5 tweets were removed to discard cases with low user engagement, while those with more than 20 tweets were trimmed to 20. After filtering, only 1% of tweets remained. Our proportion of cyberbullying-related tweets was consistent with the statistics from other research. For example, Yang’s work [29] found that among 1,006,434 utterances collected from Chinese live-streaming chat rooms on Twitch, a live-streaming platform, only 14,950 (1%) contained profane words.
In total, the final dataset contains 14,115 tweets organized into 1668 dialogues, which were annotated manually (see Section 2.4 for details). Each dialogue session is labeled with a boolean flag indicating the occurrence of a cyberbullying incident, and each tweet is labeled with a category field indicating the nature of its content. Table 2 details the distribution of dialogue lengths (i.e., the number of tweets in a thread) and the associated ratios of aggressive tweets. It can be observed that around 25–30% of tweets are categorized as aggressive for most dialogue lengths. Training on dialogue data of diverse lengths is expected to enhance model robustness, as authentic daily dialogues naturally vary in length.

2.2. Data Cleansing

Two types of Chinese characters—simplified and traditional—were present in the tweets. To unify the Chinese characters for further processing, simplified Chinese text was converted into traditional Chinese text. In addition, unnecessary data such as URLs and hash codes were removed from the text content, as they were considered irrelevant for content-understanding tasks.

2.3. Data Anonymization

Data anonymization is an essential process for mitigating risk to individuals, as tweet senders’ usernames may contain personally identifiable information. In this study, usernames were replaced with randomly generated dummy names to protect users’ privacy. For English names, the replacements were composite phrases defined by a random process: two words were randomly selected from lists of predefined adjectives and nouns, respectively, and were then concatenated (either consecutively or with a connecting underscore or hyphen symbol). Chinese names were anonymized using randomly generated dummy names according to the commonly used Chinese surname and given-name format, containing either two or three Chinese characters. Table 3 shows some examples of anonymized names.
In addition to username anonymization, the content of the tweets was examined, and well-known identifiers—such as the names of famous celebrities, politicians, and organizations—were manually located. They were anonymized using the technique mentioned above.

2.4. Human Annotation

Two native Cantonese-speaking, graduate-level male annotators, aged over twenty and forty, respectively, were involved in the annotation process. The retrieved data were separated into sessions, and an annotator read all the messages in a session before labeling the corresponding fields. If disagreement occurred, the two annotators discussed the case and agreed on a set of new labels.
Wang’s highly cited work [30] proposed a fine-grained approach to categorizing cyberbullying tweet contents instead of simply labeling whether the tweets were part of cyberbullying incidents or not. Due to the similarity of data sources and the intention to achieve more explicit labeling, five categories of aggression labels were initially used in our dataset. However, with this system, the occurrence rates of ‘religious discrimination’ and ‘racism’ were low, accounting for only < 0.1 % and 1.0% of all data, respectively. To avoid the existence of extremely rare minority class labels, these two labels were grouped with the ‘other insults’ label, resulting in a new ‘insult’ category. As a result, there were three possible categories for each tweet: ‘no aggression’, ‘insult’, and ‘sexual harassment’. Table 4 shows the mapping of Wang’s five categories to the three categories used in this work.
Moreover, the annotators were required to identify the abusers and victims from aggressive tweets, with some tweets having victims that might not be explicitly mentioned in the text. When messaging on social media platforms, users may not explicitly mention the people to whom they are referring, instead using pronouns to represent their target users. Users may even send messages with incomplete sentence structure for the sake of convenience and response speed.
One of the key issues when annotating session-based data is handling pronouns. People often use pronouns to refer to someone or something that they have already described, in order to avoid redundancy and make the text more readable. Annotators are required to read a complete dialogue session at a time, comprehend the conversation’s content, and then identify the corresponding abusers and victims. As shown in Figure 1, Message 1 was classified as aggressive, and the victim was labeled as ‘HappyMan’. When annotating Message 2, the annotator classified it as aggressive and realized that the pronoun ‘He’ was referring to ‘HappyMan’ from Message 1, meaning that the victim in Message 2 would, thus, also be ‘HappyMan’.
Our dataset was annotated by two native Cantonese speakers with university graduate-level education, aged around 20 and 40, respectively. The Cohen’s Kappa coefficient values for category labels, abusers, and victims were 0.967, 0.973, and 0.965, respectively. There was strong agreement between the two annotators on the annotation results. Confidence or trust scores for labels were not collected during annotation. Therefore, whenever annotators disagreed on a label, they were required to discuss it to reach consensus on the appropriate label(s).

2.5. Threshold-Based Cyberbullying Binary Label

The existence of abuser and victim annotations, together with follower counts, enables measurement of the intensity of bullying acts within a dialogue session. The literature [1,2] defines a cyberbullying incident as a repetitive act of aggression that involves a power imbalance between the victim(s) and perpetrator(s). In other words, cyberbullying can only happen if there is a victim (or victims) being attacked by multiple users with an imbalance in their social power. In this study, cyberbullying incident binary labels were assigned to dialogue sessions that fulfilled three criteria: (1) at least one victim was being attacked (aggression criteria); (2) the victim was being attacked by at least two distinct users with at least three aggressive tweets (frequency criteria); and (3) a power imbalance existed between the abusers and the victim.
User follower counts were used to estimate the existence of a power imbalance, with a power score assigned to each user according to their number of followers. Scores of 1, 2, 3, and 4 were mapped to counts below the first, second, and third quartiles and the maximum value, respectively. The power difference was defined as the sum of the power scores of all distinct abusers for a particular victim minus the victim’s power score. In this study, a power difference of at least 6 was used to indicate the presence of a cyberbullying incident. In total, 448 sessions (26.6%) in the annotated dataset indicated the presence of cyberbullying according to the defined criteria.

3. Data Description

This section describes the format and key characteristics of the PCCD dataset. Our data is organized in a session-based manner, with identifiers that allow tweets to be grouped for processing. Detailed definitions for each field are provided, and the corresponding statistics are presented in tabular form.

3.1. Session-Based Messages

The dataset contains tweets collected through the Twitter API, with each row storing the information from one tweet. Table 5 briefly describes the columns in the dataset. Each tweet in the dataset has a field called ‘TweetID’, which stores a 32-character hash code acting as a unique identifier. As our objective is to create a session-based dataset, each Tweet has another identifier field named ‘ThreadID’. For each thread, we take the ‘TweetID’ of its first tweet as the identifier for the thread, such that all tweets from the same thread have the same ‘ThreadID’, facilitating downstream processing tasks. The field ‘RowCount’ stores the number of tweets in the corresponding thread for the sake of convenience in data processing. The field ‘TweetCreatedAt’ stores the date and time that the tweet was published, using the format YYYYMMDD HH:MM:SS, and the tweets in each thread are arranged in chronological order.
The field ‘UserScreenName’ stores anonymized dummy identifiers using the procedure mentioned in Section 2.3. ‘TweetText’ stores the textual content of a tweet, and its characteristics are further elaborated in Section 3.2. A Twitter user can have other users as their followers, and the poster’s number of followers is stored in the ‘UserFollowerCount’ field. Users with more followers are considered to be more influential than users with fewer followers, and those users with the most followers could be considered key opinion leaders on social media platforms. The human-annotated fields ‘TextProfane’, ‘Abuser’, and ‘Victim’ store the profane words found in the textual content and the origin(s) and target(s) of aggression, respectively. Another annotated field, ‘Category’, stores the aggression category defined in Section 2.4, while the field ‘IsBullying’ stores the threshold-based cyberbullying binary label representing the presence of cyberbullying in the corresponding thread. A detailed definition of this label was given in Section 2.5.

3.2. Characteristics of Tweet Textual Content

The ‘TweetText’ field stores the textual content of a tweet, which can include English words, Chinese words, emojis, numbers, and punctuation marks. Due to the tradition of replying to tweets on Twitter, the text “@username” is frequently used in tweets to explicitly specify certain relevant users, which is helpful for understanding conversations and finding potential abusers and victims if aggression is present. Nevertheless, there are some cases in which the tweet sender does not specify any target users, which requires thread-based understanding to determine the users that are implicitly concerned. Table 6 shows some hypothetical tweet examples in a discussion among three users, namely, HappyMan, lucky-bird, and FastMonkey. The first tweet explicitly mentions an external user, jumping_girl, and shows potential aggression. The second message, which is from lucky-bird, explicitly expresses aggression towards HappyMan. Although the third message does not specify any users, it immediately follows the second message and is from the same user, lucky-bird. Thus, we can infer that lucky-bird is expressing aggression towards HappyMan implicitly. The fourth and fifth messages are similar cases.

3.3. Data Statistics

The ‘Category’ field stores the class of aggression found in each tweet, with three possible labels: ‘no aggression’, ‘insult’, and ‘sexual harassment’. The ratio among the three categories was 69.6%:23.7%:6.7%, and the statistics are tabulated in Table 7. It is observed that large portions of victims are implicitly mentioned across different session lengths (i.e., the minimum value is 45.1% for length = 5). Considering all the session lengths as a whole, our dataset contains 4450 victims in 1309 aggressive dialogues, of which 2014 victims are explicitly mentioned while 2436 victims are implicitly referred to. Table 8 shows the proportions of implicitly and explicitly mentioned victims in aggressive dialogues of various lengths. The ‘Average number of victims’ equals the total number of victims appearing in a particular dialogue (session) divided by the number of sessions, representing the expected number of victims mentioned in threads of a particular length. Note that, if the same user was attacked twice, the victim count was incremented by 2. On the other hand, long dialogues (i.e., those with length greater than 8) tend to have more implicitly mentioned victims, with percentages greater than 50%. We believe that the discussions in larger threads are more complicated, as people tend to use more pronouns and do not repeat specific names too often due to linguistic preferences.
The field ‘TweetCreatedAt’ contains the date and time when a message was sent. We calculated the time difference between consecutive messages and, as shown in Figure 2, the interquartile range of the time differences for messages in cyberbullying sessions was around 149 min, while that in non-cyberbullying sessions was around 233 min. The variation indicates the existence of distinct temporal patterns that may be useful for developing temporal-aware machine learning models.

3.4. Validation of Proposed Dataset Construction Approach

Machine learning models are commonly trained on raw data without removing personally identifiable information. It remains uncertain whether our proposed anonymization method will degrade the performance of machine learning tasks. To determine the existence of cyberbullying incidents, a model needs to be able to identify the victim(s) in an aggressive message effectively. The annotated tweets in our dataset were re-grouped into two classes, non-aggressive and aggressive, for a victim identification task. The aggressive class comprised tweets labeled as Category 1 or 2, while the non-aggressive class comprised those labeled as Category 0. A three-fold cross-validation analysis was conducted, splitting the dataset into three folds with approximately even numbers of sessions, tweets, and aggressive-to-non-aggressive ratios. As shown in Table 9, each fold had around 562 threads and 4700 tweets, with around 70% of tweets being non-aggressive and around 30% categorized as aggressive. Two folds were selected as training data, while the remaining fold was used as testing data.
The experiment was conducted with the Llama 3.1 8B model [31], an open-source LLM, which was fine-tuned using the rank-stabilized low-rank adaptation (rsLoRA) approach [32,33]. This approach is optimized to achieve better gradient stability in high-rank scenarios while maintaining lower computational resource requirements. Other closed-source LLMs, such as GPT-4 [34] and Gemini [35], were not selected due to their technical restrictions regarding the implementation of LoRA techniques.
LLMs have been shown to deliver unprecedented success in text processing and generation. By using LLMs with suitable prompt templates, we were able to avoid traditional text preprocessing techniques such as stop word removal and stemming. Figure 3 shows an example of running an LLM with a prompt template for victim recognition. The ‘system prompt’ is given for role-playing, describing the background of the task. The ‘input prompt’ is separated into three parts: the ‘Input Instructions’ teach the model how to digest the input text, the ‘Output Instructions’ illustrate the output format for the model, and the ‘Input Data’ contains the Twitter conversations to be processed, with each tweet as a row and the corresponding senders provided. At the end, the LLM performs category classification together with victim and abuser identification, and the responses are returned in the ‘Output Prompt’ section.
When using real names in training data, bias can be introduced as the model tries to learn and memorize the names that appear frequently as victims. Additionally, it is inevitable that some names will appear in both the training and validation data when using a real-name dataset, introducing forward-looking bias and making the predicted results overly optimistic. It is speculated that this bias would diminish the generalization ability of the model, reducing its ability to handle unseen participants. Additionally, real usernames can be semantically complex and noisy. The use of anonymous names can simplify the semantics of usernames, making it easier for the model to understand and digest the content. To eliminate bias and reduce the complexity of names, the use of an anonymized dataset is recommended. In our anonymized dataset, different dummy names are used in place of identical real names that appeared in multiple dialogue sessions, helping to eliminate forward-looking bias and simplify usernames. To evaluate the correctness of victim identification, a substring matching strategy was employed, with a predicted victim only considered correct when their name is a substring of the ground truth text (or vice versa).
Table 10 tabulates the victim identification performances of models trained using real names (without anonymization) and anonymized data (with anonymization). Precision, recall, and F1 scores were used as the metrics and calculated in an information retrieval manner. The precision score was the ratio of the number of correctly identified victims to the number of victims identified by the model. Incorrect identification of victims in non-aggressive tweets would reduce the precision score. The recall score was the ratio of the number of correctly identified victims to the total number of victims in the set. The F1 score was the harmonic mean of the precision and recall scores. The model trained with real names provided lower mean scores for precision (at 0.465), recall (at 0.427), and F1 (at 0.445). In contrast, the model trained with anonymized data obtained scores of 0.499, 0.488, and 0.493, respectively, showing improvements of 0.034 (7.31%), 0.061 (14.3%), and 0.048 (10.8%). Table 11 shows the p-values of the one-tailed paired t-test and Cohen’s d effect size statistics. The p-values for precision, recall, and F1 following the one-tailed paired t-test were 0.0583, 0.0739, and 0.0458, respectively, reflecting moderate-to-strong evidence that the anonymization process improves performance. In addition, all Cohen’s d effect size statistics were above 1, supporting the observation of strong effects in the experiment. Overall, the empirical results show that using anonymized names helped to improve the model’s generalization ability, as evidenced by better precision, recall, and F1 performance.

4. User Notes

This section addresses the key limitations of our dataset, including the effect of session trimming, potential biases, and the validity of the time span, while summarizing the main findings of the study, reflecting on their implications, and their potential applications.

4.1. Limitations

This dataset has potential limitations due to biases arising in data collection, processing, and annotation. The approach of trimming down long dialogues to their first 20 tweets may introduce bias as some information is lost, especially when bullying escalates over longer dialogues. However, we observed that discussions often shifted to unrelated hot topics in subsequent replies. Accordingly, annotators chose to retain only the first 20 tweets, which affected approximately 3% of the dialogue sessions. We believe this approach reasonably preserves the core completeness of each dialogue.
Moreover, the lack of annotator background diversity might have introduced bias during annotation. The annotators were both male, and profanity related to females, especially sexual harassment, might be missed. As the annotators shared similar cultural backgrounds, content considered aggressive in other cultures might be incorrectly classified as non-aggressive by the annotators. This annotation bias could be mitigated by having more annotators with diverse backgrounds.
Additionally, profanity-list filtering may fail to detect context-dependent forms of bullying, such as irony, sarcasm, or culturally specific slurs. As a result, overtly aggressive content tends to be over-represented, while subtler or relational forms of bullying are under-represented.
Furthermore, there are some Mandarin profanities in our dataset. Our Twitter API’s geographical parameters largely cover Cantonese-speaking populations, but they inevitably include Mandarin-speaking communities in those regions as well. As a result, it was unavoidable that both Cantonese and Mandarin profanities were present in the dialogues. Given that Cantonese and Mandarin belong to the family of Chinese languages, the existence of Mandarin content should not lead to major bias. Cantonese and English are still the major languages in the dataset.
Last but not least, Twitter updates its content moderation policies regularly, and language patterns, topics, and posting behaviors evolve accordingly. As the data were collected between December 2021 and June 2023, the relevance of the dataset could be a concern. However, the objective of the dataset was to provide code-mixed Cantonese–English content for machine learning. Apart from developing a real-time model for the Twitter platform, our dataset can be used for other applications, including early cyberbullying detection, hate speech identification, and online abuse moderation on other social media platforms, which might benefit from the inclusion of historical annotated low-resource data.

4.2. Conclusions

This paper presents a privacy-aware code-mixed cyberbullying dataset (PCCD) with user–victim identity annotations, which is the first dataset of this kind. It facilitates the development of robust cyberbullying detection models to accurately measure three aspects at a time: aggressiveness, attack frequency, and abuser–victim power imbalance. Our dataset, which consists of 14,115 annotated tweets organized into 1668 sessions, is comparable in size to other publicly available datasets for academic research. In addition, our empirical study demonstrated that data anonymization can not only mitigate the risk of exposing personally identifiable information but also help to eliminate forward-looking bias and reduce name complexity, thus improving the performance of trained models.
We hope that our fine-grained PCCD will be useful for people aiming to develop advanced methods to detect and counter cyberbullying, thus creating a more peaceful social environment. The data preparation and anonymization procedures presented in this paper may serve as a reference for building datasets for other low-resource languages.

Author Contributions

Conceptualization, C.C.F.C., C.C.H.T., C.H.C., D.P.K.C., and S.C.L.; methodology, C.C.F.C. and C.C.H.T.; software, C.C.F.C. and C.C.H.T.; validation, C.C.F.C., C.C.H.T., C.H.C., and S.C.L.; formal analysis, C.C.F.C. and C.C.H.T.; investigation, C.C.F.C. and C.C.H.T.; resources, C.C.F.C., and C.H.C.; data curation, C.C.F.C. and C.C.H.T.; writing—original draft preparation, C.C.F.C. and C.C.H.T.; writing—review and editing, C.C.F.C., C.C.H.T., C.H.C., D.P.K.C., and S.C.L.; visualization, C.C.H.T.; supervision, C.C.F.C.; project administration, C.C.F.C.; funding acquisition, C.C.F.C., C.H.C., and D.P.K.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Faculty Development Scheme of the HKSAR University Grants Committee, UGC/FDS14/E02/22.

Data Availability Statement

The privacy-aware code-mixed cyberbullying dataset (PCCD) is available at https://doi.org/10.5281/zenodo.18860813, accessed on 26 February 2026.

Acknowledgments

The authors would like to thank the senior research assistant, Ernest Kan Lam Kwong, for support in crawling raw data using the Twitter API.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Dinakar, K.; Jones, B.; Havasi, C.; Lieberman, H.; Picard, R. Common sense reasoning for detection, prevention, and mitigation of cyberbullying. ACM Trans. Interact. Intell. Syst. 2012, 2, 1–30. [Google Scholar] [CrossRef] [Scilit]
  2. Van Hee, C.; Jacobs, G.; Emmery, C.; Desmet, B.; Lefever, E.; Verhoeven, B.; De Pauw, G.; Daelemans, W.; Hoste, V. Automatic detection of cyberbullying in social media text. PLoS ONE 2018, 13, e0203794. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Kowalski, R.M.; Limber, S.P. Psychological, physical, and academic correlates of cyberbullying and traditional bullying. J. Adolesc. Health 2013, 53, S13–S20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Khine, A.T.; Saw, Y.M.; Htut, Z.Y.; Khaing, C.T.; Soe, H.Z.; Swe, K.K.; Thike, T.; Htet, H.; Saw, T.N.; Cho, S.M.; et al. Assessing risk factors and impact of cyberbullying victimization among university students in Myanmar: A cross-sectional study. PLoS ONE 2020, 15, e0227051. [Google Scholar] [CrossRef] [Scilit]
  5. Aoyama, I.; Saxon, T.F.; Fearon, D.D. Internalizing problems among cyberbullying victims and moderator effects of friendship quality. Multicult. Educ. Technol. J. 2011, 5, 92–105. [Google Scholar] [CrossRef] [Scilit]
  6. Hinduja, S.; Patchin, J.W. Bullying, cyberbullying, and suicide. Arch. Suicide Res. 2010, 14, 206–221. [Google Scholar] [CrossRef] [Scilit]
  7. Sampasa-Kanyinga, H.; Roumeliotis, P.; Xu, H. Associations between Cyberbullying and School Bullying Victimization and Suicidal Ideation, Plans and Attempts among Canadian Schoolchildren. PLoS ONE 2014, 9, e102145. [Google Scholar] [CrossRef] [Scilit]
  8. Salawu, S.; He, Y.; Lumsden, J. Approaches to Automated Detection of Cyberbullying: A survey. IEEE Trans. Affect. Comput. 2017, 11, 3–24. [Google Scholar] [CrossRef] [Scilit]
  9. Rosa, H.; Pereira, N.; Ribeiro, R.; Ferreira, P.C.; Carvalho, J.P.; Oliveira, S.; Coheur, L.; Paulino, P.; Simão, A.M.V.; Trancoso, I. Automatic cyberbullying detection: A systematic review. Comput. Hum. Behav. 2018, 93, 333–345. [Google Scholar] [CrossRef] [Scilit]
  10. Cheng, L.; Silva, Y.N.; Hall, D.; Liu, H. Session-Based Cyberbullying Detection: Problems and challenges. IEEE Internet Comput. 2020, 25, 66–72. [Google Scholar] [CrossRef] [Scilit]
  11. Yi, P.; Zubiaga, A. Session-based cyberbullying detection in social media: A survey. Online Soc. Netw. Media 2023, 36, 100250. [Google Scholar] [CrossRef] [Scilit]
  12. Maity, K.; Sen, T.; Saha, S.; Bhattacharyya, P. MTBullyGNN: A Graph neural Network-Based multitask framework for cyberbullying detection. IEEE Trans. Comput. Soc. Syst. 2022, 11, 849–858. [Google Scholar] [CrossRef] [Scilit]
  13. Jacobs, G.; Van Hee, C.; Hoste, V. Automatic classification of participant roles in cyberbullying: Can we detect victims, bullies, and bystanders in social media text? Nat. Lang. Eng. 2020, 28, 141–166. [Google Scholar] [CrossRef] [Scilit]
  14. Dewani, A.; Memon, M.A.; Bhatti, S.; Sulaiman, A.; Hamdi, M.; Alshahrani, H.; Alghamdi, A.; Shaikh, A. Detection of Cyberbullying Patterns in Low Resource Colloquial Roman Urdu Microtext using Natural Language Processing, Machine Learning, and Ensemble Techniques. Appl. Sci. 2023, 13, 2062. [Google Scholar] [CrossRef] [Scilit]
  15. Yadav, J.; Kumar, D.; Chauhan, D. Cyberbullying Detection using Pre-Trained BERT Model. In Proceedings of the 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC); IEEE: Piscataway, NJ, USA, 2020; pp. 1096–1100. [Google Scholar] [CrossRef] [Scilit]
  16. López-Vizcaíno, M.F.; Nóvoa, F.J.; Carneiro, V.; Cacheda, F. Early detection of cyberbullying on social media networks. Future Gener. Comput. Syst. 2021, 118, 219–229. [Google Scholar] [CrossRef] [Scilit]
  17. Zampieri, M.; Malmasi, S.; Nakov, P.; Rosenthal, S.; Farra, N.; Kumar, R. Predicting the Type and Target of Offensive Posts in Social Media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 1415–1420. [Google Scholar] [CrossRef] [Scilit]
  18. Zampieri, M.; Morgan, S.; North, K.; Ranasinghe, T.; Simmmons, A.; Khandelwal, P.; Rosenthal, S.; Nakov, P. Target-Based Offensive Language Identification. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 762–770. [Google Scholar] [CrossRef] [Scilit]
  19. Pavlopoulos, J.; Sorensen, J.; Laugier, L.; Androutsopoulos, I. SEMEVAL-2021 Task 5: Toxic Spans Detection. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
  20. Lu, J.; Xu, B.; Zhang, X.; Min, C.; Yang, L.; Lin, H. Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 16235–16250. [Google Scholar] [CrossRef] [Scilit]
  21. Jiang, A.; Yang, X.; Liu, Y.; Zubiaga, A. SWSR: A Chinese dataset and lexicon for online sexism detection. Online Soc. Netw. Media 2021, 27, 100182. [Google Scholar] [CrossRef] [Scilit]
  22. Yang, Q.; Chen, Y.; Xu, Z.; Shang, Y.-M.; Guo, S.; Zhang, X. SCCD: A session-based dataset for Chinese cyberbullying detection. In Proceedings of the 31st International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 9533–9545. Available online: https://aclanthology.org/2025.coling-main.639 (accessed on 26 February 2026).
  23. Gupta, V.; Roychowdhury, S.; Das, M.; Banerjee, S.; Saha, P.; Mathew, B.; Vanchinathan, H.; Mukherjee, A. MACD: Multilingual Abusive Comment Detection at Scale for Indic Languages. Adv. Neural Inf. Process. Syst. 2022, 35, 26176–26191. [Google Scholar]
  24. Muhammad, S.H.; Abdulmumin, I.; Ayele, A.A.; Adelani, D.I.; Ahmad, I.S.; Aliyu, S.M.; Röttger, P.; Oppong, A.; Bukula, A.; Chukwuneke, C.I.; et al. AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 1854–1871. [Google Scholar] [CrossRef] [Scilit]
  25. Oriola, O.; Kotze, E. Evaluating machine learning techniques for detecting offensive and hate speech in South African tweets. IEEE Access 2020, 8, 21496–21509. [Google Scholar] [CrossRef] [Scilit]
  26. Chavan, T.; Gokhale, O.; Kane, A.; Patankar, S.; Joshi, R. My Boli: Code-mixed Marathi-English Corpora, Pretrained Language Models and Evaluation Benchmarks. In Findings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 242–249. [Google Scholar] [CrossRef] [Scilit]
  27. Alhazmi, A.; Mahmud, R.; Idris, N.; Abo, M.E.M.; Eke, C.I. Code-mixing unveiled: Enhancing the hate speech detection in Arabic dialect tweets using machine learning models. PLoS ONE 2024, 19, e0305657. [Google Scholar] [CrossRef] [Scilit]
  28. Kaware, P.D.; Raut, A.B. Indo-HateSpeech Analysis: A Multi-Level hate speech classification framework using BERT features and machine learning models. Adv. Nonlinear Var. Inequalities 2025, 28, 491–503. [Google Scholar] [CrossRef] [Scilit]
  29. Yang, H.; Lin, C.J. TOCP: A Dataset for Chinese Profanity Processing. In Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying; European Language Resources Association (ELRA): Paris, France, 2020; pp. 6–12. [Google Scholar]
  30. Wang, J.; Fu, K.; Lu, C.-T. SOSNET: A Graph Convolutional Network Approach to Fine-Grained Cyberbullying Detection. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data); IEEE: Piscataway, NJ, USA, 2020; pp. 1699–1708. [Google Scholar] [CrossRef] [Scilit]
  31. Llama Team, AI @ Meta. The Llama 3 Herd of Models. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
  32. Kalajdzievski, D. A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
  33. Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Chen, W. LORA: Low-Rank adaptation of Large Language Models. arXiv 2022, arXiv:2106.09685. [Google Scholar]
  34. OpenAI. GPT-4 Technical Report. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
  35. Gemini Team Google. Gemini: A Family of Highly Capable Multimodal Models. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Illustration of handling pronouns during victim annotation (The offensive words are partially masked with the asterisk symbol *.).
Figure 1. Illustration of handling pronouns during victim annotation (The offensive words are partially masked with the asterisk symbol *.).
Data 11 00051 g001
Figure 2. Distribution of time differences between consecutive messages in cyberbullying and non-cyberbullying sessions.
Figure 2. Distribution of time differences between consecutive messages in cyberbullying and non-cyberbullying sessions.
Data 11 00051 g002
Figure 3. A prompt example for the identification of victims and perpetrators (The offensive words are partially masked with the asterisk symbol *. Translated English texts are in italics and placed within brackets.).
Figure 3. A prompt example for the identification of victims and perpetrators (The offensive words are partially masked with the asterisk symbol *. Translated English texts are in italics and placed within brackets.).
Data 11 00051 g003
Table 1. Comparison of existing cyberbullying text datasets.
Table 1. Comparison of existing cyberbullying text datasets.
DatasetPlatformLanguageSizeData
Collection
Period
Annotation
Type
Session-
Based
Profane
Word
Annotation
Abuser-
Victim
Identity
OLID [17]TwitterEnglish14,100
messages
Not
mentioned
Manual
(Crowdsourcing)
TBO [18]TwitterEnglish4673
messages
Not
mentioned
Manual
Toxic Span
Detection [19]
Civil
Comments
English10,629
messages
2016–2017Manual
(Crowdsourcing)
TOXICN [20]Zhihu,
Tieba
Chinese12,011
messages
Not
mentioned
Manual
SWSR [21]WeiboChinese1527
sessions
2015–2020Manual
SCCD [22]WeiboChinese677
sessions
2018–2024Machine-
assisted
MACD [23]ShareChatHindi, Tamil, Telugu,
Malayalam, Kannada
152,422
comments
2021–2022Manual
AfriHate [24]Twitter15 African languages90,437
messages
2012–2023Manual
South African
Tweets [25]
TwitterCode-mixed
Afrikaans, IsiZulu,
Sesotho, English
21,350
messages
2019Manual
L3Cube-
MeHate [26]
Twitter,
YouTube
Code-mixed
Marathi–English
2763
messages
2008–2023Manual
Large Transliterated
Arabic Corpus [27]
TwitterCode-mixed
Arabic–English
23,587
messages
2019–2023Manual,
Machine-assisted
Indo-
HateSpeech [28]
InstagramCode-mixed
Hindi–English
77,926
comments
2019–2024Manual
PCCD (ours)TwitterCode-mixed
Cantonese–English
1668
sessions
2021–2023Manual
Table 2. Distribution of dialogue length and ratios of aggressive tweets.
Table 2. Distribution of dialogue length and ratios of aggressive tweets.
Dialogue LengthCountNo. (%) of Non-
Aggressive Tweets
No. (%) of
Aggressive Tweets
54441551(69.9%)669(30.1%)
62941249(70.8%)515(29.2%)
7205982(68.4%)453(31.6%)
8139814(73.2%)298(26.8%)
9121716(65.7%)373(34.3%)
1084593(70.6%)247(29.4%)
1169520(68.5%)239(31.5%)
1264478(62.2%)290(37.8%)
1342370(67.8%)176(32.2%)
1434330(69.3%)146(30.7%)
1531311(66.9%)154(33.1%)
1620220(68.8%)100(31.2%)
1714185(77.7%)53(22.3%)
1817226(73.9%)80(26.1%)
1923309(70.7%)128(29.3%)
20 or above67971(72.5%)369(27.5%)
Total16689825(69.6%)4290(30.4%)
Table 3. Examples of anonymized names.
Table 3. Examples of anonymized names.
LanguageAnonymized Name
Englishconfident_morse
hardcorehamilton
ElasticWilbur
magical-kowalevski
Chinese白美賢
袁蓉
Table 4. Comparison of tweet category definitions.
Table 4. Comparison of tweet category definitions.
CategoryDefinition in This PaperWang’s Definition [30] (% in the Dataset)
0No AggressionNo Aggression (69.6%)
1InsultRacism (1.0%)
Religious Discrimination (< 0.1 %)
Other Insults (e.g., Ageism, Disability) (22.6%)
2Sexual HarassmentSexism (6.7%)
Table 5. Descriptions of the presented dataset’s columns.
Table 5. Descriptions of the presented dataset’s columns.
ColumnDescription
TweetIDThe unique identifier of a tweet
ThreadIDThe unique identifier of a thread
RowCountThe number of tweets belonging to a thread
TweetCreatedAtThe date and time of publishing of a tweet
UserScreenNameThe username of the sender of a tweet
TweetTextThe content of a tweet
UserFollowerCountThe number of followers that a tweet’s publisher has
TextProfaneThe profane word(s) found in the content
AbuserThe perpetrator(s) of aggression
VictimThe target(s) of aggression, without considering the target type
CategoryThe aggression category of a tweet
IsBullyingIndicator of the presence of cyberbullying in a given thread
Table 6. An example dialogue session.
Table 6. An example dialogue session.
SenderTextVictimsImplicit/
Explicit
HappyManjumping_girl絕密sex影片Data 11 00051 i001大家一齊睇Data 11 00051 i002 (Secret sex video featuring jumping_girl. Let’s watch together.)jumping_girlExplicit
lucky-bird@HappyMan Rubbish 這樣也叫絕密? (@HappyMan Rubbish. This video could be called as secret?)HappyManExplicit
lucky-bird白痴垃圾Data 11 00051 i001去死 (Retarded rubbish. Die.)HappyManImplicit
FastMonkey@HappyMan Fuck! 狗up騙子 (@HappyMan F*ck. Swindler speaking like a dog.)HappyManExplicit
FastMonkey去用條sh*t片草你自己媽逼啦 (To f*ck your mother’s c*nt with this sh*t video.)HappyManImplicit
The offensive words are partially masked with the asterisk symbol *. English translated texts are in italics and placed within brackets.
Table 7. Tweet aggression categories.
Table 7. Tweet aggression categories.
Category LabelDefinition% in Dataset
0No Aggression69.6%
1Insult23.7%
2Sexual Harassment6.7%
Table 8. Distribution of the average number of victims mentioned implicitly and explicitly in aggressive dialogues of various lengths.
Table 8. Distribution of the average number of victims mentioned implicitly and explicitly in aggressive dialogues of various lengths.
Session LengthAverage No. of Victims% Implicit Victims% Explicit Victims
52.19845.1%54.9%
62.36449.2%50.8%
72.93253.9%46.1%
82.94545.2%54.8%
93.61562.2%37.8%
103.78955.0%45.0%
114.04961.1%38.9%
125.46370.5%29.5%
135.13966.5%33.5%
144.96757.7%42.3%
155.85760.4%39.6%
165.88261.0%39.0%
175.72757.1%42.9%
185.14352.8%47.2%
197.26360.9%39.1%
20 or above6.83651.9%48.1%
Table 9. Dataset fold statistics.
Table 9. Dataset fold statistics.
FoldNo. of SessionsNo. of Tweets% of Non-Aggressive% of Aggressive
1562472370.2%29.8%
2562469969.4%30.6%
3561469369.3%30.7%
Table 10. Results of victim identification without vs. with data anonymization.
Table 10. Results of victim identification without vs. with data anonymization.
Without AnonymizationWith Anonymization
FoldsPrecisionRecallF1PrecisionRecallF1
Fold 10.4810.4710.4760.5140.4800.496
Fold 20.4200.3860.4020.4760.4770.477
Fold 30.4930.4240.4560.5050.5060.506
Mean0.4650.4270.4450.4990.4880.493
Table 11. p-value of one-tailed paired t-test and Cohen’s d effect size statistics.
Table 11. p-value of one-tailed paired t-test and Cohen’s d effect size statistics.
PrecisionRecallF1
p-value0.05830.07390.0458
Cohen’s d1.53931.34941.7745
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chu, C.C.F.; Tong, C.C.H.; Chiu, C.H.; Chan, D.P.K.; Lam, S.C. Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data 2026, 11, 51. https://doi.org/10.3390/data11030051

AMA Style

Chu CCF, Tong CCH, Chiu CH, Chan DPK, Lam SC. Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data. 2026; 11(3):51. https://doi.org/10.3390/data11030051

Chicago/Turabian Style

Chu, Carlin Chun Fai, Calvin Chun Ho Tong, Chun Hung Chiu, David Po Kin Chan, and Simon Ching Lam. 2026. "Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis" Data 11, no. 3: 51. https://doi.org/10.3390/data11030051

APA Style

Chu, C. C. F., Tong, C. C. H., Chiu, C. H., Chan, D. P. K., & Lam, S. C. (2026). Privacy-Aware Code-Mixed Cyberbullying Dataset for Session-Based Analysis. Data, 11(3), 51. https://doi.org/10.3390/data11030051

Article Metrics

Back to TopTop