Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research
Abstract
1. Introduction
1.1. Literature Review
1.1.1. Computational Language Models: Background
1.1.2. Pretrained Language Models (PLMs)
1.1.3. Large Language Models (LLMs)
1.1.4. Evaluation Framework: Validity, Reliability, and Transferability
2. Materials and Methods
2.1. Model Selection
2.2. Training Dataset Selection
2.3. Documentation Extraction and Coding (RQ1–RQ2)
2.4. Cross-Dataset Testing
2.5. GPT-5 Benchmark (LLM Prompting)
3. Results
3.1. Additional Training Datasets
3.2. Hate Speech Detection Models
3.3. Threat Detection Models
3.4. Bully Detection Models
3.5. Offensive Speech Detection Models
3.6. Racism Detection Models
4. Discussion
4.1. Implications for Researchers, Reviewers, and Editors
4.2. Limitations and Further Research
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Mathew, B.; Saha, P.; Yimam, S.M.; Biemann, C.; Goyal, P.; Mukherjee, A. HateXplain: A benchmark dataset for explainable hate speech detection with annotated rationales. Proc. AAAI Conf. Artif. Intell. 2021, 35, 14867–14875. [Google Scholar] [CrossRef] [Scilit]
- Schmidt, A.; Wiegand, M. A survey on hate speech detection using natural language processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, Valencia, Spain, 3 April 2017; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
- Kshetri, N.; Carter, W.; Kern, S.; Mensah, R.; Pokharel, B.P. hateUS—Analysis, impact of social media use and hate speech over university student platforms: Case study, problems, and solutions. arXiv 2024, arXiv:2410.20070. [Google Scholar] [CrossRef] [Scilit]
- Müller, K.; Schwarz, C. Fanning the flames of hate: Social media and hate crime. J. Eur. Econ. Assoc. 2021, 19, 2131–2167. [Google Scholar] [CrossRef] [Scilit]
- Olteanu, A.; Castillo, C.; Boy, J.; Varshney, K.R. The effect of extremist violence on hateful speech online. In Proceedings of the International AAAI Conference on Web and Social Media, 25–28 June 2018, Palo Alto, CA, USA; Volume 12, pp. 221–230. [CrossRef] [Scilit]
- Ramos, G.; Batista, F.; Ribeiro, R.; Fialho, P.; Moro, S.; Fonseca, A.; Guerra, R.; Carvalho, P.; Marques, C.; Silva, C. A comprehensive review on automatic hate speech detection in the age of the transformer. Soc. Netw. Anal. Min. 2024, 14, 204. [Google Scholar] [CrossRef] [Scilit]
- Stoll, A.; Yu, J.; Andrich, A.; Domahidi, E. Classification bias of LLMs in detecting incivility towards female and male politicians in German social media discourse. Commun. Methods Meas. 2025, 19, 350–368. [Google Scholar] [CrossRef] [Scilit]
- Birkenmaier, L.; Lechner, C.M.; Wagner, C. The search for solid ground in text as data: A systematic review of validation practices and practical recommendations for validation. Commun. Methods Meas. 2024, 18, 249–277. [Google Scholar] [CrossRef] [Scilit]
- Bernhard-Harrer, J.; Ashour, R.; Eberl, J.-M.; Tolochko, P.; Boomgaarden, H. Beyond standardization: A comprehensive review of topic modeling validation methods for computational social science research. Political Sci. Res. Methods 2025, 1–19. [Google Scholar] [CrossRef] [Scilit]
- Eisele, O.; Heidenreich, T.; Litvyak, O.; Boomgaarden, H.G. Capturing a news frame—Comparing machine-learning approaches to frame analysis with different degrees of supervision. Commun. Methods Meas. 2023, 17, 205–226. [Google Scholar] [CrossRef] [Scilit]
- Van Atteveldt, W.; van der Velden, M.A.C.G.; Boukes, M. The validity of sentiment analysis: Comparing manual annotation, crowd-coding, dictionary approaches, and machine learning algorithms. Commun. Methods Meas. 2021, 15, 121–140. [Google Scholar] [CrossRef] [Scilit]
- Bender, E.M.; Friedman, B. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Trans. Assoc. Comput. Linguist. 2018, 6, 587–604. [Google Scholar] [CrossRef] [Scilit]
- Malik, J.S.; Qiao, H.; Pang, G.; van den Hengel, A. Deep learning for hate speech detection: A comparative study. Int. J. Data Sci. Anal. 2025, 20, 3053–3068. [Google Scholar] [CrossRef] [Scilit]
- Radford, J.; Joseph, K. Theory in, theory out: The uses of social theory in machine learning for social science. Front. Big Data 2020, 3, 18. [Google Scholar] [CrossRef] [Scilit]
- Galke, L.; Scherp, A.; Diera, A.; Karl, F.; Lin, B.X.; Khera, B.; Meuser, T.; Singhal, T. Are we really making much progress in text classification? A comparative review. arXiv 2022, arXiv:2204.03954. [Google Scholar] [CrossRef] [Scilit]
- Poletto, F.; Basile, V.; Sanguinetti, M.; Bosco, C.; Patti, V. Resources and benchmark corpora for hate speech detection: A systematic review. Lang. Resour. Eval. 2021, 55, 477–523. [Google Scholar] [CrossRef] [Scilit]
- Fortuna, P.; Nunes, S. A survey on automatic detection of hate speech in text. ACM Comput. Surv. 2018, 51, 85. [Google Scholar] [CrossRef] [Scilit]
- Osborne, C.; Ding, J.; Kirk, H.R. The AI community building the future? A quantitative analysis of development activity on Hugging Face Hub. J. Comput. Soc. Sci. 2024, 7, 2067–2105. [Google Scholar] [CrossRef] [Scilit]
- Fortuna, P.; Soler, J.; Wanner, L. How well do hate speech, toxicity, abusive and offensive language classification models generalize across datasets? Inf. Process. Manag. 2021, 58, 102524. [Google Scholar] [CrossRef] [Scilit]
- Berelson, B. Content Analysis in Communication Research; Free Press: Glencoe, IL, USA, 1952. [Google Scholar] [CrossRef] [Scilit]
- Krippendorff, K. Content Analysis: An Introduction to Its Methodology, 2nd ed.; Sage Publications: Thousand Oaks, CA, USA, 2004. [Google Scholar] [CrossRef] [Scilit]
- Hase, V.; Bachl, M.; TeBlunthuis, N. Critical, but constructive: Defining, detecting, and addressing bias in computational social science. Commun. Methods Meas. 2025, 19, 281–293. [Google Scholar] [CrossRef] [Scilit]
- Ross, B.; Rist, M.; Carbonell, G.; Cabrera, B.; Kurowsky, N.; Wojatzki, M. Measuring the reliability of hate speech annotations: The case of the European refugee crisis. In NLP4CMC III: 3rd Workshop on Natural Language Processing for Computer-Mediated Communication; Dipper, S., Ed.; Bochumer Linguistische Arbeitsberichte: Bochum, Germany, 2016; Volume 17, pp. 6–9. [Google Scholar] [CrossRef]
- Fussell, R.K.; Mazrui, A.; Holmes, N.G. Machine learning for automated content analysis: Characteristics of training data impact reliability. In Proceedings of the 2022 Physics Education Research Conference (PERC 2022), Grand Rapids, MI, USA, 13–14 July 2022; Frank, B.W., Jones, D.L., Ryan, Q.X., Eds.; American Association of Physics Teachers: College Park, MD, USA, 2022; pp. 176–181. [Google Scholar] [CrossRef] [Scilit]
- Lavelle-Hill, R.; Smith, G.; Murayama, K. Bridging traditional-statistics and machine-learning approaches in psychology: Navigating small samples, measurement error, nonindependent observations, and missing data. Adv. Methods Pract. Psychol. Sci. 2025, 8, 1–31. [Google Scholar] [CrossRef] [Scilit]
- Davidson, T.; Warmsley, D.; Macy, M.; Weber, I. Automated hate speech detection and the problem of offensive language. In Proceedings of the International AAAI Conference on Web and Social Media, 15–18 May 2017, Montreal, QC, Canada; 2017; Volume 11, pp. 512–515. [Google Scholar] [CrossRef] [Scilit]
- Yu, Z.; Sen, I.; Assenmacher, D.; Samory, M.; Fröhling, L.; Dahn, C.; Nozza, D.; Wagner, C. The unseen targets of hate: A systematic review of hateful communication datasets. Soc. Sci. Comput. Rev. 2025, 43, 965–989. [Google Scholar] [CrossRef] [Scilit]
- Verma, K.; Milosevic, T.; Cortis, K.; Davis, B. Benchmarking language models for cyberbullying identification and classification from social-media texts. In Proceedings of the First Workshop on Language Technology and Resources for a Fair, Inclusive, and Safe Society within the 13th Language Resources and Evaluation Conference; Adebayo, K., Nanda, R., Verma, K., Davis, B., Eds.; European Language Resources Association: Marseille, France, 2022; pp. 26–31. Available online: https://aclanthology.org/2022.lateraisse-1.4/ (accessed on 15 April 2026).
- Mansur, Z.; Omar, N.; Tiun, S. Twitter hate speech detection: A systematic review of methods, taxonomy analysis, challenges, and opportunities. IEEE Access 2023, 11, 16226–16249. [Google Scholar] [CrossRef] [Scilit]
- Burscher, B.; Vliegenthart, R.; de Vreese, C.H. Using supervised machine learning to code policy issues: Can classifiers generalize across contexts? Ann. Am. Acad. Political Soc. Sci. 2015, 659, 122–130. [Google Scholar] [CrossRef] [Scilit]
- Keum, B.T.; Miller, M.J. Racism on the internet: Conceptualization and recommendations for research. Psychol. Violence 2018, 8, 782–791. [Google Scholar] [CrossRef] [Scilit]
- Cinelli, M.; De Francisci Morales, G.; Galeazzi, A.; Quattrociocchi, W.; Starnini, M. The echo chamber effect on social media. Proc. Natl. Acad. Sci. USA 2021, 118, e2023301118. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dutton, W.H.; Reisdorf, B.C. Cultural divides and digital inequalities: Attitudes shaping internet and social media divides. Inf. Commun. Soc. 2019, 22, 18–38. [Google Scholar] [CrossRef] [Scilit]
- Aytac, U. Digital domination: Social media and contestatory democracy. Political Stud. 2024, 72, 6–25. [Google Scholar] [CrossRef] [Scilit]
- Schmid, U.K. Humorous hate speech on social media: A mixed-methods investigation of users’ perceptions and processing of hateful memes. New Media Soc. 2025, 27, 1588–1606. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Burstein, J., Doran, C., Solorio, T., Eds.; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
- Waseem, Z.; Hovy, D. Hateful symbols or hateful people? Predictive features for hate speech detection on Twitter. In Proceedings of the NAACL Student Research Workshop; Andreas, J., Choi, E., Lazaridou, A., Eds.; Association for Computational Linguistics: San Diego, CA, USA, 2016; pp. 88–93. [Google Scholar] [CrossRef] [Scilit]
- Founta, A.; Djouvas, C.; Chatzakou, D.; Leontiadis, I.; Blackburn, J.; Stringhini, G.; Vakali, A.; Sirivianos, M.; Kourtellis, N. Large scale crowdsourcing and characterization of Twitter abusive behavior. In Proceedings of the International AAAI Conference on Web and Social Media, 25–28 June 2018, Palo Alto, CA, USA; Volume 12, pp. 491–500. [CrossRef] [Scilit]
- Wiegand, M.; Ruppenhofer, J.; Kleinbauer, T. Detection of abusive language: The problem of biased datasets. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Burstein, J., Doran, C., Solorio, T., Eds.; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; pp. 602–608. [Google Scholar] [CrossRef] [Scilit]
- Kennedy, B.; Atari, M.; Davani, A.M.; Yeh, L.; Omrani, A.; Kim, Y.; Coombs, K., Jr.; Havaldar, S.; Portillo-Wightman, G.; Gonzalez, E.; et al. Introducing the Gab Hate Corpus: Defining and applying hate-based rhetoric to social media posts at scale. Lang. Resour. Eval. 2022, 56, 79–108. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Fu, K.; Lu, C.T. SOSNet: A graph convolutional network approach to fine-grained cyberbullying detection. In Proceedings of the 2020 IEEE International Conference on Big Data (Big Data), Atlanta, GA, USA, 10–13 December 2020; pp. 1212–1220. [Google Scholar] [CrossRef] [Scilit]
- Grimmer, J.; Stewart, B.M. Text as data: The promise and pitfalls of automatic content analysis methods for political texts. Political Anal. 2013, 21, 267–297. [Google Scholar] [CrossRef] [Scilit]
- Baden, C.; Pipal, C.; Schoonvelde, M.; van der Velden, M.A.C.G. Three gaps in computational text analysis methods for social sciences: A research agenda. Commun. Methods Meas. 2022, 16, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Laurer, M.; van Atteveldt, W.; Casas, A.; Welbers, K. On measurement validity and language models: Increasing validity and decreasing bias with instructions. Commun. Methods Meas. 2025, 19, 46–62. [Google Scholar] [CrossRef] [Scilit]
- Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
- Koo, T.K.; Li, M.Y. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J. Chiropr. Med. 2016, 15, 155–163, Erratum in J. Chiropr. Med. 2017, 16, 346.. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elsafoury, F.; Katsigiannis, S.; Pervez, Z.; Ramzan, N. When the timeline meets the pipeline: A survey on automated cyberbullying detection. IEEE Access 2021, 9, 103541–103563. [Google Scholar] [CrossRef] [Scilit]
- Waseem, Z. Are you a racist or am I seeing things? Annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science; Bamman, D., Doğruöz, A.S., Eisenstein, J., Hovy, D., Jurgens, D., O’Connor, B., Oh, A., Tsur, O., Volkova, S., Eds.; Association for Computational Linguistics: Austin, TX, USA, 2016; pp. 138–142. [Google Scholar] [CrossRef] [Scilit]
- Vidgen, B.; Derczynski, L. Directions in abusive language training data, a systematic review: Garbage in, garbage out. PLoS ONE 2020, 15, e0243300. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ljubešić, N.; Mozetič, I.; Kralj Novak, P. Quantifying the impact of context on the quality of manual hate speech annotation. Nat. Lang. Eng. 2023, 29, 1481–1494. [Google Scholar] [CrossRef] [Scilit]
- Kovács, G.; Alonso, P.; Saini, R. Challenges of hate speech detection in social media. SN Comput. Sci. 2021, 2, 95. [Google Scholar] [CrossRef] [Scilit]
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 2022, 35, 27730–27744. Available online: https://proceedings.neurips.cc/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html (accessed on 15 April 2026).
- TeBlunthuis, N.; Hase, V.; Chan, C.H. Misclassification in automated content analysis causes bias in regression. Can we fix it? Yes we can! Commun. Methods Meas. 2024, 18, 278–299. [Google Scholar] [CrossRef] [Scilit]

| Model Name | Dataset Name | Dataset Information | Annotator Guidelines | Inter-Annotator Agreement ** | Final Label Decision | |
|---|---|---|---|---|---|---|
| Hate | hatexplain | HateXplain | 250 Amazon MTurk 20 k tweets and GAB comments | Definition examples | K = 0.46 (not acceptable) | Majority |
| facebook_roberta | DynaHate | Not mentioned 40 k tweets and comments | Definition | K = 0.52 (not acceptable) | Majority | |
| dehatebert_ non-English | TDavidson | Appen 25 k tweets | Definition | A = 0.92 (reliable) | Majority | |
| distilroberta_of-fensive_hate | Badmatr11x | Not mentioned 57 k tweets | Not mentioned | Not mentioned | Not mentioned | |
| roberta_large_hate | TweetHate | Amazon MTurk 7 k tweets | Not mentioned | K = 0.53 (not acceptable) | Agreement of at least 2 of 5 annotators | |
| Racism | deberta | TDavidson | Appen 25 k tweets | Definition | A = 0.92 (reliable) | Majority |
| Bullying | lstm_bert cyberbully-ing_bert | Kaggle Cyberbullying Dataset | Not mentioned 47 k tweets | Not mentioned | Not mentioned | Not mentioned |
| Offensive | distilroberta_olid twitter_roberta offensive_speech_ detector | OLID | Appen 14,100 tweets | Test Set | F = 0.83 (al- most perfect) | Majority |
| hatexplain | HateXplain | 250 Amazon MTurk 20 k tweets and GAB comments | Definition examples | K = 0.46 (not acceptable) | Majority | |
| distilroberta_ offensive_hate | Badmatr11x | Not mentioned 57 k tweets | Not mentioned | Not mentioned | Not mentioned | |
| llama | TDavidson | Appen 25 k tweets | Definition | A = 0.92 (reliable) | Majority | |
| Threat | ensemble lstm_glove nb_svm lstm_rnn detoxify | Toxic Comment Classification | Not mentioned 160 k tweets | Not mentioned | Not mentioned | Not mentioned |
| detoxify_unbiased | Jigsaw Unintended Toxicity Bias | Appen 1.05 M comments | Not mentioned | Not mentioned | Majority |
| Label | Dataset Name | Dataset & Annotators Information | Annotator Guidelines | Inter-annotator Agreement * | Final Label Decision |
|---|---|---|---|---|---|
| Hate | Large-Scale Hate Speech | 20 Student annotators 68.5 k tweets | Definition Examples | K = 0.395 (not acceptable) | Unanimous |
| Implicit Hate Corpus | Amazon MTurk 19 k tweets | Definition Examples Test | ICC = 0.616 (moderate) | Majority | |
| New-Wave-Hate | Authors (Amazon MTurk -> authors) | Definition Codebook | F = 0.84 (almost perfect) | Single | |
| Racism | BiasCorp | Amazon MTurk 45 k comments | Not mentioned | Not mentioned | Not mentioned |
| Are You Racist, Or Am I Seeing Things | Appen, Experts 7 k tweets | Test | F = 0.57 (moderate) | Majority | |
| MMHS150k | Amazon MTurk annotators 150 k tweets | Definition Examples | Not mentioned | Majority | |
| Twitter Racism Dataset | Not mentioned 13.5 k tweets | Not mentioned | Not mentioned | Not mentioned | |
| Bully | Cyberbullying Twitter | 17 annotators personally contacted 62.5 k tweets | Not mentioned | K = 0.67 (tentatively acceptable) | Majority |
| Online Harassment Dataset | Not mentioned. 8.2 k tweets | Not mentioned | Not mentioned | Not mentioned | |
| Kaggle Insults Dataset | Not mentioned 2 k comments | Not mentioned | Not mentioned | Not mentioned | |
| Bullying Traces Dataset | Not mentioned 1.7 k tweets | Not mentioned | Not mentioned | Not mentioned | |
| Automated Cyberbullying Detection | Pre-labeled. 15 k messages | Definition Codebook | Not mentioned | Threshold | |
| Online Harassment Research | Authors 35 k tweets | Definition | C = 0.84 (almost perfect) | Majority | |
| Offensive | OffensiveLang | Amazon MTurk annotators 8270 posts generated by ChatGPT | Definition | C = 0.54 (moderate) | Majority |
| Multimodal Meme Dataset | Not mentioned 743 meme content | Examples Codebook | F = 0.4–0.5 (moderate) | Single | |
| HASOC | Authors 6 k posts | Definition | % of agreement: 72% (acceptable) | Not mentioned | |
| Large-Scale Hate Speech | 20 Student annotators 68.5 k tweets | Definition Examples | K = 0.395 (not acceptable) | Unanimous | |
| Threat | Combined Toxicity Profanity | Not mentioned 710 k comments | Not mentioned | Not mentioned | Not mentioned. Comments can have multiple labels. |
| Suspicious Tweets | Not mentioned 59 k tweets | Not mentioned | Not mentioned | Not mentioned |
| Dataset Model | Dyna-Hate | HateX-Plain | Badmatr11x | TweetHate | TDavidson | Large-Scale Hate Speech | Implicit Hate Corpus | New-Wave-Hate | Average |
|---|---|---|---|---|---|---|---|---|---|
| 0.75 | 0.55 | 0.69 | 0.63 | ||||||
| facebook_roberta | (0.76) | (0.55) | 0.62 (0.71) | 0.73 (0.74) | 0.25 (0.33) | 0.85 (0.90) | 0.62 (0.59) | (0.68) | (0.66) |
| 0.54 | 0.84 | 0.60 | 0.76 | ||||||
| hatexplain | (0.43) | (0.84) | 0.92 (0.90) | 0.73 (0.63) | 0.93 (0.92) | 0.98 (0.97) | 0.51 (0.37) | (0.47) | (0.70) |
| distilroberta_offen- | 0.56 | 0.65 | 0.63 | 0.74 | |||||
| sive_hate | (0.50) | (0.65) | 0.96 (0.96) | 0.69 (0.66) | 0.97 (0.97) | 0.91 (0.94) | 0.53 (0.52) | (0.57) | (0.72) |
| 0.56 | 0.70 | 0.72 | 0.73 | ||||||
| roberta_large_hate | (0.50) | (0.70) | 0.77 (0.82) | 0.98 (0.98) | 0.59 (0.69) | 0.96 (0.96) | 0.56 (0.48) | (0.68) | (0.73) |
| dehate- | 0.60 | 0.55 | 0.57 | 0.62 | |||||
| bert_mono_english | (0.57) | (0.56) | 0.63 (0.72) | 0.71 (0.72) | 0.76 (0.73) | 0.90 (0.93) | 0.60 (0.55) | (0.57) | (0.64) |
| 0.56 | 0.54 | 0.60 | 0.61 | ||||||
| hatebert | (0.50) | (0.52) | 0.50 (0.46) | 0.75 (0.71) | 0.39 (0.39) | 0.94 (0.56) | 0.60 (0.55) | (0.58) | (0.53) |
| 0.50 | 0.50 | 0.54 | 0.56 | ||||||
| distilbert_hate | (0.49) | (0.47) | 0.59 (0.42) | 0.57 (0.50) | 0.67 (0.42) | 0.71 (0.43) | 0.48 (0.45) | (0.50) | (0.46) |
| 0.81 | 0.79 | 0.74 | 0.75 | ||||||
| ChatGPT-5 | (0.81) | (0.79) | 0.69 (0.66) | 0.69 (0.64) | 0.89 (0.85) | 0.64 (0.58) | 0.73 (0.73) | (0.70) | (0.72) |
| Dataset Model | Toxic Comments Dataset | Jigsaw Unintended Toxicity Bias | Combined Toxicity Profanity | Suspicious Tweets | Average |
|---|---|---|---|---|---|
| ensemble | 0.54 (0.42) | 0.60 (0.01) | 0.61 (0.03) | 0.60 (<0.01) | 0.59 (0.11) |
| lstm_glove | 0.90 (0.90) | 0.70 (0.43) | 0.68 (0.49) | 0.62 (0.25) | 0.73 (0.52) |
| nb_svm | 0.89 (0.88) | 0.60 (0.02) | 0.65 (0.22) | 0.60 (<0.01) | 0.69 (0.28) |
| bert_threatening | 0.81 (0.76) | 0.62 (0.08) | 0.66 (0.26) | 0.60 (<0.01) | 0.67 (0.28) |
| lstm_rnn | 0.75 (0.67) | 0.63 (0.14) | 0.66 (0.27) | 0.60 (<0.01) | 0.66 (0.27) |
| detoxify | 0.95 (0.95) | 0.58 (0.29) | 0.97 (0.03) | <0.001 (<0.001) | 0.63 (0.32) |
| detoxify_unbi-ased | 0.82 (0.78) | 0.76 (0.68) | 0.80 (0.76) | <0.001 (<0.001) | 0.60 (0.56) |
| ChatGPT-5 | 0.92 (0.92) | 0.72 (0.64) | 0.75 (0.70) | 0.60 (0.38) | 0.75 (0.72) |
| Model Dataset | Kaggle Cyberbullying Dataset | Automated Cyberbullying Detection | Bullying Traces Dataset | Online Harassment Dataset | Cyberbul- Lying | Kaggle Insults Dataset | Online Harassment Corpus | Average |
|---|---|---|---|---|---|---|---|---|
| lstm_bert | 0.95 (0.95) | 0.65 (0.63) | 0.75 (0.54) | 0.93 (0.64) | 0.53 (0.40) | 0.64 (0.57) | 0.47 (0.46) | 0.70 (0.60) |
| fair_cyber-bully | 0.61 (0.55) | 0.81 (0.76) | 0.64 (0.61) | 0.80 (0.56) | 0.47 (0.38) | 0.72 (0.69) | 0.53 (0.51) | 0.65 (0.58) |
| distilbert_ cyberbully | 0.40 (0.34) | 0.46 (0.44) | 0.41 (0.41) | 0.24 (0.21) | 0.45 (0.35) | 0.44 (0.37) | 0.42 (0.41) | 0.40 (0.36) |
| cyberbully-ing_bert | 0.86 (0.86) | 0.64 (0.48) | 0.69 (0.50) | 0.80 (0.50) | 0.72 (0.47) | 0.55 (0.54) | 0.69 (0.48) | 0.71 (0.55) |
| ChatGPT-5 | 0.50 (0.48) | 0.83 (0.80) | 0.54 (0.48) | 0.96 (0.79) | 0.64 (0.48) | 0.81 (0.80) | 0.54 (0.53) | 0.75 (0.66) |
| Dataset Model | OLID | Offensive-Lang | TDavidson | HASOC | Large-Scale Hate Speech | Multimodal Meme Dataset | Average |
|---|---|---|---|---|---|---|---|
| arash_bert | 0.90 (0.90) | 0.30 (0.29) | 0.79 (0.72) | 0.59 (0.48) | 0.86 (0.68) | 0.61 (0.56) | 0.67 (0.60) |
| distilroberta_olid | 0.86 (0.77) | 0.26 (0.26) | 0.79 (0.72) | 0.68 (0.51) | 0.89 (0.83) | 0.61 (0.56) | 0.68 (0.61) |
| twitter_roberta | 0.86 (0.78) | 0.28 (0.26) | 0.87 (0.80) | 0.62 (0.49) | 0.88 (0.82) | 0.60 (0.54) | 0.69 (0.61) |
| offensive_speech_de-tector | 0.75 (0.71) | 0.26 (0.26) | 0.75 (0.83) | 0.70 (0.52) | 0.91 (0.85) | 0.61 (0.52) | 0.66 (0.61) |
| hatexplain | 0.75 (0.68) | 0.28 (0.19) | 0.88 (0.93) | 0.87 (0.53) | 0.98 (0.97) | 0.59 (0.40) | 0.73 (0.62) |
| llama | 0.73 (0.66) | 0.27 (0.26) | 0.90 (0.88) | 0.68 (0.50) | 0.87 (0.76) | 0.62 (0.60) | 0.68 (0.61) |
| distilroberta_ offensive_hate | 0.73 (0.61) | 0.22 (0.19) | 0.97 (0.95) | 0.82 (0.42) | 0.89 (0.78) | 0.59 (0.38) | 0.70 (0.57) |
| ChatGPT-5 | 0.72 (0.61) | 0.48 (0.47) | 0.84 (0.83) | 0.51 (0.51) | 0.72 (0.72) | 0.51 (0.49) | 0.52 (0.49) |
| Model Dataset | Are You Racist, Or Am I Seeing Things? | BiasCorp | MMHS150k | Twitter Racism Dataset | Average |
|---|---|---|---|---|---|
| deberta | 0.62 (0.60) | 0.52 (0.52) | 0.50 (0.48) | 0.61 (0.41) | 0.56 (0.50) |
| xlm_r_rac-ismo | 0.63 (0.58) | 0.51 (0.32) | 0.51 (0.38) | 0.63 (0.46) | 0.57 (0.43) |
| ChatGPT-5 | 0.84 (0.82) | 0.53 (0.53) | 0.63 (0.63) | 0.84 (0.83) | 0.71 (0.70) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Himelboim, I.; Baid, M. Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research. Big Data Cogn. Comput. 2026, 10, 143. https://doi.org/10.3390/bdcc10050143
Himelboim I, Baid M. Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research. Big Data and Cognitive Computing. 2026; 10(5):143. https://doi.org/10.3390/bdcc10050143
Chicago/Turabian StyleHimelboim, Itai, and Mudit Baid. 2026. "Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research" Big Data and Cognitive Computing 10, no. 5: 143. https://doi.org/10.3390/bdcc10050143
APA StyleHimelboim, I., & Baid, M. (2026). Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research. Big Data and Cognitive Computing, 10(5), 143. https://doi.org/10.3390/bdcc10050143

