Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data
Abstract
1. Introduction
- We construct and curate a corpus of Ecuadorian Spanish social-media conversations for inappropriate-content detection, including conversational context via a fixed window ().
- We propose a controlled synthetic data augmentation protocol (seed selection, few-shot generation, exact deduplication, and cosine-similarity filtering) to mitigate extreme class imbalance while preventing evaluation leakage.
- We evaluate monolingual (BETO) and multilingual (mBERT) Transformer baselines and demonstrate the impact of imbalance-handling strategies (original, undersampling, oversampling, and cost-sensitive training) on per-class recovery.
- We report a cost-sensitive BETO configuration that improves minority-class detection under a strict train/validation/test split and provides detailed error analysis (confusion matrices, ROC, and Precision–Recall curves).
2. Background (State of the Art)
2.1. Digital Violence Dynamics and Linguistic Challenges
2.2. Data Acquisition: The Post-API Paradigm
2.3. Synthetic Data in NLP: Generation and Filtering
- Seed Selection: A Human-in-the-Loop approach is used, strictly relying on real examples manually validated (Gold Standard). This is justified in the literature because generative models tend to amplify input biases; consequently, using high-quality seeds is the only mechanism to anchor generation to the linguistic reality of the target dialect and avoid semantic degradation [22].
- Controlled Generation (In-Context Learning): The Few-Shot Prompting is employed to leverage the ability of LLMs to recognize patterns at inference time. Specifically, previously validated instances are used as seed examples within structured prompts that instruct the model to preserve the communicative intent, semantic category, and linguistic characteristics of Ecuadorian Spanish while generating new variants. This approach enables the induction of local slang, morphology, and discourse patterns from a small number of demonstrative examples (k-shots) without modifying the model’s weights [23], thereby mitigating the semantic degradation commonly observed in zero-shot scenarios. The generated samples are subsequently subjected to validation, deduplication, and semantic similarity filtering before being incorporated into the final corpus.
- Filtering and Refinement: Synthetic generation often produces high redundancy rates. Based on evidence that language models memorize duplicated data, degrading their generalization ability [24], we implement exact deduplication and cosine-similarity filters. This ensures that synthetic data contributes real semantic variability to training, rather than merely adding volume that induces overfitting.
2.4. State-of-the-Art Architectures and Domain Adaptation
- mBERT (Multilingual): Multilingual BERT in its cased variant is built on the base-sized Transformer architecture and pretrained in a self-supervised manner on raw Wikipedia text in 104 languages [14,26]. To mitigate resource imbalance across languages, pretraining applies a sampling scheme that reduces the dominance of high-resource languages and increases exposure to low-resource languages. The model uses WordPiece tokenization with a shared vocabulary (110,000 subwords), is case-sensitive, and represents segment pairs as [CLS] Sentence A [SEP] Sentence B [SEP] (maximum length 512). Learning follows the standard objectives of Masked Language Modeling (MLM) and Next Sentence Prediction (NSP): in MLM, 15% of tokens are masked (80% [MASK], 10% random replacement, and 10% unchanged), and in NSP the model predicts whether B is the continuation of A (50% real pairs and 50% random pairs) [14].
- BETO (Monolingual): Considered the gold standard for Spanish NLP, it corresponds to the “dccuchile/bert-base-spanish-wwm-cased” variant, pretrained exclusively on Spanish using Whole Word Masking. Comparative studies show that it captures the language’s morphology, idioms, and syntactic regularities with higher fidelity, consistently outperforming mBERT in Spanish discourse-classification tasks on social media comments [13,27,28].
Fine-Tuning and Cost-Sensitive Learning
2.5. Related Work and Rationale for the Taxonomy
- Level 0 (Non-Offensive): Texts that do not contain inappropriate or hostile content [25].
- Level 1 (Dark Humor): Ironic expressions, sarcasm, or aggressive jokes that normalize hostility without constituting direct threats [25]. This category poses a significant classification challenge because its pragmatic meaning often depends on the immediate conversational context, making isolated comments difficult to distinguish from neutral or benign interactions. To reduce this ambiguity, the proposed framework incorporates the context window () described in Section 4.1, enabling the model to recover the preceding conversational cues required to correctly infer the speaker’s intent.
- Level 2 (Intolerance): Discriminatory, dehumanizing, or stigmatizing language directed at groups or identities [25].
- Level 3 (Violence): Direct threats, incitement, or explicit descriptions of physical/sexual harm [25].
2.6. Metrics and Evaluation Protocols
- Global Performance Metrics:
- Accuracy: Provides a measure of the model’s overall correctness across all categories [32].
- Weighted Precision, Recall, and F1-Score: Due to the inherent imbalance in social media data, weighted means are used to adjust each class’s contribution according to its representativeness (support), preventing performance in the majority class (Non-Offensive) from masking deficiencies in critical classes [33,34].
- Per-Class Classification Report: To assess the model’s granularity within the four-class taxonomy, we individually analyze Precision, Recall, and F1-Score for each label [35]. This analysis is vital to identify ambiguity phenomena, such as those previously observed in the Dark Humor category, where distinguishing aggressive irony/sarcasm from neutral comments is often difficult [8,36]. Support is included to validate the statistical significance of the results in each category.
2.7. Research Questions
- RQ1: Which techniques enable appropriate selection and curation of data in Latin American Spanish?
- RQ2: Which methodological approach can be used for model selection in Spanish?
3. Methodology
3.1. Studied Problem (Task Definition)
3.2. Methodological Framework
- Problem and Data Understanding: Definition of the contextual digital-violence classification task and exploratory analysis of class distribution (severe imbalance) in the Ecuadorian digital ecosystem.
- Data Preparation and Engineering: Design of a hybrid (Human–AI) pipeline that goes beyond traditional cleaning. It includes post-API corpus acquisition, manual label validation after the failure of generic classifiers, and the implementation of a synthetic data generation and filtering protocol to correct class imbalance.
- Modeling: Selection of a monolingual Transformer architecture optimized for Spanish to execute supervised fine-tuning. This phase involves a critical performance comparison between training with the original imbalanced data versus our curated hybrid corpus, specifically testing the hypothesis that semantic deduplication and filtering are essential for effective domain adaptation in regional violence detection.
- Evaluation: Based on theoretical recommendations on balanced partitions [39], a two-stage stratified scheme was implemented, resulting in a final distribution of 70% for training, 15% for validation, and 15% for testing, performed before any synthetic augmentation (Section 4). Specifically: (i) a first stratified split with test_size = 0.30 reserves 70% for training; and (ii) the remaining 30% is stratified-split with test_size = 0.50, yielding two equal halves (15% validation, 15% test). Validation is used for model selection (early stopping), and the test set is reserved for the final evaluation, prioritizing per-class metrics (precision, recall, and F1) and error analysis via confusion matrices.
3.3. Detailed Workflow (Step-by-Step)
- Data acquisition and thread reconstruction: collect posts and their conversational context under post-API restrictions, and build context windows with preceding interactions (Section 4).
- Initial labeling and validation: perform a zero-shot labeling attempt and then apply manual validation to establish a reliable ground truth and quantify model failure under dialectal variation.
- Train/validation/test split (before augmentation): split the validated corpus into 70%/15%/15% using stratification, ensuring that any synthetic data derived from seeds remains within the training partition (leakage prevention).
- Synthetic data generation (minority classes only): generate class-conditioned candidates from validated seeds using few-shot prompting (Section 4).
- Quality control and filtering: remove exact duplicates and eliminate semantically redundant candidates via cosine-similarity filtering, keeping only samples that contribute semantic variability.
- Dataset assembly: construct the final training corpus by leveling the minority classes to a target operational threshold while keeping the validation and test partitions purely original.
- Model selection and training scenarios: compare BETO and mBERT under a controlled baseline and train BETO under multiple imbalance-handling strategies (original, undersampling, oversampling, and cost-sensitive learning) (Section 5).
- Evaluation and analysis: report global and per-class metrics, confusion matrices, and threshold-robust curves (ROC and Precision–Recall), emphasizing minority-class recovery in imbalanced settings (Section 6).
4. Development: Data Acquisition and Processing
4.1. Acquisition and Contextualization Pipeline
- Source: High-visibility and highly polarized profiles in Ecuador (politics, entertainment, crime news).
- Volume: 267,315 raw records were obtained, consolidated into 238,007 unique records after normalization.
4.2. Labeling Analysis: Zero-Shot Evaluation and Human Validation
- Automatic Prediction: 18,531 records classified as violence.
- Validated Reality: Only 347 records were genuine violence.
- Error Rate: 97.5% of alerts were cultural-interpretation errors (use of coarse slang in non-violent contexts).
4.3. Synthetic Data Engineering to Correct Class Imbalance
4.3.1. Split Protocol (Preventing Data Leakage)
4.3.2. Generation and Filtering Strategy
- Seeding: Use of the 1404 real violence records and 2998 real dark-humor records as a base.
- Mass Generation: Production of class-conditioned variations.
- Cleaning (Deduplication + Cosine Similarity): To mitigate the redundancy inherent to LLM-based generation, we applied a two-stage cleaning procedure combining exact deduplication and semantic similarity filtering. After removing identical records, sentence embeddings were generated using the hiiamsid/sentence_similarity_ spanish_es model, and a cosine-similarity threshold of was applied to identify semantically redundant samples. As shown in Figure 5, approximately 45% of the generated records (around 17,000 instances) were discarded because, despite exhibiting superficial lexical differences, they conveyed essentially the same semantic content as previously existing samples. Retaining these instances would have artificially increased the size of the corpus without contributing new information, reducing linguistic diversity and increasing the risk of overfitting. Consequently, the resulting hybrid corpus preserves genuine semantic variability, providing empirical support for the semantic richness introduced through the synthetic data generation process.
- Violence: From 16,825 generated → 9154 unique.
- Dark Humor: From 20,821 generated → 11,540 unique.
4.3.3. Final Dataset Composition (Gap Filling)
- Augmented critical classes: Violence and Dark Humor are raised to 25 k by injecting 29,306 curated synthetic examples.
- Purely real classes: Intolerance (36k) and Non-Offensive (196 k) remain entirely real to stabilize training and reduce false positives.
5. Development: Model Selection and Modeling
5.1. Experimental Setup (Environment and Tools)
- Hardware: GPU: T4, RAM: [High-RAM].
- Software: Python [3.12], PyTorch [2.5.1], Transformers [4.46.1], scikit-learn [1.5.2].
5.2. Model Selection
5.3. Fine-Tuning with Mixed Data
5.3.1. Preprocessing
5.3.2. Training Scenarios
5.3.3. Hyperparameter Selection and Final Configuration
5.3.4. Evaluation Protocol
6. Results
7. Discussion
7.1. Data Selection and Curation in Latin American Spanish (RQ1)
7.2. Model Selection in Spanish (RQ2)
7.3. Critical Analysis and Limitations
7.4. Practical Implications
- (a)
- keep real Validation and Test sets strictly separated,
- (b)
- report stability across multiple seeds and confidence intervals, and
- (c)
- validate the model in out-of-domain scenarios and/or other dialectal variants.
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| API | Application Programming Interface |
| BERT | Bidirectional Encoder Representations from Transformers |
| BETO | Spanish BERT model |
| CRISP-DM | Cross-Industry Standard Process for Data Mining |
| F1 | F1-score |
| IA | Artificial Intelligence |
| LDA | Latent Dirichlet Allocation |
| LLM | Large Language Model |
| LLaMA | Large Language Model Meta AI |
| mBERT | Multilingual BERT |
| MLM | Masked Language Modeling |
| NLP | Natural Language Processing |
| NSP | Next Sentence Prediction |
| PLM | Pretrained Language Model |
| ROC | Receiver Operating Characteristic |
References
- Dhanaraj, A. The Evolution of Cyber Threats: From Traditional Attacks to AI-Powered Challenges. Eur. J. Comput. Sci. Inf. Technol. 2025, 13, 50–61. [Google Scholar] [CrossRef] [Scilit]
- Leymann, H. The Content and Development of Mobbing at Work. Eur. J. Work. Organ. Psychol. 1996, 5, 165–184. [Google Scholar] [CrossRef] [Scilit]
- Laboy-Vélez, L.; Ríos-Steiner, A.I.; Flores-Suárez, W. La violencia digital como amenaza a un ambiente laboral seguro. Forum Empres. 2021, 26, 99–112. [Google Scholar] [CrossRef] [Scilit]
- Plaza-del Arco, F.; Molina-González, M.; Ureña López, L.; Martín-Valdivia, M. Integrating Implicit and Explicit Linguistic Phenomena via Multi-task Learning for Offensive Language Detection. Knowl.-Based Syst. 2022, 258, 109965. [Google Scholar] [CrossRef] [Scilit]
- Mabula, M.; Mambeti, S.; Mbelwa, J. Cyberbullying Detection: Exploring Datasets, Technologies, and Approaches on Social Media Platforms. arXiv 2024, arXiv:2407.12154. [Google Scholar]
- Báez, P.; Villena, F.; Durán, M. Detecting Cyberbullying in Spanish Texts throughout Deep Learning Techniques. Int. J. Data Min. Model. Manag. 2022, 14, 234–247. [Google Scholar] [CrossRef] [Scilit]
- Plaza-del Arco, F.; Casavantes, M.; Escalante, H.; Martín-Valdivia, M.; Montejo-Ráez, A.; Montes, M.; Jarquín-Vásquez, H.; Villaseñor-Pineda, L. Overview of MeOffendEs at IberLEF 2021: Offensive Language Detection in Spanish Variants. Proces. Leng. Nat. 2021, 67, 183–194. [Google Scholar]
- Plaza-del Arco, F.; Molina-González, M.; Ureña López, L.; Martín-Valdivia, M. Comparing Pre-trained Language Models for Spanish Hate Speech Detection. Expert Syst. Appl. 2021, 166, 114120. [Google Scholar] [CrossRef] [Scilit]
- Akosa, J. Predictive Accuracy: A Misleading Performance Measure for Highly Imbalanced Data. In Proceedings of the SAS Global Forum 2017, Cary, NC, USA, 2–5 April 2017. Paper 942-2017. [Google Scholar]
- Dai, H.; Liu, Z.; Liao, W.; Huang, X.; Cao, Y.; Wu, Z.; Zhao, L.; Xu, S.; Zeng, F.; Liu, W.; et al. AugGPT: Leveraging ChatGPT for Text Data Augmentation. IEEE Trans. Big Data 2025, 11, 907–918. [Google Scholar] [CrossRef] [Scilit]
- Jahan, M.; Oussalah, M.; Beddia, D.; Mim, J.; Arhab, N. A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Zambrano, P.; Torres, J.; Anchundia, C.; Illicachi, J. Mobbing: AI-Powered Cyberthreat Behavior Analysis and Modeling. In Proceedings of the 2024 8th Cyber Security in Networking Conference (CSNet); IEEE: Piscataway, NJ, USA, 2024; pp. 278–281. [Google Scholar] [CrossRef] [Scilit]
- Cañete, J.; Chaperon, G.; Fuentes, R.; Ho, J.; Kang, H.; Pérez, J. Spanish Pre-Trained BERT Model and Evaluation Data. In Proceedings of the PML4DC at International Conference on Learning Representations 2020, Virtual Conference, 26 April 2020. [Google Scholar]
- Devlin, J.; Chang, M.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Bruns, A. After the ‘APIcalypse’: Social media platforms and their fight against critical scholarly research. Inf. Commun. Soc. 2019, 22, 1544–1566. [Google Scholar] [CrossRef] [Scilit]
- Plaza, L.; Carrillo-de Albornoz, J.; Arcos, I.; Rosso, P.; Spina, D.; Amigó, E.; Gonzalo, J.; Morante, R. Overview of EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2026; Volume 16089, pp. 266–289. [Google Scholar]
- Maqbool, N. Sexism Identification in Social Networks: Advances in Automated Detection—A Report on the Exist Task at CLEF; CEUR Workshop Proceedings; CEUR-WS.org: Aachen, Germany, 2024; Volume 3740, pp. 1098–1106. [Google Scholar]
- Urios Alacreu, E.; Rosso, P. Identification of Racial and Sexist Stereotypes in Spanish: A Learning with Disagreements Approach. Proces. Del Leng. Nat. 2025, 74, 15–31. [Google Scholar]
- Kharitonova, K.; Pérez-Fernández, D.; Gutiérrez-Hernando, J.; Gutiérrez-Fandiño, A.; Callejas, Z.; Griol, D. EsCorpiusBias: The Contextual Annotation and Transformer-Based Detection of Racism and Sexism in Spanish Dialogue. Future Internet 2025, 17, 340. [Google Scholar] [CrossRef] [Scilit]
- Castillo-López, G.; Riabi, A.; Seddah, D. Analyzing Zero-Shot transfer Scenarios across Spanish variants for Hate Speech Detection. In Proceedings of the 10th Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023), Dubrovnik, Croatia, 5 May 2023; pp. 1–13. [Google Scholar]
- Mitchell, R. Web Scraping with Python: Collecting More Data from the Modern Web, 2nd ed.; O’Reilly Media: Sebastopol, CA, USA, 2018. [Google Scholar]
- Yoo, K.; Park, D.; Kang, J.; Lee, S.W.; Park, W. GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event/Punta Cana, Dominican Republic, 7–11 November 2021; pp. 2225–2239. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. In Proceedings of the Advances in Neural Information Processing Systems; Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 1877–1901. [Google Scholar]
- Lee, K.; Ippolito, D.; Nystrom, A.; Zhang, C.; Eck, D.; Callison-Burch, C.; Carlini, N. Deduplicating Training Data Makes Language Models Better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; pp. 8424–8445. [Google Scholar] [CrossRef] [Scilit]
- Zambrano, P.; Torres, J.; Anchundia, C.; Illicachi, J. Gender-Based Violence as a Cyberthreat: The Impact of Social Media on Digital Security. In Computer Science and Computational Intelligence (CSCI 2024); Communications in Computer and Information Science (CCIS); Springer: Cham, Switzerland, 2025; Volume 2509, pp. 237–256. [Google Scholar] [CrossRef] [Scilit]
- Pires, T.; Schlinger, E.; Garrette, D. How Multilingual is Multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 4996–5001. [Google Scholar] [CrossRef] [Scilit]
- Nozza, D.; Bianchi, F.; Hovy, D. What the [MASK]? Making Sense of Language-Specific BERT Models. arXiv 2020. [Google Scholar] [CrossRef] [Scilit]
- Ponnusamy, K.; Vegupatti, M.; Kumaresan, P.; Priyadharshini, R.; Buitelaar, P.; Chakravarthi, B. VEL@IberLEF 2024: Hope Speech Detection in Spanish Social Media Comments using BERT Pre-trained Model. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2024); CEURWorkshop Proceedings; CEUR-WS.org: Aachen, Germany, 2024; Volume 3756. [Google Scholar]
- García-Díaz, J.A.; Jiménez-Zafra, S.M.; Valencia-García, R. UMUTeam at HOMO-MEX 2023: Fine-tuning Large Language Models integration for solving hate-speech detection in Mexican Spanish. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2023); CEURWorkshop Proceedings; CEUR-WS.org: Aachen, Germany, 2023; Volume 3496. [Google Scholar]
- Sanchez-Gomez, J.M.; Batista, F.; Vega-Rodríguez, M.A.; Pérez, C.J. A transformer-based deep learning approach for detecting online hate speech in Spanish. Appl. Soft Comput. 2026, 187, 114259. [Google Scholar] [CrossRef] [Scilit]
- Hosseini, H.; Kannan, S.; Zhang, B.; Poovendran, R. Deceiving Google’s Perspective API Built for Detecting Toxic Comments. In Proceedings of the 2017 Points of View in Computer Vision Workshop, Venice, Italy, 22–29 October 2017. [Google Scholar]
- Grandini, M.; Bagli, E.; Visani, G. Metrics for multi-class classification: An overview. arXiv 2020, arXiv:2008.05756. [Google Scholar]
- Davidson, T.; Warmsley, D.; Macy, M.; Weber, I. Automated hate speech detection and the problem of offensive language. In Proceedings of the International AAAI Conference on Web and Social Media, Montreal, QC, Canada, 15–18 May 2017; Volume 11. [Google Scholar]
- Haixiang, G.; Yijing, L.; Jennifer, S.; Mingyun, G.; Yuanyuan, H.; Bing, G. Learning from class-imbalanced data: Review of methods and applications. Expert Syst. Appl. 2017, 73, 220–239. [Google Scholar] [CrossRef] [Scilit]
- Zampieri, M.; Malmasi, S.; Nakov, P.; Rosenthal, S.; Farra, N.; Kumar, R. Predicting the Type and Target of Offensive Posts in Social Media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, Minneapolis, MN, USA, 2–7 June 2019. [Google Scholar]
- Barbieri, F.; Camacho-Collados, J.; Neves, L.; Espinosa-Anke, L. TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16–20 November 2020. [Google Scholar]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef] [Scilit]
- Gholamy, A.; Kreinovich, V.; Olaya, O. Why 70/30 or 80/20 Relation Between Training and Testing Sets: A Pedagogical Explanation; Technical Report UTEP-CS-18-09; University of Texas at El Paso: El Paso, TX, USA, 2018. [Google Scholar]
- Gilardi, F.; Alizadeh, M.; Kubli, M. ChatGPT outperforms crowd workers for text-annotation tasks. Proc. Natl. Acad. Sci. USA 2023, 120, e2305016120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Heaton, J. Genetic Programming and Evolvable Machines; Springer: Berlin/Heidelberg, Germany, 2017; Volume 19. [Google Scholar] [CrossRef] [Scilit]
- Kaufman, S.; Rosset, S.; Perlich, C. Leakage in Data Mining: Formulation, Detection, and Avoidance. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and 845 Data Mining, San Diego, CA, USA, 21–24 August 2011; Volume 6, pp. 556–563. [Google Scholar] [CrossRef] [Scilit]












| Metric | Value |
|---|---|
| Accuracy | 94.39% |
| Weighted F1 | 0.9429 |
| Macro F1 | 0.9022 |
| F1 Non-Offensive | 0.9669 |
| F1 Violence | 0.9028 |
| F1 Intolerance | 0.8801 |
| F1 Dark Humor | 0.8590 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Rodríguez, P.X.Z.; Aguayo, M.P.S.; Valencia, C.E.A.; Manzano, J.S.I.; Calahorrano, A.D.O.; Montenegro, A.E.P.; Espinosa, J.S.L. Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data. Informatics 2026, 13, 151. https://doi.org/10.3390/informatics13090151
Rodríguez PXZ, Aguayo MPS, Valencia CEA, Manzano JSI, Calahorrano ADO, Montenegro AEP, Espinosa JSL. Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data. Informatics. 2026; 13(9):151. https://doi.org/10.3390/informatics13090151
Chicago/Turabian StyleRodríguez, Patricio Xavier Zambrano, Marco Polo Sánchez Aguayo, Carlos Eduardo Anchundia Valencia, Johan Sebastian Illicachi Manzano, Andrea Damarys Oña Calahorrano, Adrian Esteban Paguay Montenegro, and Juan Sebastián León Espinosa. 2026. "Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data" Informatics 13, no. 9: 151. https://doi.org/10.3390/informatics13090151
APA StyleRodríguez, P. X. Z., Aguayo, M. P. S., Valencia, C. E. A., Manzano, J. S. I., Calahorrano, A. D. O., Montenegro, A. E. P., & Espinosa, J. S. L. (2026). Inappropriate Content Classification Model for Digital Violence Detection Using Hybrid Data. Informatics, 13(9), 151. https://doi.org/10.3390/informatics13090151

