Abstract
This study investigates how sentiment polarity relates to offensive function in Chinese online attacks, focusing on cases where offence is expressed implicitly rather than through explicit negative wording. Using a 1400-item Chinese online-text corpus derived from TOXICN and its pragmatics-oriented annotation extension, the study examines sentiment distributions across non-offensive texts, explicit attacks, and implicit attacks, and evaluates how these patterns are reflected in model-generated sentiment labels. The implicit-attack subset is analyzed through four pragmatic strategies: Irony, Trope, Indirectness, and Exaggeration. A stratified 400-item validation subset was independently annotated by three human annotators for sentiment polarity, and four NLP models were used to assign sentiment labels to the full corpus: Claude Haiku 4.5, OpenAI GPT-4.1-mini, a Chinese BERT-based sentiment classifier, and Multilingual DistilBERT. Human reference annotations showed that negative sentiment was strongly associated with offensive-language labels but did not map onto them perfectly: some non-offensive texts were perceived as negative, while some implicit attacks were assigned neutral sentiment. Pragmatic strategy did not significantly predict human reference sentiment polarity in the validation subset. Model–human agreement varied across the four selected systems, with Claude Haiku 4.5 and OpenAI GPT-4.1-mini showing higher agreement than Chinese BERT and Multilingual DistilBERT in the validation subset. Full-corpus mixed-effects logistic regression showed that model-generated negative labels varied significantly by text category and model, whereas evidence for stable model-specific pragmatic-strategy effects was limited. These findings provide empirical evidence from Chinese online attacks that sentiment polarity can inform, but cannot replace, offensive-language analysis. They also show that model-generated sentiment labels require human validation and cautious interpretation when offence is implicit, figurative, homophonic, or context-dependent.
1. Introduction
With the rapid development of social media and online platforms, offensive language has become an important research topic in natural language processing (NLP) and content moderation. Hate speech is commonly understood as language that expresses hostility, derogation, or harmful intent toward individuals or groups on the basis of identity attributes, such as race, gender, region, religion, sexual orientation, or other social identity categories [1]. Existing studies often formulate offensive language detection as a text classification task, namely determining whether a text contains insults, discrimination, derogation, or hateful content. Machine learning and deep learning methods have been widely applied to offensive language and hate speech detection, with substantial progress especially in identifying explicit abuse, insults, and discriminatory expressions [2,3,4].
However, offensive language does not always appear in explicit forms. Implicit attacks may convey offensive meaning without direct insults or overtly abusive vocabulary, relying instead on insinuation, irony, analogy, presupposition, coded expressions, or shared contextual knowledge [5,6,7,8,9]. In Chinese online discourse, this difficulty is further intensified by homophones, character variation, internet memes, region-specific expressions, and other forms of indirect or altered wording [10,11].
Sentiment analysis and offensive language detection are related but distinct tasks. Sentiment analysis focuses on the affective or evaluative orientation expressed in a text, typically using categories such as negative, neutral, and positive [12,13]. Offensive language detection, by contrast, concerns whether a text performs an attacking, derogatory, abusive, or exclusionary function. Although the two dimensions often overlap in explicit attacks, this study treats their distinction as the conceptual premise for examining how sentiment polarity and offensive function are empirically related in Chinese online attacks.
This issue is especially important for implicit attacks, where an utterance may appear as a factual statement, rhetorical question, joke, analogy, exaggeration, or apparently positive expression while still conveying offensive function toward a target. Pragmatic theories of implicature, indirect speech acts, relevance, and pragmatic enrichment provide useful resources for understanding how hearers infer meanings beyond literal wording [14,15,16,17], while corpus-based work on implicit and covert offensive language shows that such meanings are often realized through socially and culturally situated linguistic mechanisms [6,7,8,9].
Despite this connection, relatively little is known about how sentiment polarity behaves across different pragmatic forms of Chinese implicit attacks, or how different NLP systems assign sentiment labels to such texts. Prior work on Chinese offensive language resources, including COLD and TOXICN, has advanced dataset construction and model evaluation for Chinese toxic and offensive language [18,19]. However, the relationship among human-perceived sentiment polarity, original offensive-language labels, pragmatic strategies, and model-generated sentiment labels remains underexplored.
The present study addresses this issue using a 1400-item Chinese online-text corpus derived from TOXICN [18] and its pragmatics-oriented annotation extension. The corpus includes 600 non-offensive texts, 400 explicit attacks, and 400 implicit attacks. The implicit-attack subset is further analyzed according to four pragmatic strategies: Irony, Trope, Indirectness, and Exaggeration. To provide a human reference point for sentiment polarity, a stratified 400-item validation subset was independently annotated by three human annotators for negative, neutral, or positive sentiment. In addition, four NLP systems were used to assign sentiment labels to the full corpus: Claude Haiku 4.5, OpenAI GPT-4.1-mini, a Chinese BERT-based sentiment classifier, and Multilingual DistilBERT.
The study is guided by four research questions:
- RQ1: How are human reference sentiment labels distributed across non-offensive texts, explicit attacks, and implicit attacks?
- RQ2: Do different pragmatic strategies in implicit attacks correspond to different human reference sentiment distributions?
- RQ3: How do sentiment labels generated by the four selected NLP models compare with human reference sentiment labels in the validation subset?
- RQ4: How do model-generated negative sentiment labels vary across text categories, pragmatic strategies, and models in the full corpus?
Accordingly, this study makes three main contributions. First, it empirically examines how the known distinction between sentiment polarity and offensive function appears in Chinese online offensive-language data, especially in implicit attacks. Second, it connects this distinction to a pragmatics-based taxonomy of Chinese implicit attacks, examining whether Irony, Trope, Indirectness, and Exaggeration correspond to different sentiment-polarity patterns. Third, it compares four selected NLP systems against human reference sentiment labels and evaluates full-corpus model-generated negative-label patterns using repeated-measures mixed-effects analysis. Through this design, the study clarifies when sentiment polarity can serve as an informative cue to offensive function and when it should be treated as a distinct dimension requiring pragmatic and contextual interpretation.
2. Theoretical Background and Related Work
2.1. Online Offensive Language and Implicit Attacks
Online offensive language detection is an important research area in natural language processing. Existing studies have typically focused on the automatic detection of hate speech, offensive language, or toxic language, noting that such texts often involve hostility, derogation, exclusion, or harmful expressions targeting the identity attributes of individuals or groups [2,3]. In the Chinese context, datasets such as COLD and TOXICN have advanced the detection of Chinese offensive and toxic language, providing important resources for model training and evaluation [18,19].
However, offensive language does not always appear in explicit forms. ElSherief et al. show that implicit hate speech often conveys offensive meaning through insinuation, coded expressions, analogy, or contextual inference rather than direct abuse or explicit offensive terms [5]. Recent research on subtle or implicit hate speech has similarly shown that implicit attacks may rely on indirect expression, metaphor, irony, presupposition, and shared cultural knowledge [7,8]. In Chinese online discourse, homophones, variant characters, internet memes, emojis, and region-specific expressions further increase this complexity. The analysis of implicit offensive language therefore needs to move beyond surface lexical cues and coarse-grained labels and attend to pragmatic mechanisms.
2.2. Pragmatic Strategies in Implicit Attacks
Pragmatic theory provides an important perspective for understanding implicit attacks. Grice’s theory of conversational implicature shows how speakers can convey intended meanings through inference beyond literal wording [14]. Searle’s theory of indirect speech acts explains how surface forms such as questions, suggestions, or statements can perform communicative functions such as blaming, mocking, or attacking [15]. Relevance theory emphasizes that discourse interpretation depends on contextual assumptions and pragmatic enrichment [16]. These frameworks help explain why the offensive meaning of an implicit attack may emerge from the relation between literal form, inferred speaker meaning, and available context.
The present study uses a four-way pragmatic taxonomy of Chinese implicit attacks: Irony, Trope, Indirectness, and Exaggeration. This taxonomy is consistent with prior work showing that implicit and covert offensive language may be realized through sarcasm, metaphor, allusion, presupposition, overstatement, and understatement [5,7,8,9]. In this section, the taxonomy is introduced as a theoretical and descriptive basis for analyzing implicit offence; the operational definitions used for annotation are provided in Section 3.2. Together, these strategies highlight that implicit attacks do not necessarily rely on explicit negative terms, but may instead be constructed through contextual inference, rhetorical mapping, and non-literal meaning.
2.3. Distinguishing Sentiment Polarity from Offensive Function
Sentiment analysis and offensive language detection are related but distinct tasks. Sentiment analysis usually focuses on the emotion or evaluative orientation expressed in a text and classifies it into categories such as negative, neutral, and positive [12,13]. By contrast, offensive language detection focuses on whether a text performs a derogatory, exclusionary, stigmatizing, or harmful function. The two tasks may overlap in explicit attack texts; for example, direct abuse is often likely to be classified as negative. However, such overlap does not mean that sentiment polarity can substitute for offensiveness.
On the one hand, negative texts are not necessarily offensive, as complaints, sadness, or criticism of events may also express negative sentiment. On the other hand, offensive texts do not necessarily display negative polarity. Implicit attacks may perform an offensive function through apparent praise, rhetorical questions, homophones, tropes, or neutral statements, and may therefore be classified by sentiment models as neutral or positive. Thus, sentiment polarity reflects only the affective orientation of a text and cannot be directly equated with offensive function.
2.4. NLP Models and Sentiment Classification of Implicit Meaning
Different NLP models may assign sentiment labels to implicitly expressed meaning differently because they vary in architecture, training objectives, training data, label systems, language coverage, and inference procedures. Discriminative models such as BERT-based classifiers typically use pretrained encoders followed by supervised fine-tuning for classification tasks [20]. Their outputs therefore depend on the training distributions and label definitions of the corresponding fine-tuning data. Generative large language models, by contrast, can perform classification through task instructions and contextual input, but their outputs may also vary with prompt formulation, model version, and inference settings. The present study does not assume that either model type is inherently better at recovering implicit meaning. Instead, it compares the sentiment labels generated by four selected systems under the present experimental settings.
2.5. Research Gap and Positioning of This Study
In summary, existing research has advanced offensive-language detection, Chinese toxic-language resources, implicit-offence analysis, and sentiment analysis models. However, the relationship between offensive-language labels and sentiment polarity remains insufficiently examined, especially when offence is conveyed implicitly through pragmatic strategies. Existing work has also rarely compared model-generated sentiment labels with independent human sentiment annotations in this setting.
Building on the research questions proposed in the Introduction, this study positions sentiment polarity and offensive function as related but non-equivalent dimensions of online language. Its novelty lies in combining human reference sentiment annotation, pragmatic classification of Chinese implicit attacks, and full-corpus model-output analysis. The study first examines whether human-perceived negative sentiment aligns with offensive-language labels, then analyses whether pragmatic strategies correspond to different human sentiment distributions, and finally compares model-generated sentiment labels with human reference labels and full-corpus offensive-language categories. Through this design, the study avoids treating model outputs as gold-standard sentiment labels and instead evaluates how human and model sentiment judgments relate to offensive function in Chinese online attacks.
3. Methods
This study examines how sentiment polarity relates to offensive function in Chinese online texts, especially when offence is conveyed implicitly through pragmatic strategies. Offensive-language labels and sentiment-polarity labels are treated as distinct analytical dimensions. The analysis first structures the source offensive-language and pragmatic-strategy labels, then adds human reference sentiment labels for a validation subset, and finally compares these labels with model-generated sentiment outputs from four NLP systems.
3.1. Dataset
This study uses a 1400-item Chinese online-text corpus derived from TOXICN [18] and its pragmatics-oriented annotation extension. TOXICN is a publicly available Chinese toxic-language dataset collected from Zhihu and Baidu Tieba. The dataset and code are available at https://github.com/DUT-lujunyu/ToxiCN (accessed on 20 July 2026), and the original paper is available through the ACL Anthology at https://aclanthology.org/2023.acl-long.898/ (accessed on 6 September 2026). According to the original documentation, TOXICN focuses on four sensitive topics: gender, race, region, and LGBTQ. Keyword-based filtering was used to collect 15,442 comments, of which 12,011 were retained after removing duplicated, irrelevant, or semantically insufficient samples. The dataset is released under the CC BY-NC-ND 4.0 licence for scientific research and non-commercial use. The original paper identifies the collection platforms and filtering procedure but does not specify a precise collection period.
TOXICN uses the Monitor Toxic Frame, a hierarchical annotation scheme that distinguishes toxic and non-toxic content, general offensive language and hate speech, targeted group, and expression category. In the present study, toxic language refers to the broader label framework of TOXICN, while offensive language refers to text that performs an attacking, insulting, derogatory, or abusive function. Attack is used as the operational category and includes explicit attacks, where offensive function is conveyed through direct abusive or derogatory wording, and implicit attacks, where offensive function is conveyed indirectly through pragmatic strategies. Hate speech is treated as a narrower category within toxic or offensive language, involving attacks based on social attributes such as gender, race, region, or LGBTQ identity.
The original TOXICN labels were produced through hierarchical manual annotation. According to Lu et al. [18], nine annotators with linguistics-related training participated, and each comment was labelled by at least three annotators. Final labels were determined by majority vote. The original paper reports Fleiss’ kappa values of 0.62 for toxic identification, 0.75 for toxic type, 0.65 for targeted group, and 0.68 for expression category. In the present study, the non-offensive versus offensive distinction is most closely related to toxic identification, while the explicit versus implicit distinction is most closely related to expression category, which distinguishes explicitness, implicitness, and reporting. These source labels were adopted as corpus-level grouping variables and were not re-annotated.
The four pragmatic-strategy labels were not part of the original TOXICN annotation scheme but were introduced through the pragmatics-oriented annotation extension. These labels were obtained through the annotation and expert-adjudication procedure described in Section 3.2 and were used as the pragmatic-strategy annotations for the present analysis.
The working corpus contains 600 non-offensive texts, 400 explicit attacks, and 400 implicit attacks. The 1400-item working corpus was constructed by randomly selecting samples within these three predefined categories. In the working label format, output = 0 denotes non-offensive content, output = 1 denotes explicit attack, and output = 1|strategy denotes implicit attack with one or more pragmatic-strategy labels. The four standard pragmatic strategies are Irony, Trope, Indirectness, and Exaggeration. For analyses requiring one primary strategy, the primary strategy identified through the annotation and adjudication procedure was used. In compound labels, this primary strategy was listed first and represented the dominant mechanism for conveying offensive function; labels outside the four standard categories were grouped as UNKNOWN.
Because the unit of analysis is a single online text item, full conversational context was not available for all cases. This is particularly relevant for implicit attacks, whose interpretation may depend on shared background knowledge, prior discourse, target reference, or interactional cues. The pragmatic-strategy labels are therefore treated as source annotations made under available-text conditions rather than as context-complete determinations of speaker intent. The present study did not conduct new data scraping or analyze private messages or non-public user communications. The working dataset did not include usernames, user IDs, profile links, URLs, or other directly identifying information. Because verbatim online texts can sometimes be searchable even after direct identifiers are removed, qualitative examples are used sparingly and only to illustrate linguistic and pragmatic patterns relevant to the analysis.
3.2. Pragmatic Strategy Annotation and Reliability Assessment
This study adopts a pragmatics-based annotation framework for Chinese implicit offensive texts. The framework draws on several traditions in pragmatics, but these traditions are not treated as interchangeable. Grice’s account of conversational implicature helps explain how speakers may convey evaluative or offensive meanings beyond explicit wording [14]. Searle’s theory of indirect speech acts is relevant to cases where an utterance performs an attacking, blaming, or derogating function through a surface form that does not directly encode that function [15]. Relevance theory emphasizes the role of contextual assumptions, shared background knowledge, and inferential effort in deriving speaker meaning [16]. Bach’s notion of conversational impliciture is useful for cases where literal content requires pragmatic enrichment before its evaluative force can be recovered [17].
These theoretical perspectives motivate the pragmatic orientation of the study, but they are not used as separate annotation categories. The operational framework classifies implicit attacks according to four corpus-level pragmatic strategies: Irony, Trope, Indirectness, and Exaggeration. These categories are practical coding categories rather than theoretically exhaustive or mutually exclusive types. In naturally occurring online discourse, the same text may involve more than one pragmatic mechanism; for example, an ironic utterance may also be indirect, figurative, or exaggerated. The taxonomy therefore identifies the dominant pragmatic mechanism available in the text while allowing secondary strategies to be recorded.
The four categories were defined as follows. Irony refers to cases in which the intended evaluative stance contrasts with the surface or literal meaning of the text, such as apparent praise used to convey criticism. Trope refers to comparison-based figurative expression, including metaphor, simile, analogy, and related associative mappings [21], such as degrading comparisons between a target and an animal. Indirectness refers to cases in which offensive function is conveyed through indirect speech acts, rhetorical questions, insinuation, presupposition, or euphemistic formulation rather than through directly derogatory wording, such as a rhetorical question that implies a negative evaluation without asserting it explicitly. These phenomena are grouped operationally because they convey offensive function indirectly, but they are not treated as theoretically identical; presupposed content, for example, is not treated as a subtype of indirect speech act. Exaggeration refers to cases in which evaluation is intensified or downplayed through overstatement or understatement in a way that contributes to offensive function. Following Bączkowska et al. [6], these two distinct figures of speech were grouped under Exaggeration because both instantiate the same underlying condition of exaggeration: overstatement increases the evaluative value or importance assigned to a feature, whereas understatement reduces or downplays it. Representative annotated examples of all four categories, including both metaphor and simile under Trope and both overstatement and understatement under Exaggeration, are provided in Appendix B (Table A8).
For the pragmatics-oriented annotation extension, pragmatic strategy annotation was conducted by nine annotators with backgrounds in linguistics or computational linguistics. The annotators were divided into three independent groups and received standardized training on category definitions, representative examples, boundary cases, and multi-strategy annotation. The unit of annotation was a single implicit-attack text item. Each group produced one annotation set for each item after internal discussion, including one primary strategy and, where relevant, up to two secondary strategies. The primary strategy was defined as the dominant mechanism for conveying offensive function under the available-text condition.
Reliability was calculated at the level of the three independent group-level annotations. For primary pragmatic strategies, Fleiss’ kappa was κ = 0.5842, indicating appreciable but imperfect agreement across the three group-level annotations. For multi-label annotations, average pairwise Jaccard similarity was calculated between the strategy sets assigned by the three groups, yielding 0.6358. Remaining disagreements were resolved through expert adjudication by two experts with experience in online offensive-language research. The adjudicated labels were retained as the final pragmatic-strategy labels and used for all subsequent descriptive and inferential analyses. The reliability statistics therefore characterize agreement prior to expert adjudication.
During data processing, compound labels such as irony + exaggeration were split into primary and secondary strategies. The main statistical analyses used the primary pragmatic strategy, while secondary strategies were retained descriptively. Labels outside the four-category system, such as other and marked form, were grouped as UNKNOWN and reported as supplementary categories. Other refers to implicit attacks whose pragmatic mechanism did not fit one of the four predefined categories. Marked form refers to cases in which offensiveness is conveyed through altered or non-standard linguistic form, such as homophonic substitution, character variation, spelling deformation, or visually marked wording. The distributions of primary, secondary, and non-standard labels are reported in Appendix A Table A1 and Table A2.
3.3. Sentiment Polarity Annotation
Because the source corpus contains offensive-language labels but not independently established sentiment-polarity labels, this study used two complementary sources of sentiment annotation: human reference sentiment annotation for a stratified validation subset and model-generated sentiment annotation for the full corpus.
3.3.1. Human Reference Sentiment Annotation
A stratified subset of 400 texts was selected from the full corpus for human sentiment annotation. The subset included 100 non-offensive texts, 100 explicit attacks, and 200 implicit attacks. The implicit-attack subset contained 56 Irony cases, 42 Trope cases, 50 Indirectness cases, and 52 Exaggeration cases.
Three annotators independently assigned one sentiment label to each text: negative, neutral, or positive. Annotators did not have access to model-generated sentiment labels and were instructed to annotate expressed sentiment polarity rather than offensiveness or target-directed stance. Expressed sentiment polarity was defined as the overall affective or evaluative orientation conveyed by the utterance as a whole. Annotators were explicitly instructed not to treat offensiveness as equivalent to negative sentiment, because implicitly offensive texts may convey offensive function through apparently neutral or positive surface wording.
The annotation scheme distinguished expressed sentiment polarity from related but unannotated dimensions: literal lexical valence, speaker stance, target-directed evaluation, and communicative/offensive function. Negative referred to texts expressing negative affect, criticism, anger, dissatisfaction, contempt, hostility, rejection, complaint, or negative evaluation. Neutral referred to factual statements, informational expressions, questions, quotations, weakly affective texts, or cases where positive and negative cues were ambiguous or balanced. Positive referred to praise, approval, liking, support, gratitude, amusement, friendly tone, or positive evaluation.
Inter-annotator agreement was assessed using Fleiss’ kappa. The three annotators reached complete agreement on 303 of the 400 items, corresponding to a complete agreement rate of 75.75%. Fleiss’ kappa was κ = 0.554 (95% bootstrap CI [0.480, 0.623]), indicating non-trivial agreement alongside appreciable annotation uncertainty. Final human reference sentiment labels were determined by majority vote. For the single item in which all three annotators selected different labels, the final label was determined through adjudication according to the annotation guidelines. Because agreement was moderate rather than high, these labels are treated as human reference labels rather than definitive gold-standard sentiment labels.
3.3.2. Model-Generated Sentiment Annotation
For the full corpus of 1400 texts, sentiment polarity was also annotated using four NLP models: Claude Haiku 4.5 (claude-haiku-4-5-20251001), OpenAI GPT-4.1-mini, a Chinese BERT-based sentiment classifier (senlou/weibo-sentiment-chinese-bert), and Multilingual DistilBERT (lxyuan/distilbert-base-multilingual-cased-sentiments-student). The purpose was to examine how different systems assign sentiment labels across offensive-language and pragmatic categories and to compare model outputs with human reference labels in the validation subset. Model-generated labels were therefore analyzed as system outputs rather than as independently verified sentiment labels.
All four models were mapped onto the same three-label output scheme: negative, neutral, and positive. This shared vocabulary was used for harmonization, but it does not imply that the labels had identical operational meanings across systems. For the generative models, labels were elicited through zero-shot prompt instructions defining sentiment polarity as the overall affective or evaluative orientation of the utterance. For the discriminative classifiers, labels reflect the annotation schemes and training distributions of their fine-tuning corpora.
Claude Haiku 4.5 and OpenAI GPT-4.1-mini were run through API-based batch inference. The models were instructed to classify each Chinese text into one of the three sentiment categories and return only the text ID and sentiment label in structured JSON format. Each batch contained only the text ID and Chinese text. Outputs were automatically checked for missing IDs, duplicated IDs, and invalid labels. Both generative models were run with temperature set to 0. The original full-corpus inference used one prompt formulation and one complete inference run for each generative model; results are therefore interpreted under this specific prompt and batch setting. A limited repeated-run stability check on the 400-item validation subset is reported in Appendix A Table A4.
The two discriminative classifiers were run locally using Hugging Face. Each text was passed to the corresponding tokenizer with maximum sequence length set to 256, and the final sentiment label was determined by the class with the highest softmax probability. The Chinese BERT-based classifier was included as a Chinese-language, Weibo-oriented exploratory baseline. Its model card reports 87.79% accuracy and macro-F1 = 0.8783 on a 10,000-item test split, but its neutral category was added through rule-based supplementation to an originally binary sentiment resource. Its neutral outputs are therefore interpreted cautiously. Multilingual DistilBERT was included as a multilingual comparative baseline. Its model card reports 88.29% agreement between student and teacher predictions, but no Chinese-specific benchmark or Chinese social-media performance is reported. Model identifiers, inference settings, and reproducibility details are provided in Appendix A Table A3, Table A4, Table A5, Table A6 and Table A7.
For statistical analysis, the three sentiment labels were encoded as negative, neutral, and positive. In analyses focusing on whether offensive texts were assigned negative sentiment, labels were also reduced to a binary variable: negative versus non-negative, where neutral and positive were combined as non-negative. This binary operationalisation was used only as an analytical contrast to negative sentiment and should not be interpreted as a positive evaluative function, positive speaker stance, or absence of offensive function.
3.4. Analytical Procedure
The dependent variable was defined separately for each research question. For RQ1 and RQ2, the outcome was the human reference sentiment label in the validation subset. For RQ3, the outcome was model–human agreement between model-generated and human reference sentiment labels. For RQ4, the outcome was the model-generated binary label, negative versus non-negative, in the full corpus.
For annotation reliability, Fleiss’ κ was used to assess agreement among the three independent annotations. For the human sentiment annotation, the 95% confidence interval for Fleiss’ κ was estimated using 10,000 nonparametric bootstrap resamples at the item level, with the three annotator ratings retained within each resampled item.
The analysis proceeded in four steps. First, human reference sentiment labels in the 400-item validation subset were analyzed across non-offensive texts, explicit attacks, and implicit attacks. Three-way sentiment distributions (negative, neutral, and positive) were reported descriptively. For inferential analysis, sentiment labels were reduced to negative versus non-negative (neutral and positive combined), consistent with the study’s analytical focus on whether offensive texts were necessarily associated with negative sentiment. The overall association between text category and this binary sentiment distinction was evaluated using a Pearson chi-square test of independence, with Cramér’s V reported as an effect-size measure. Where the omnibus association was significant, pairwise Pearson chi-square tests without continuity correction were conducted between the three text categories, with Holm adjustment applied to control the family-wise error rate across the three comparisons. The binary labels were also used to evaluate human-perceived negative sentiment as a proxy for offensive-language labels, for which accuracy, sensitivity/recall, specificity, precision, and F1 were reported.
Second, within the implicit-attack portion of the validation subset, human reference sentiment labels were compared across the four standard pragmatic strategies: Irony, Trope, Indirectness, and Exaggeration. The association between pragmatic strategy and the binary sentiment distinction (negative versus non-negative) was evaluated using a Pearson chi-square test of independence, with Cramér’s V reported as an effect-size measure. Given the small number of non-negative observations in some strategy categories, Fisher’s exact test was additionally used as a robustness check.
Third, model-human agreement was assessed in the validation subset using accuracy, Cohen’s kappa, and macro-F1. These analyses were conducted for the full validation subset and separately by text category.
Fourth, model-generated sentiment labels were analyzed across the full 1400-item corpus. Descriptive distributions of negative, neutral, and positive labels were calculated for each model across text categories. A composite text-category variable distinguished non-offensive texts, explicit attacks, and implicit attacks by primary pragmatic strategy: implicit-irony, implicit-trope, implicit-indirectness, implicit-exaggeration, and implicit-unknown. The implicit-unknown category contained 47 of the 400 implicit-attack items and was retained in the full-corpus text-category analysis but excluded from the standard pragmatic-strategy analysis.
To account for the repeated-measures structure of the model outputs, mixed-effects logistic regression was then used, since each text was classified by all four models. The binary outcome was whether a model assigned a negative sentiment label. For the full corpus, negative label assignment was modelled as a function of text category, model, and their interaction, with a random intercept for item:
negative_label ~ model × text_category + (1 | item)
For the standard implicit-attack subset, excluding implicit-unknown cases, negative label assignment was modelled as a function of pragmatic strategy, model, and their interaction:
negative_label ~ model × pragmatic_strategy + (1 | item)
As a sensitivity analysis, the implicit-only mixed-effects model was repeated after excluding compound-label items and UNKNOWN cases. Likelihood-ratio tests were used to evaluate interaction terms. Predicted probabilities were inspected descriptively, and Wilson 95% confidence intervals were reported for the principal negative-label proportions. Qualitative examples were selected from the 400-item validation subset to illustrate recurrent types of model-human divergence and sentiment-offence mismatch; they were not treated as statistically representative cases or as evidence of model-internal mechanisms.
4. Results
4.1. RQ1: Human Reference Sentiment Labels and Offensive-Language Labels
The first analysis examined the distribution of human reference sentiment labels in the 400-item validation subset. As shown in Table 1, negative sentiment was the dominant human-perceived polarity in the validation subset, but it was not the only polarity assigned by human annotators. Across the full subset, 310 items were assigned negative sentiment, accounting for 77.5% of the sample. Neutral sentiment accounted for 81 items (20.25%), while positive sentiment accounted for 9 items (2.25%).
Table 1.
Distribution of human reference sentiment labels by text category.
When the validation subset was divided by text category, explicit attacks showed the highest negative-sentiment rate: 98 of the 100 explicit attacks (98.0%) were assigned negative sentiment and 2 were assigned neutral sentiment. Among the 200 implicit attacks, 184 (92.0%) were assigned negative sentiment and 16 were assigned neutral sentiment. By contrast, the 100 non-offensive texts showed a more heterogeneous distribution, with 28 negative (28.0%), 63 neutral (63.0%), and 9 positive (9.0%) labels.
For inferential analysis, sentiment labels were reduced to negative versus non-negative. The overall association between text category and the negative versus non-negative sentiment distinction was statistically significant, χ2(2, N = 400) = 188.73, p < 0.001, Cramér’s V = 0.687. Holm-adjusted pairwise Pearson chi-square tests showed that negative sentiment was significantly more frequent in explicit attacks than in non-offensive texts, χ2(1) = 105.11, adjusted p < 0.001, and in implicit attacks than in non-offensive texts, χ2(1) = 131.73, adjusted p < 0.001. Explicit attacks also showed a higher negative-sentiment rate than implicit attacks (98.0% vs. 92.0%), although this difference was considerably smaller, χ2(1) = 4.26, adjusted p = 0.039.
Taken together, these results indicate a strong association between human-perceived negative sentiment and offensive-language category. Both explicit and implicit attacks were substantially more likely than non-offensive texts to receive negative sentiment labels. Although explicit attacks showed a somewhat higher negative rate than implicit attacks, this difference was comparatively small. More importantly, the relationship between sentiment polarity and offensive function was not one-to-one.
Using the same negative versus non-negative distinction, we next evaluated how well human-perceived negative sentiment functioned as a proxy for the original offensive-language labels. Negative sentiment was treated as a predicted offensive label and compared with the original offensive-language labels. As shown in Table 2, this proxy analysis produced high recall but lower specificity. The high recall indicates that most offensive texts were assigned negative sentiment by human annotators, whereas the lower specificity indicates that negative sentiment also appeared in a non-negligible proportion of non-offensive texts.
Table 2.
Performance of human negative sentiment as a proxy for offensive-language labels.
Overall, RQ1 shows that the prevalence of human-perceived negative sentiment differed significantly across the three text categories, with negative sentiment substantially more prevalent in explicit and implicit attacks than in non-offensive texts. Nevertheless, the presence of negative sentiment in 28.0% of non-offensive texts and neutral sentiment in 8.0% of implicit attacks demonstrates that sentiment polarity and offensive function are strongly associated but not equivalent.
4.2. RQ2: Pragmatic Strategies and Human Reference Sentiment in Implicit Attacks
The second analysis examined whether different pragmatic strategies in implicit attacks corresponded to different distributions of human reference sentiment labels. This analysis focused on the 200 implicit attacks in the human validation subset, including Irony, Trope, Indirectness, and Exaggeration. As shown in Table 3, human annotators assigned negative sentiment to the majority of implicit attacks across all four pragmatic strategies. No implicit-attack item in the validation subset received a final positive sentiment label.
Table 3.
Distribution of human reference sentiment labels by pragmatic strategy in implicit attacks.
The four pragmatic strategies showed highly similar sentiment distributions. Indirectness had the highest negative rate, with 47 of 50 items assigned negative sentiment (94.0%). Exaggeration followed with 48 of 52 items assigned negative sentiment (92.3%), while Irony and Trope had negative rates of 91.1% and 90.5%, respectively. Neutral sentiment appeared in a small proportion of cases across all four strategies, ranging from 6.0% for Indirectness to 9.5% for Trope.
Because no positive labels occurred in the implicit-attack subset, the statistical comparison was conducted on the binary distinction between negative and neutral sentiment. A chi-square test did not show a significant association between pragmatic strategy and human reference sentiment distribution, χ2(3, N = 200) = 0.48, p = 0.924, Cramér’s V = 0.049. Because some expected cell counts were small, Fisher’s exact test was also used as a robustness check and likewise showed no significant association, p = 0.940.
These results suggest that, in human reference annotations, implicit attacks were generally perceived as negative regardless of the pragmatic strategy used. Unlike the model-generated sentiment labels, which showed more variation across systems and strategies, human reference labels did not provide strong evidence that Irony, Trope, Indirectness, and Exaggeration correspond to substantially different sentiment-polarity distributions. The small number of neutral cases nevertheless remains theoretically relevant, because it shows that some implicit attacks can be perceived as non-negative even by human annotators when sentiment cues are weak, indirect, or context-dependent.
Overall, RQ2 is answered cautiously. In the human validation subset, pragmatic strategy did not significantly predict human reference sentiment polarity. The main pattern was not strategy-specific divergence, but the overall predominance of negative sentiment across all four types of implicit attack, accompanied by a small set of neutral cases. This finding suggests that pragmatic strategy may play a larger role in explaining model-output variation than in explaining human reference sentiment distributions, a point examined further in the model-human comparison below.
4.3. RQ3: Model-Human Agreement in the Validation Subset
The third analysis examined how the sentiment labels produced by the four selected NLP models compared with the human reference sentiment labels in the 400-item validation subset. Unlike the full-corpus model-output analysis reported below, this analysis was restricted to the subset for which human sentiment annotations were available. Model–human agreement was evaluated using accuracy, Cohen’s kappa, and macro-F1.
As shown in Table 4, Claude Haiku 4.5 showed the highest agreement with the human reference labels, with an accuracy of 78.50%, Cohen’s κ = 0.493, and macro-F1 = 0.632. The OpenAI GPT-4.1-mini model followed with an accuracy of 75.00%, Cohen’s κ = 0.457, and macro-F1 = 0.598. Chinese BERT and Multilingual DistilBERT showed lower agreement with the human reference labels in this validation subset. Chinese BERT reached 61.00% accuracy, Cohen’s κ = 0.274, and macro-F1 = 0.472, while Multilingual DistilBERT showed the lowest agreement, with 53.75% accuracy, Cohen’s κ = 0.076, and macro-F1 = 0.335. These differences are reported descriptively and should not be interpreted as inferential evidence of systematic differences among the four models or, more broadly, between model architectures.
Table 4.
Agreement between model-generated sentiment labels and human reference sentiment labels in the validation subset.
Model-human agreement also varied by text category, as shown in Table 5. Agreement was highest for explicit attacks in the two generative models, especially OpenAI GPT-4.1-mini, which reached 92.0%. For implicit attacks, Claude Haiku 4.5 showed the highest agreement at 78.0%, whereas Chinese BERT showed the lowest agreement at 45.5%. The non-offensive subset showed a different pattern: Chinese BERT reached the highest category-level agreement at 72.0%, while Multilingual DistilBERT showed substantially lower agreement at 31.0%.
Table 5.
Model–human agreement by text category in the validation subset.
These category-level results show that model-human agreement was not uniform across text categories. The two generative models aligned more closely with human reference labels in explicit attacks, while agreement became less stable in implicit attacks. The lower agreement of Chinese BERT in the implicit-attack subset suggests that it more often diverged from human reference judgments when offence was conveyed indirectly. By contrast, the low non-offensive accuracy of Multilingual DistilBERT indicates a different type of divergence, namely a tendency to assign sentiment labels that did not match human judgments in non-offensive texts.
Overall, the four selected models differed substantially in their agreement with human reference sentiment labels. Claude Haiku 4.5 and OpenAI GPT-4.1-mini were more closely aligned with human annotations than Chinese BERT and Multilingual DistilBERT in this validation subset, but this pattern should be interpreted as model-specific rather than as evidence of general model-family differences.
4.4. RQ4: Full-Corpus Repeated-Measures Analysis of Model-Generated Negative Labels
The fourth analysis examined model-generated negative sentiment labels across the full corpus of 1400 texts. Because each text was classified by all four models, model outputs were treated as repeated observations clustered within items. The dependent variable was whether a model assigned a negative sentiment label, with neutral and positive labels combined as non-negative.
Descriptive negative-label rates by text category and model are shown in Table 6. Explicit attacks were more consistently assigned negative labels than implicit attacks across all four models. Non-offensive texts also received negative labels at non-negligible rates, especially from Multilingual DistilBERT. These results show that model-generated negative sentiment labels did not map cleanly onto offensive-language categories.
Table 6.
Model-generated negative-label rates by text category and model in the full corpus.
Table 7 reports 95% Wilson confidence intervals for the strategy-specific negative-label rates to quantify the uncertainty associated with these descriptive proportions. The intervals are particularly informative for the smaller Trope (n = 43), Indirectness (n = 57), and Exaggeration (n = 75) subgroups, for which the estimated proportions are less precise.
Table 7.
Model-generated negative-label rates and 95% Wilson confidence intervals by pragmatic strategy and model.
To test these differences while accounting for repeated observations over the same texts, a mixed-effects logistic regression was fitted with text category, model, and their interaction as fixed effects and item as a random intercept. The text-category variable distinguished non-offensive texts, explicit attacks, implicit-irony, implicit-trope, implicit-indirectness, implicit-exaggeration, and implicit-unknown cases. As shown in Table 8, the interaction between text category and model was statistically significant, likelihood-ratio χ2(18) = 262.32, p < 0.001. This indicates that model-generated negative-label probabilities differed across text categories and that these differences were not uniform across the four selected models.
Table 8.
Mixed-effects logistic regression tests for model-generated negative labels.
A second mixed-effects logistic regression focused on the standard implicit-attack subset, excluding the 47 implicit-unknown cases. In this model, negative label assignment was predicted by pragmatic strategy, model, and their interaction, again with a random intercept for item. The model × pragmatic-strategy interaction was not statistically significant, likelihood-ratio χ2(9) = 10.38, p = 0.320.
A sensitivity analysis was also conducted after excluding compound-label items and UNKNOWN cases. This single-standard-strategy subset included 256 implicit attacks: 133 Irony, 34 Trope, 41 Indirectness, and 48 Exaggeration items. The model × pragmatic-strategy interaction remained non-significant, likelihood-ratio χ2(9) = 11.15, p = 0.265.
The predicted probabilities from the implicit-only model showed some descriptive variation across pragmatic strategies. For example, Trope-based attacks tended to receive lower negative-label probabilities than Indirectness-based attacks in several models. However, because the model × pragmatic-strategy interaction was not statistically significant and some strategy groups were small, these differences should be interpreted only as descriptive tendencies rather than as evidence of stable strategy-specific model effects.
Overall, the full-corpus repeated-measures analysis shows that model-generated negative sentiment labels varied substantially across text categories and models. Within implicit attacks, however, neither the primary implicit-only model nor the single-standard-strategy sensitivity analysis provided strong evidence for stable model-specific pragmatic-strategy effects.
5. Discussion
5.1. Sentiment Polarity and Offensive Function as Related but Non-Equivalent Dimensions
The revised analysis shows that sentiment polarity and offensive function are closely related but not equivalent. In the human validation subset, negative sentiment was strongly associated with offensive-language labels: human annotators assigned negative sentiment to 98.0% of explicit attacks and 92.0% of implicit attacks. The binary proxy analysis also showed that using human-perceived negative sentiment as a proxy for offensive-language labels yielded high recall and F1. These findings indicate that sentiment polarity is highly relevant to offensiveness in the present data.
However, the relationship was not one-to-one. In the validation subset, 28.0% of non-offensive texts were also assigned negative sentiment, showing that negative sentiment is not sufficient to establish offensiveness. Conversely, 8.0% of implicit attacks were assigned neutral sentiment by human annotators, showing that offensive function is not always accompanied by clearly negative sentiment. These cases support the view that sentiment polarity and offensive function should be treated as analytically distinct dimensions of online language.
This distinction is especially important for implicit attacks. Offensive function concerns whether an utterance performs a derogatory, exclusionary, stigmatizing, or harmful social act [5,6,7,9]. Sentiment polarity, by contrast, concerns the affective or evaluative orientation expressed in the text [12,13]. The two dimensions often overlap in explicit attacks because direct insults tend to contain salient negative expressions. In implicit attacks, however, offensive meaning may be constructed through irony, figurative association, presupposition, rhetorical questioning, homophonic substitution, or contextual inference [7,8,14,15,16]. In such cases, offensive function may not be fully captured by surface sentiment polarity alone.
The human annotation results also show that sentiment polarity itself is not always straightforward to determine in short online texts. Inter-annotator agreement was moderate rather than high, suggesting that even human readers encounter ambiguity when judging sentiment in decontextualized online comments. This ambiguity may arise from weak affective cues, mixed positive and negative signals, ironic wording, or missing conversational context [22,23]. For this reason, the human labels in this study are treated as human reference labels rather than definitive gold-standard sentiment labels.
Overall, the revised results support a qualified conclusion: sentiment polarity is an informative cue to offensiveness, but it should not be used as a direct proxy for offensive function. For computational studies of offensive language, sentiment polarity is therefore best understood as a complementary analytical dimension rather than as a substitute for offensive-language annotation.
These findings both confirm and extend previous research on Chinese offensive-language detection and implicit offence. Existing resources such as COLD and TOXICN have advanced Chinese offensive- and toxic-language research by providing annotated corpora and benchmarks for classification and model evaluation [18,19]. Research on implicit and covert offence has further shown that offensive meaning may be conveyed through indirect, figurative, coded, or context-dependent forms rather than overtly abusive expressions [5,6,7,8,9]. Previous research on Chinese cyberbullying has also reported differences in sentiment polarity between bullying and non-bullying comments and between explicit and implicit bullying, with explicit bullying tending to be more negative than implicit bullying [24]. The present findings are broadly consistent with this literature: offensive texts were predominantly associated with negative sentiment, while implicit offensive meaning could not be adequately characterized by overt lexical cues or sentiment polarity alone. However, the present study extends this line of research by directly quantifying the extent to which human-perceived negative sentiment corresponds to offensive-language labels and by explicitly separating three analytical layers: original offensive-language labels, independently annotated human sentiment polarity, and model-generated sentiment labels. These layers are further linked to a pragmatics-oriented classification of implicit attacks. This design makes it possible to examine an asymmetry that is not the primary focus of conventional offensive-language classification benchmarks: negative sentiment is common among offensive texts but is not sufficient for offensiveness, while a subset of pragmatically offensive texts can receive non-negative sentiment judgments. The specific contribution of the present corpus is therefore not a new general-purpose offensive-language benchmark, but an analytically enriched test bed for examining where sentiment polarity, pragmatic form, and offensive function converge or diverge in Chinese online discourse.
5.2. Pragmatic Strategies and Sentiment Polarity in Implicit Attacks
The revised results require a cautious interpretation of the role of pragmatic strategies. In the human validation subset, implicit attacks were overwhelmingly assigned negative sentiment across all four pragmatic strategies. Negative sentiment accounted for 91.1% of Irony cases, 90.5% of Trope cases, 94.0% of Indirectness cases, and 92.3% of Exaggeration cases. The differences among strategies were small and not statistically significant. Thus, the human reference annotations do not provide strong evidence that different pragmatic strategies correspond to clearly distinct human-perceived sentiment distributions.
This finding prevents an overstatement of the strategy effect. Pragmatic strategies characterize different ways in which offensive meaning can be constructed, but this does not necessarily mean that they correspond to different sentiment-polarity labels. Human annotators generally perceived implicit attacks as negative regardless of whether the offensive meaning was conveyed through irony, trope, indirectness, or exaggeration. This is consistent with pragmatic accounts of implicit meaning, which emphasize that hearers can recover speaker meaning through contextual inference rather than literal form alone [14,15,16], and with studies showing that covert or implicit offensive language may rely on indirect, figurative, or context-dependent mechanisms [6,7,8,9].
At the same time, the 16 implicit attacks labelled as neutral by human annotators remain analytically meaningful. These cases suggest that some implicit attacks may convey offensive function without strong or stable negative sentiment cues. They are especially relevant for distinguishing sentiment polarity from offensive function, because they show that even human readers may sometimes perceive an utterance as offensive while assigning it neutral sentiment.
The qualitative examples illustrate this point but should not be treated as direct evidence of model-internal mechanisms. For instance, the trope-based item “你去广州看看,广坎达可不是空喊的” (“Go to Guangzhou and take a look; ‘Guang-kanda’ is not just an empty slogan”) was assigned a negative human reference label. The expression “广坎达” blends “广州” (“Guangzhou”) with “瓦坎达” (“Wakanda”), the fictional African country in the Black Panther franchise. In Chinese online discourse, this blend can evoke a culturally and regionally specific racialized association linked to Guangzhou. The offensive target was inferred from the explicit place reference to Guangzhou, the lexical blend, the source annotation, and the derogatory analogy activated by the Wakanda association. Under the available-text condition, the item was interpreted as an implicit regionalized and racialized attack rather than as a neutral reference to a place. This example illustrates how trope-based attacks may depend on background knowledge and associative mapping rather than direct abusive vocabulary. More broadly, figurative interpretation often depends on the interaction of linguistic form, conceptual association, discourse context, and background knowledge [20].
A similar point can be made with the indirectness-based item “性格善良的人全世界多的是,都住到中国来吗?” (“There are kind people all over the world; should they all come live in China?”), which was also assigned a negative human reference label. The negative stance is conveyed through rhetorical implication and exclusionary presupposition rather than through explicit insults. Such cases show that human annotators can often recover negative evaluation from pragmatic cues, while also showing why sentiment polarity cannot be reduced to surface lexical valence alone.
The full-corpus model-output analysis provides a more qualified picture. Descriptively, model-generated negative-label rates varied across pragmatic strategies and models, and trope-based attacks tended to receive lower negative-label probabilities than indirectness-based attacks in several models, especially Chinese BERT. This pattern is consistent with the possibility that trope-based attacks may rely more heavily on figurative association, cultural knowledge, or metaphorical mapping [7,8,20]. However, the repeated-measures mixed-effects model for the standard implicit-attack subset did not show a significant model × pragmatic-strategy interaction. Strategy-level patterns in the model outputs should therefore be interpreted as descriptive tendencies rather than demonstrated causal or mechanism-level effects.
Overall, pragmatic strategies remain useful for describing how implicit attacks are constructed linguistically, but the present evidence does not show that each strategy corresponds to a clearly distinct sentiment-polarity profile. Their role is better understood as part of a broader interpretive context that includes lexical cues, target information, cultural associations, available context, model training data, and label definitions.
5.3. Model–Human Divergence in Sentiment Polarity Judgments
The model–human comparison shows that the four selected NLP models differed in how closely their sentiment labels aligned with human reference labels. In the 400-item validation subset, Claude Haiku 4.5 showed the highest agreement with human annotations, followed by OpenAI GPT-4.1-mini, whereas Chinese BERT and Multilingual DistilBERT showed lower agreement. This pattern indicates that model-generated sentiment labels should not be treated as interchangeable with human-perceived sentiment polarity, especially when texts involve implicit offence or context-dependent evaluative meaning.
The divergence was particularly clear in the implicit-attack subset. Human annotators assigned negative sentiment to 92.0% of implicit attacks, but model-generated negative-label rates varied considerably: Claude Haiku 4.5 assigned negative labels to 75.5% of implicit attacks, OpenAI GPT-4.1-mini to 65.0%, Multilingual DistilBERT to 61.5%, and Chinese BERT to 45.0%. This indicates that all four models, especially Chinese BERT and Multilingual DistilBERT, were more likely than human annotators to assign non-negative labels to implicit attacks.
The following examples illustrate recurrent patterns of model–human divergence identified in the validation subset. They were selected after comparing human reference labels with model-generated labels across text categories and pragmatic strategies, and they are not intended as statistically representative cases or as direct evidence of model-internal mechanisms.
One illustrative example is the irony/marked-form item “写得好,爱紫病爱你
” (“Well written, ‘Aizi disease’ loves you”). The human reference label was negative, whereas all four models assigned positive sentiment. On the surface, “写得好” (“well written”) and “爱你” (“love you”) are positive expressions. However, “爱紫病” is a deliberate homophonic substitution for “艾滋病” (“AIDS”), using sound similarity and altered written form to evoke a stigmatizing disease-related insult. The divergence may therefore reflect lexical-recognition or tokenisation difficulties caused by the non-standard homophonic form, as well as misleading positive surface cues. This case is treated as an example of model-human divergence involving marked form, homophony, and surface valence, rather than as direct evidence of a specifically pragmatic failure.

” (“Well written, ‘Aizi disease’ loves you”). The human reference label was negative, whereas all four models assigned positive sentiment. On the surface, “写得好” (“well written”) and “爱你” (“love you”) are positive expressions. However, “爱紫病” is a deliberate homophonic substitution for “艾滋病” (“AIDS”), using sound similarity and altered written form to evoke a stigmatizing disease-related insult. The divergence may therefore reflect lexical-recognition or tokenisation difficulties caused by the non-standard homophonic form, as well as misleading positive surface cues. This case is treated as an example of model-human divergence involving marked form, homophony, and surface valence, rather than as direct evidence of a specifically pragmatic failure.A second example is the exaggeration-based item “从某种意义上来讲,在座各位都要感谢杨笠,她把大部分小仙女心声说了出来” (“In a sense, everyone here should thank Yang Li; she voiced what many ‘little fairies’ were thinking”). The human reference label was negative, whereas all four models assigned positive sentiment. The surface expression “感谢” (“thank”) appears positive, but the utterance uses exaggerated praise and the gendered expression “小仙女” to construct a critical or mocking stance. This example illustrates how ironic or exaggerated positive wording may produce model outputs that diverge from human reference sentiment.
A third example is “说明现在相亲市场上男的看身家女的看姿色” (“This shows that in today’s matchmaking market, men look at family wealth and women look at appearance”). The human reference label was negative, while Claude Haiku 4.5, OpenAI GPT-4.1-mini, and Chinese BERT assigned neutral sentiment, and Multilingual DistilBERT assigned positive sentiment. The sentence does not contain an explicit insult, but it expresses a generalized evaluative stance toward gender relations in the matchmaking context. This case illustrates a different kind of divergence: when negative evaluation is conveyed through social generalization rather than overt affective vocabulary, models may classify the text as neutral or positive.
As a counterexample, some implicit attacks were assigned negative sentiment by both human annotators and all four models, even when the offensive function was conveyed indirectly. For example, the indirectness-based item “咋就不问问老师为什么说河南人坏话呢?” (“Why not ask why the teacher said bad things about Henan people?”) was assigned a negative human reference label, and all four models also assigned negative sentiment. This case does not fit a simple explanation in which indirectness necessarily leads models away from negative sentiment. Instead, it suggests that some indirect attacks still contain sufficiently salient negative evaluative cues for both humans and models to classify them as negative.
The non-offensive subset showed the opposite side of the sentiment-offence distinction. Human annotators assigned negative sentiment to 28.0% of non-offensive texts, showing that negative sentiment can occur without offensive function. OpenAI GPT-4.1-mini and Claude Haiku 4.5 produced similar negative-label rates in this subset, whereas Multilingual DistilBERT assigned negative labels to 48.0% of non-offensive texts. Such cases illustrate that negative sentiment is not sufficient evidence of offensiveness.
Inter-model disagreement should also be interpreted cautiously. It does not by itself demonstrate that sentiment polarity is an unstable property of the texts. Rather, it indicates that the selected systems produced different labels under the present classification setting. These differences may arise from model-specific training data, domain fit, label definitions, calibration, prompt formulation, language coverage, or classification errors. Accordingly, the present study treats model disagreement as variation in model-generated sentiment labels, not as direct evidence about the intrinsic stability of sentiment polarity.
Overall, the model–human divergence results show that model-generated sentiment labels vary in their correspondence with human reference labels. Claude Haiku 4.5 and OpenAI GPT-4.1-mini were closer to human annotations than the two discriminative classifiers in the validation subset, but this should not be interpreted as evidence that generative models as a class are generally superior. The safer conclusion is that the four selected systems differed in their treatment of implicit offensive texts, and that human reference labels are necessary for interpreting these differences.
5.4. Implications for Modelling Implicit Offensive Language
The revised findings have several implications for computational approaches to implicit offensive language. First, sentiment polarity can provide useful information, but it should not be used as a substitute for offensive-language detection. In the human validation subset, negative sentiment was strongly associated with offensive-language labels, but the relationship was imperfect: some non-offensive texts expressed negative sentiment, and some implicit attacks were assigned neutral sentiment. The model-output results showed a similar mismatch: model-generated negative sentiment labels did not map cleanly onto offensive-language labels across the full corpus.
This distinction has direct implications for content-moderation systems. If negative sentiment is used as a proxy for offensiveness, systems may generate false positives by flagging non-offensive complaints, criticism, or other negatively valenced content as offensive. Conversely, they may generate false negatives when implicit attacks are expressed through neutral, apparently positive, ironic, figurative, or culturally coded forms. The present validation data illustrate both risks: 28.0% of the sampled non-offensive texts received negative human sentiment labels, whereas 8.0% of the sampled implicit attacks received neutral labels. Sentiment analysis may therefore be useful as one feature in a moderation pipeline, but it should not function as a stand-alone decision rule. For implicit offence in particular, a layered moderation pipeline could combine sentiment information with target identification, pragmatic and contextual cues, and dedicated offensive-language classification, while routing uncertain or context-dependent cases for additional review.
Second, the results show the value of human reference sentiment labels for interpreting model-generated outputs. Without human sentiment annotations, it is difficult to determine whether a model output is close to human-perceived sentiment polarity or instead reflects surface lexical cues, model bias, domain mismatch, or differences in label definitions. Claude Haiku 4.5 and GPT-4.1-mini showed higher agreement with the human reference labels than Chinese BERT and Multilingual DistilBERT in the present validation subset, but agreement remained imperfect across all four models. Model-generated sentiment labels should therefore be interpreted as system outputs requiring external validation, not as direct measurements of sentiment polarity [25,26].
Third, the results caution against overgeneralizing from broad model categories. The four systems differ not only in architecture, but also in training data, language coverage, domain fit, label definitions, prompt formulation, decoding parameters, and calibration. Therefore, the findings should be interpreted as differences among the selected models under the present experimental settings, rather than as general evidence that one architecture family is superior to another.
Fourth, the results highlight the importance of context and pragmatic information. Implicit offensive language often depends on inferential cues, including irony, figurative association, presupposition, rhetorical questioning, homophonic substitution, and culturally specific meanings [7,8,12,13,14]. In isolated online comments, these cues may be difficult to interpret even for human annotators, as reflected in the moderate inter-annotator agreement. For computational models, these cases may be particularly challenging when the available input or task formulation does not adequately represent speaker stance, target-directed evaluation, communicative function, or the broader discourse context.
These implications support a layered approach to modelling implicit offensive language. A text may express negative affect without being offensive, and it may perform an offensive act without strong surface-level negative sentiment. Offensive-language detection systems therefore need richer representations than sentiment polarity alone, including pragmatic strategy, implicit target, stance toward the target, contextual assumptions, and discourse-level information. Recent work on functional evaluation has similarly emphasized the need to move beyond aggregate performance measures and examine how hate-speech detection models behave on specific linguistic phenomena and challenging cases [27]. From this perspective, pragmatic strategy can serve not only as an annotation dimension but also as a useful axis for evaluating where offensive-language systems succeed or fail.
5.5. Limitations and Future Directions
Several limitations should be noted. First, the unit of analysis was a single online text item, and full conversational context was not available for all cases. This is particularly relevant to implicit attacks, whose interpretation may depend on prior discourse, target identification, shared cultural knowledge, and interactional cues. Moreover, although the study distinguishes among lexical valence, expressed sentiment, speaker stance, target-directed evaluation, and communicative/offensive function, these dimensions were not independently annotated. Future work could incorporate thread-level context and multi-layer annotation to examine them more directly.
Second, although the full corpus contains 1400 texts, the standard strategy-level analysis is restricted to 353 implicit attacks assigned to the four focal pragmatic categories, and some individual strategy groups are relatively small (Trope, n = 43; Indirectness, n = 57; Exaggeration, n = 75). These subgroup-level estimates therefore have greater uncertainty and should be interpreted cautiously. The relatively small and imbalanced strategy groups also limit statistical precision and the ability to detect modest strategy-specific or interaction effects. Larger and more balanced datasets are needed to assess the robustness of the observed strategy-level patterns.
Third, the pragmatic-strategy taxonomy is an operational coding framework rather than an exhaustive or mutually exclusive classification. Implicit attacks may involve overlapping mechanisms, including irony, figurative association, indirectness, exaggeration, and marked form. Agreement for the primary pragmatic strategy was moderate (Fleiss’ κ = 0.5842), further indicating uncertainty in assigning a single dominant strategy to some items. Although secondary strategies were retained descriptively and a single-standard-strategy sensitivity analysis was conducted, the main analysis relied on one primary strategy per item. Future work could adopt multi-label or hierarchical coding schemes to represent such overlap more directly.
Fourth, the four selected models should not be treated as representative of broader model families. They differ in architecture, training data, language coverage, domain fit, label construction, and inference procedure. This is particularly relevant to the Chinese BERT classifier, whose neutral category was added through rule-based supplementation to an originally binary sentiment resource. The observed differences should therefore be interpreted as model-specific rather than architecture-level effects, and future work should evaluate a broader range of models.
Fifth, the analysis identifies observable output patterns but cannot determine the internal processing mechanisms underlying individual model predictions. The study did not systematically test prompt formulations, label-order or batch-order effects, or repeated-run stability across the full corpus, although a limited stability check was conducted on the 400-item validation subset. The qualitative examples are therefore illustrative rather than diagnostic of model-internal mechanisms. Controlled minimal pairs, perturbation tests, probing experiments, and systematic error analysis could provide stronger evidence about how specific linguistic and pragmatic cues relate to model judgments.
Finally, because the data are drawn from Chinese social-media discourse, the findings should not be assumed to generalize across platforms, languages, or cultural contexts. Cross-platform, multilingual, and cross-cultural replication, ideally incorporating richer multi-turn conversational data, would provide a stronger test of the robustness and generalizability of the present findings.
6. Conclusions
This study examined how the distinction between sentiment polarity and offensive function is manifested in Chinese online attacks. Using a 1400-item corpus derived from TOXICN and its pragmatics-oriented annotation extension, the analysis combined human reference sentiment annotation, model-generated sentiment labels, pragmatic-strategy categories, and repeated-measures statistical modelling. By distinguishing among original offensive-language labels, human reference sentiment labels, and model-generated sentiment labels, the study avoided treating model outputs as direct evidence of either sentiment polarity or offensive function.
Four main findings emerge from the analysis. First, human-perceived negative sentiment was strongly associated with offensive-language labels in the validation subset, but the mapping was imperfect: some non-offensive texts were perceived as negative, while some implicit attacks were assigned neutral sentiment. Second, pragmatic strategies did not produce clearly distinct human sentiment distributions, although they remained useful for describing how implicit offensive meaning was constructed. Third, Claude Haiku 4.5 and OpenAI GPT-4.1-mini aligned more closely with human reference sentiment labels than Chinese BERT and Multilingual DistilBERT in the validation subset, but these differences should be interpreted as model-specific patterns rather than general evidence about model families. Fourth, the full-corpus mixed-effects analysis revealed a significant interaction between model and text category, whereas neither the primary implicit-only analysis nor the sensitivity analysis provided strong evidence of stable model-specific differences across pragmatic strategies.
Overall, the findings provide empirical evidence from Chinese online offensive-language data that sentiment polarity can inform, but cannot replace, offensive-language analysis. Sentiment labels should not be treated as direct substitutes for offensive-language labels, especially when offence is implicit, figurative, homophonic, or context-dependent. The study therefore supports a layered approach to analyzing implicit offensive language, in which sentiment polarity is considered alongside pragmatic strategy, contextual information, speaker stance, target-directed evaluation, and offensive function.
This layered analytical framework can also be adapted in future studies to evaluate implicit offensive language across different datasets, platforms, languages, and model settings, particularly where sentiment-based signals may diverge from pragmatic or offensive function.
Author Contributions
Conceptualization, D.Z. and S.C.; methodology, D.Z. and S.C.; software, D.Z.; validation, D.Z. and S.C.; formal analysis, D.Z.; investigation, D.Z.; resources, D.Z. and S.C.; data curation, D.Z.; writing—original draft preparation, D.Z.; writing—review and editing, D.Z. and S.C.; visualization, D.Z.; project administration, S.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by the Institute of Modern Languages and Linguistics, Fudan University (IDH4307380/005).
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Informed consent was obtained from all annotators involved in the annotation process.
Data Availability Statement
The data supporting the findings of this study are available from the corresponding author upon reasonable request.
Acknowledgments
During the preparation and revision of this manuscript, the authors used ChatGPT (GPT-5.6 Sol, OpenAI) for language editing and formatting support only. The authors have reviewed and edited the output and take full responsibility for the content of this publication. The authors would also like to thank the three anonymous reviewers for their careful reading and constructive comments, which helped improve the clarity and quality of the manuscript.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Reproducibility Materials
To support reproducibility, this appendix documents the data source, sampling procedure, label mapping, annotation files, model inference configurations, prompt text, repeated-run stability checks, file-integrity checksums, and software environment information used in the revised analysis. SHA256 checksums are reported as file-integrity checks rather than as standalone evidence of reproducibility.
Appendix A.1. Data Access and Annotation Files
The source data were derived from the publicly available TOXICN dataset and its pragmatics-oriented annotation extension. Where redistribution of processed online-text data is limited by licencing or ethical considerations, the accompanying materials provide source links, selection criteria, preprocessing steps, and label mappings for reconstructing the 1400-item corpus and the 400-item validation subset.
The pragmatic-strategy distributions for the 400 implicit-attack items are reported in Table A1 and Table A2. These tables document the primary strategy labels, the first secondary strategy labels, and the non-standard labels that were grouped as UNKNOWN for statistical analysis.
Table A1.
Primary pragmatic-strategy labels.
Table A2.
(A) Secondary and non-standard pragmatic-strategy labels. (B) Raw non-standard labels grouped as UNKNOWN.
Appendix A.2. Model Inference Configuration
The model identifiers and available inference configurations used for sentiment polarity annotation are reported in Table A3. For the OpenAI GPT-4.1-mini full-corpus run, the available summary file records 1400 successful outputs and 0 failed outputs. The sentiment-label distribution was 762 negative labels, 486 neutral labels, and 152 positive labels. Token usage was 59,153 input tokens and 18,208 output tokens, for a total of 77,361 tokens. The revised GPT output file name records the run timestamp as 2026-09-03 13:06:33 +08:00.
Table A3.
Model identifiers and inference configurations used for sentiment polarity annotation.
Appendix A.3. Prompt for Generative Models
The following sentiment analysis prompt was used for the generative models:
You are a Chinese text sentiment analysis annotator. Please determine the overall sentiment polarity expressed by the speaker in each text.
Task requirements:
1. Classify each Chinese text into one of three categories: negative, neutral, or positive.
2. negative indicates negative sentiment, neutral indicates neutral sentiment, and positive indicates positive sentiment.
3. Only the three labels negative, neutral, and positive are allowed.
4. Return a JSON array, with each item in the following format:
{“id”: number, “sentiment”: “negative|neutral|positive”}
5. Do not omit any id.
6. Do not output explanations, reasons, or any additional text.
Texts to annotate:
[
{“id”: 0, “text”: “…”},
{“id”: 1, “text”: “…”}
]
Appendix A.4. Generative-Model Repeated-Run Stability Check
A limited repeated-run stability check was conducted for the two generative models in the 400-item validation subset. The repeated-run labels were compared with the corresponding full-corpus-run labels for the same items. This analysis evaluates within-model output stability under the documented inference setup, but it does not replace a full prompt-robustness analysis involving alternative prompts, label orders, item orders, or batch compositions.
Table A4.
Repeated-run stability check for the two generative models in the 400-item validation subset.
Appendix A.5. File Integrity Checksums
SHA256 checksums are provided as file-integrity checks for the input files, model outputs, scripts, repeated-run stability outputs, and statistical files used in the revised analysis.
Table A5.
SHA256 checksums for input data, intermediate data, model outputs, and statistical test files.
Table A6.
SHA256 checksums for the inference scripts used in sentiment polarity annotation.
Appendix A.6. Software and Environment Information
Table A7.
Software and environment information used for local inference and statistical analysis.
Appendix B. Representative Examples of Pragmatic Strategies
Table A8 provides representative examples of the four pragmatic-strategy categories used in the annotation scheme. The examples illustrate how each category was operationalized in the present study, including major subtypes of Trope and Exaggeration.
Table A8.
Representative examples of the pragmatic-strategy categories used for implicit attacks.
References
- Bilewicz, M.; Soral, W. Hate speech epidemic: The dynamic effects of derogatory language on intergroup relations and political radicalization. Political Psychol. 2020, 41, 3–33. [Google Scholar] [CrossRef] [Scilit]
- Schmidt, A.; Wiegand, M. A survey on hate speech detection using natural language processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, Valencia, Spain, 3 April 2017; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
- Davidson, T.; Warmsley, D.; Macy, M.; Weber, I. Automated hate speech detection and the problem of offensive language. Proc. Int. AAAI Conf. Web Soc. Media 2017, 11, 512–515. [Google Scholar] [CrossRef] [Scilit]
- Caselli, T.; Basile, V.; Mitrović, J.; Granitzer, M. HateBERT: Retraining BERT for abusive language detection in English. In Proceedings of the 5th Workshop on Online Abuse and Harm, Online, 6 August 2021; pp. 17–25. [Google Scholar] [CrossRef] [Scilit]
- ElSherief, M.; Ziems, C.; Muchlinski, D.; Anupindi, V.; Seybolt, J.; De Choudhury, M.; Yang, D. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, 7–11 November 2021; pp. 345–363. [Google Scholar] [CrossRef] [Scilit]
- Bączkowska, A.; Lewandowska-Tomaszczyk, B.; Žitnik, S.; Liebeskind, C.; Trojszczak, M.; Valunaite Oleskeviciene, G. Implicit offensive language taxonomy. Lodz Pap. Pragmat. 2024, 20, 463–483. [Google Scholar] [CrossRef] [Scilit]
- Ocampo, N.B.; Sviridova, E.; Cabrio, E.; Villata, S. An in-depth analysis of implicit and subtle hate speech messages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Dubrovnik, Croatia, 2–6 May 2023; pp. 1997–2013. [Google Scholar] [CrossRef] [Scilit]
- Parvaresh, V. Covertly communicated hate speech: A corpus-assisted pragmatic study. J. Pragmat. 2023, 205, 63–77. [Google Scholar] [CrossRef] [Scilit]
- Lewandowska-Tomaszczyk, B.; Bączkowska, A.; Liebeskind, C.; Valunaite Oleskeviciene, G.; Žitnik, S. An integrated explicit and implicit offensive language taxonomy. Lodz Pap. Pragmat. 2023, 19, 7–48. [Google Scholar] [CrossRef] [Scilit]
- Xiao, Y.; Hu, Y.; Choo, K.T.W.; Lee, R.K.-W. ToxiCloakCN: Evaluating robustness of offensive language detection in Chinese with cloaking perturbations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Miami, FL, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Ji, S.; Zhong, K.; Peng, H.; Xiao, Z.; Liu, X.; Wei, W. Enhancing Chinese offensive language detection with homophonic perturbation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Suzhou, China, 2025; pp. 22660–22675. [Google Scholar] [CrossRef] [Scilit]
- Pang, B.; Lee, L. Opinion mining and sentiment analysis. Found. Trends Inf. Retr. 2008, 2, 1–135. [Google Scholar] [CrossRef] [Scilit]
- Liu, B. Sentiment Analysis and Opinion Mining; Morgan & Claypool Publishers: San Rafael, CA, USA, 2012. [Google Scholar] [CrossRef] [Scilit]
- Grice, H.P. Logic and conversation. In Syntax and Semantics, Vol. 3: Speech Acts; Cole, P., Morgan, J.L., Eds.; Academic Press: New York, NY, USA, 1975; pp. 41–58. [Google Scholar]
- Searle, J.R. Expression and Meaning: Studies in the Theory of Speech Acts; Cambridge University Press: Cambridge, UK, 1979. [Google Scholar]
- Sperber, D.; Wilson, D. Relevance: Communication and Cognition, 2nd ed.; Blackwell: Oxford, UK, 1995. [Google Scholar]
- Bach, K. Conversational impliciture. Mind Lang. 1994, 9, 124–162. [Google Scholar] [CrossRef] [Scilit]
- Lu, J.; Xu, B.; Zhang, X.; Min, C.; Yang, L.; Lin, H. Facilitating fine-grained detection of Chinese toxic language: Hierarchical taxonomy, resources, and benchmarks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, ON, Canada, 9–14 July 2023; pp. 16235–16250. [Google Scholar] [CrossRef] [Scilit]
- Deng, J.; Zhou, J.; Sun, H.; Zheng, C.; Mi, F.; Meng, H.; Huang, M. COLD: A benchmark for Chinese offensive language detection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 11580–11599. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT 2019, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
- Gibbs, R.W.; Colston, H.L. Interpreting Figurative Meaning; Cambridge University Press: Cambridge, UK, 2012. [Google Scholar] [CrossRef] [Scilit]
- Burgers, C.; van Mulken, M.; Schellens, P.J. Type of evaluation and marking of irony: The role of perceived complexity and comprehension. J. Pragmat. 2012, 44, 231–242. [Google Scholar] [CrossRef] [Scilit]
- Frenda, S.; Cignarella, A.T.; Basile, V.; Bosco, C.; Patti, V.; Rosso, P. The unbearable hurtfulness of sarcasm. Expert Syst. Appl. 2022, 193, 116398. [Google Scholar] [CrossRef] [Scilit]
- Zhong, J.; Qiu, J.; Sun, M.; Jin, X.; Zhang, J.; Guo, Y.; Qiu, X.; Xu, Y.; Huang, J.; Zheng, Y. To Be Ethical and Responsible Digital Citizens or Not: A Linguistic Analysis of Cyberbullying on Social Media. Front. Psychol. 2022, 13, 861823. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Das, A.; Zhang, Z.; Hasan, N.; Sarkar, S.; Jamshidi, F.; Bhattacharya, T.; Rahgouy, M.; Raychawdhary, N.; Feng, D.; Jain, V.; et al. Investigating annotator bias in large language models for hate speech detection. arXiv 2024, arXiv:2406.11109. [Google Scholar] [CrossRef] [Scilit]
- Zhang, M.; He, J.; Ji, T.; Lu, C.-T. Don’t go to extremes: Revealing the excessive sensitivity and calibration limitations of LLMs in implicit hate speech detection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Bangkok, Thailand, 11–16 August 2024; pp. 12073–12086. [Google Scholar] [CrossRef] [Scilit]
- Röttger, P.; Vidgen, B.; Nguyen, D.; Waseem, Z.; Margetts, H.; Pierrehumbert, J. HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 41–58. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.