Despite advances in computational social science, online conflict discourse is still commonly analyzed through sentiment, stance, or topic models, leaving limited insight into the social and psychological mechanisms embedded in digital war narratives. This study develops a theory-informed probabilistic framework for detecting discourse-level
[...] Read more.
Despite advances in computational social science, online conflict discourse is still commonly analyzed through sentiment, stance, or topic models, leaving limited insight into the social and psychological mechanisms embedded in digital war narratives. This study develops a theory-informed probabilistic framework for detecting discourse-level indicators of Social Identity Theory, Moral Foundations Theory, Threat Appraisal, Cognitive Distortion, and Deindividuation in Russia–Ukraine war discourse. The empirical design uses 48,201 tweets in total: 10,815 tweets collected between 1 January and 28 June 2022 for model development and primary analysis, as well as and an external validation corpus of 37,386 Russia–Ukraine cyberwar-related tweets—collected from 30,706 users across 54 languages between October 2022 and April 2023—for temporal robustness assessment. The primary corpus contained 10,815 unique tweet identifiers, 10,229 unique textual records, 586 repeated textual items, a textual uniqueness rate of 94.58%, 6646 English tweets (61.45%), and 32,260 retweet engagements. Methodologically, the framework combines contextual language representations, theory-aligned linguistic cues, temporal signals, engagement features, and graph-based indicators. These signals are used to infer latent constructs and are evaluated through calibration, ablation testing, human validation, and cascade comparison. Empirically, Deindividuation was the dominant construct (1654 posts, 15.29%), followed by Cognitive Distortion (525, 4.85%) and Threat Appraisal (503, 4.65%). Co-activation analysis showed the strongest overlap between Deindividuation and Cognitive Distortion (Jaccard = 0.26). Validation diagnostics indicated internal lexical consistency (
for Deindividuation), 93% rumor calibration, 91% bootstrap stability, and improved baseline performance (F1 = 0.72; Brier = 0.12; cascade log-likelihood = −865). The findings demonstrate that theoretically grounded probabilistic modeling can provide scalable, interpretable, and temporally validated insight into psychological patterns in digital conflict discourse.
Full article