1. Introduction
In recent years, computational social science has made major strides in modeling online behavior, yet much of this work remains limited to sentiment, stance, or topic-based analyses that insufficiently capture the deeper psychological mechanisms driving collective action in conflict discourse [
1,
2,
3,
4]. Prior studies show that identity signaling and intergroup dynamics shape online polarization [
5,
6,
7] and that moral framing amplifies diffusion and virality in contentious debates [
8,
9,
10,
11]. However, challenges persist in translating nuanced constructs such as deindividuation, cognitive distortions, threat appraisal, and rumor transmission into computationally tractable models [
12,
13]. The Russia–Ukraine conflict, with its intense digital information warfare, highlights this gap: while research has documented polarization, cascades, and misinformation dynamics [
14,
15,
16,
17,
18], few studies systematically connect these observed patterns to established social–psychological theories [
19,
20,
21]. This motivates the present study, which integrates probabilistic modeling with established behavioral theories to bridge the divide between psychological constructs and computational analysis, thereby offering a more theoretically grounded explanation of cyberwar discourse.
For readability, the contribution can be separated into three linked gaps. The theoretical gap is that established social psychological constructs are rarely operationalized in computational analyses of war discourse. The methodological gap is that sentiment, stance, and topic models provide limited explanation of identity, threat, moral framing, and crowd based discourse mechanisms. The empirical gap is that these constructs require validation through corpus diagnostics, human coding, temporal comparison, and robustness testing rather than through lexical matching alone.
As seen from
Figure 1, this study develops a systematic framework to analyze social media posts in order to detect underlying psychological and social theories of behavior. The process begins with the collection of raw posts, which are carefully preprocessed to extract both linguistic information (such as words, grammar, and entities) and contextual metadata (such as timing and user networks). These inputs are then transformed into a rich representation that combines textual features, lexical cues, temporal markers, and social-network signals. In practical terms, textual features capture what is said, temporal and engagement features capture when and how strongly messages spread, and network features capture how messages cluster across discourse communities.
From this representation, we infer latent constructs that capture fundamental behavioral dimensions, including hostility, deindividuation, mobilization readiness, and threat perception. Auxiliary layers are applied to identify moral foundations (such as fairness, loyalty, or authority) and framing roles (diagnostic, prognostic, motivational). Together, these latent constructs and auxiliary layers feed into specialized detection modules that operationalize established theories—including Social Identity Theory, Realistic Conflict Theory, Moral Foundations Theory, General Aggression Model, Threat Appraisal Theory, Theory of Planned Behavior, and others. To address the temporal boundedness of the primary corpus, the study also incorporates an external temporal validation comparison using a previously published dataset of 37,386 Russia–Ukraine cyberwar-related tweets [
22]. The primary dataset captures the onset and early escalation phase of the conflict, whereas the validation corpus represents a later and more routinized phase of online war discourse. This design allows the analysis to retain its focus on the psychologically intense onset period while examining whether the observed construct patterns remain visible, attenuate, or diffuse over time.
The framework also examines polarization and diffusion at the network level. Model outputs are checked through calibration, reliability testing, human coding, and temporal validation. Together, these steps connect computational text analysis with social psychological theory and support a clearer interpretation of online collective behavior. The selected theories are appropriate for this context because they correspond to recurrent mechanisms in online conflict discourse. Social Identity Theory captures in-group and out-group categorization, which is central to war-related polarization. Moral Foundations Theory captures the moral vocabularies through which users frame harm, justice, loyalty, authority, and purity. Threat Appraisal explains how users represent danger, vulnerability, and response urgency under conditions of uncertainty. Cognitive Distortion captures simplified, absolutist, or exaggerated interpretations of complex geopolitical events. Deindividuation is relevant because hashtag-driven, synchronized, and crowd-oriented online discourse can reduce individual salience while amplifying collective expression. The study therefore treats these theories not as directly observable psychological states but as discourse-level constructs that can be probabilistically inferred from textual, temporal, and network indicators.
The following describe the core contributions of this study:
The study provides a theoretical contribution by operationalizing Social Identity Theory, Moral Foundations Theory, Threat Appraisal, Cognitive Distortion, and Deindividuation within a probabilistic detection framework, bridging psychological theory with computational modeling.
Deindividuation emerged as the dominant construct with 15.3% of posts, followed by Cognitive Distortion with 4.9% and Threat Appraisal with 4.7%, reflecting the psychological underpinnings of conflict discourse.
Construct co-occurrence analysis revealed systematic overlaps, most notably between Deindividuation and Distortions with a Jaccard index of 0.26, while reliability diagnostics confirmed robustness with lexical alignment up to , a calibration of 93%, and a bootstrap stability of 91%.
Cascade dynamics highlighted heavy-tailed diffusion patterns where the majority of posts generated little amplification but a minority triggered disproportionately large retweet cascades, underscoring asymmetric influence in cyberwarfare discourse.
The findings demonstrate significant implications for crisis informatics, counter-disinformation strategies, and platform governance by enabling theoretically grounded and empirically validated monitoring of psychological constructs in digital conflict.
3. Methodology
3.7. Practical Operationalization Pipeline
The computational pipeline followed eight steps. First, raw tweet records were cleaned by removing empty records and standardizing text fields. Second, repeated textual content was identified to avoid inflation of construct prevalence. Third, language labels, hashtags, mentions, URLs, and engagement variables were extracted. Fourth, transformer-based contextual embeddings were generated from tweet text. Fifth, theory-aligned lexical and syntactic cues were extracted, including collective pronouns, threat terms, moral vocabulary, conditional threat templates, and mobilization verbs. Sixth, these features were combined into a multiview representation. Seventh, theory-specific probabilistic heads estimated discourse-level construct activation. Eighth, construct outputs were evaluated through co-occurrence, calibration, bootstrap stability, and diffusion association analyses.
Figure 2 depicts the overall pipeline. To further clarify the computational implementation, Algorithms 1–3 provide pseudocode descriptions of the end-to-end procedure, theory-specific construct inference, and validation workflow.
| Algorithm 1 End-to-End Pipeline for Theory-Informed Social Media Construct Detection |
Require: Raw tweet corpus Require: Theory set , lexicons , interaction graphs Ensure: Construct activations , latent scores , validation diagnostics 1: Remove empty records and standardize text fields 2: Identify duplicate tweet identifiers and repeated textual content 3: Extract language labels, hashtags, mentions, URLs, timestamps, and engagement variables 4: Generate contextual embedding for each tweet using a Twitter-oriented transformer encoder 5: Extract linguistic cues: , , and theory-aligned lexical indicators 6: Extract temporal features from and graph features from and 7: Construct multiview feature vector 8: Infer continuous latent constructs 9: Estimate moral foundation proportions and frame role indicators 10: for each theory do 11: Estimate posterior activation probability 12: Convert posterior probability into binary activation using threshold 13: Compute construct prevalence, co-activation, temporal dynamics, and diffusion diagnostics 14: Return and aggregate validation outputs
|
| Algorithm 2 Theory-Specific Probabilistic Construct Inference |
Require: Feature vector , latent constructs , moral proportions , frame indicators Require: Theory-specific parameters Ensure: Posterior probabilities and construct decisions 1: for each tweet do 2: Estimate hostility, deindividuation, mobilization readiness, and threat appraisal scores
3: Estimate moral foundation distribution 4: Detect diagnostic and prognostic frame indicators from frame role sequence 5: for each theory do 6: Combine textual, latent, moral, and framing evidence 7: Estimate theory activation probability 8: Compute posterior odds 9: if then 10: 11: else 12: 13: Return
|
| Algorithm 3 Validation, Robustness, and External Temporal Comparison |
Require: Construct activations , posterior probabilities , timestamps , retweet counts, interaction graphs Require: External validation corpus Ensure: Reliability, calibration, diffusion, and temporal validation diagnostics 1: Compute construct prevalence for each theory k 2: Estimate pairwise construct co-activation and Jaccard indices 3: Assess internal diagnostic consistency between construct scores and theory-aligned indicators 4: Calibrate posterior probabilities using temperature scaling where language group size is sufficient 5: Evaluate robustness through bootstrap resampling and stability checks 6: Estimate temporal construct trajectories by aggregating over weekly intervals 7: Model cascade dynamics using retweet engagement and cascade association diagnostics 8: Compare early-onset corpus with external validation corpus 9: Interpret whether construct signals remain stable, attenuate, or become more diffuse over time 10: Return diagnostic metrics, temporal patterns, and external validation summary
|
To improve methodological transparency for non-specialist readers, the pipeline can be read as four simple stages: data preparation, feature construction, construct inference, and validation. Data preparation removes empty or repeated records; feature construction converts each tweet into textual, lexical, temporal, engagement, and network signals; construct inference estimates whether theory-aligned discourse indicators are present; and validation compares these estimates with robustness tests, temporal comparison, and human coding.
Theory-aligned lexicons were constructed by mapping each target construct to high-precision cues derived from the relevant theory and then refining the cue lists for conflict discourse. For example, deindividuation cues included collective pronouns, crowd references, hashtag uniformity, and synchronized expression; threat appraisal cues included danger terms and conditional threat templates; cognitive distortion cues included absolutist, overgeneralized, and exaggerated claims; and rumor cues included uncertainty and breaking-news markers. These lexicons were not used as stand-alone classifiers but were combined with BERTweet embeddings, syntactic features, temporal markers, engagement variables, and graph indicators.
Algorithms 1–3 clarify that the proposed framework does not infer individual psychological states directly. Rather, it estimates discourse-level construct activations by combining transformer-based representations, theory-aligned lexical and syntactic cues, latent construct scores, moral foundation proportions, frame indicators, and diffusion diagnostics. The pseudocode also distinguishes the primary inference pipeline from the external temporal validation procedure, thereby improving reproducibility and interpretability without altering the mathematical formulation of the model. The implementation details corresponding to Algorithms 1–3 are provided in the public reproducibility repository at
https://github.com/DrSufi/RU_Social_Psychological (accessed on 7 July 2026). The repository includes the algorithmic codebase for feature construction, theory-specific probabilistic construct inference, validation metric computation, temporal diagnostics, and aggregate reproducibility outputs, while excluding raw tweet text, user identifiers, and reconstructable network data in accordance with the ethical safeguards of this study.
4. Results
The empirical analysis examined 10,815 Russia–Ukraine related tweet records collected between 1 January and 28 June 2022. The dataset contained 10,815 unique tweet identifiers and 10,229 unique textual records, indicating that 586 records contained repeated tweet text. The average tweet length was 155.32 characters, with a median of 143 characters and a range from 8 to 320 characters. Retweet engagement was highly skewed: 3303 tweets received at least one retweet, while the corpus generated 32,260 total retweet engagements. This section presents descriptive corpus diagnostics, theoretical construct prevalence, interactional structures, validation evidence, and diffusion properties.
Author Contributions
Conceptualization, F.S.; methodology, F.S.; software, F.S.; validation, F.S. and F.Z.; formal analysis, F.S.; investigation, F.S.; resources, F.Z.; data curation, F.S.; writing—original draft preparation, F.S.; writing—review and editing, F.S. and F.Z.; visualization, F.S.; supervision, F.Z.; project administration, F.Z.; All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study received an exemption from formal ethics approval from the Department of Management, University of Dhaka, because the research relied solely on publicly available X (formerly Twitter) posts, involved no direct interaction with users, and reported findings only in aggregated and anonymized form.
Informed Consent Statement
The study received an exemption from formal informed consent from the Department of Management, University of Dhaka, because the research relied solely on publicly available X (formerly Twitter) posts, involved no direct interaction with users, and reported findings only in aggregated and anonymized form.
Data Availability Statement
The analytical outputs supporting this study are available from the corresponding author upon reasonable request and subject to platform terms and privacy restrictions. Public sharing is limited to aggregated statistics, summary tables, derived construct counts, and reproducible code where permitted. Raw tweet text, user identifiers, user tweet mappings, and reconstructable network data are not released. The theory-aligned coding protocol, construct validation template, and pseudocode-based implementation notes are available at:
https://github.com/DrSufi/RU_Social_Psychological (accessed on 7 July 2026).
Acknowledgments
Autonomous social data acquisition and structuring mechanism was facilitated by the COEUS Institute’s GERA Platform:
https://coeus.institute/gera/ (accessed on 8 May 2026).
Conflicts of Interest
Author Fahim Sufi was employed by the company COEUS Institute. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| SIT | Social Identity Theory |
| RCT | Realistic Conflict Theory |
| MFT | Moral Foundations Theory |
| TPB | Theory of Planned Behavior |
| NLP | Natural Language Processing |
| CR | F Conditional Random Field |
| JS | Jensen–Shannon |
| LDA | Latent Dirichlet Allocation |
| ECE | Expected Calibration Error |
| BERT | Bidirectional Encoder Representations from Transformers |
References
- Emmert-Streib, F.; Dehmer, M. Data-Driven Computational Social Network Science: Predictive and Inferential Models for Web-Enabled Scientific Discoveries. Front. Big Data 2021, 4, 591749. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Egger, R.; Yu, J. A Topic Modeling Comparison Between LDA, NMF, Top2Vec, and BERTopic to Demystify Twitter Posts. Front. Sociol. 2022, 7, 886498. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shokouhyar, S.; Ahmadi, S.; Ashrafzadeh, M. Promoting a novel method for warranty claim prediction based on social network data. Reliab. Eng. Syst. Saf. 2021, 216, 108010. [Google Scholar] [CrossRef] [Scilit]
- Sufi, F.K.; Razzak, I.; Khalil, I. Tracking anti-vax social movement using AI-based social media monitoring. IEEE Trans. Technol. Soc. 2022, 3, 290–299. [Google Scholar] [CrossRef] [Scilit]
- Vahed, S.; Galván, E.P.; Sanfey, A.G. Computational modeling of social decision-making. Curr. Opin. Psychol. 2024, 60, 101884. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cheng, Z.; Lin, Z.; Yang, Y.; Wei, Z.; Chen, S. Interactive simulation and visual analysis of social media event dynamics with LLM-based multi-agent modeling. Vis. Inform. 2025, 9, 100260. [Google Scholar] [CrossRef] [Scilit]
- Melnikoff, D.E.; Carlson, R.W.; Stillman, P.E. A computational theory of the subjective experience of flow. Nat. Commun. 2022, 13, 2252. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sipilä, J.; Tarkiainen, A.; Levänen, J. Exploration of public discussion around sustainable consumption on social media. Resour. Conserv. Recycl. 2024, 204, 107505. [Google Scholar] [CrossRef] [Scilit]
- Ognibene, D.; Wilkens, R.; Taibi, D.; Hernández-Leo, D.; Kruschwitz, U.; Donabauer, G.; Theophilou, E.; Lomonaco, F.; Bursic, S.; Lobo, R.A.; et al. Challenging social media threats using collective well-being-aware recommendation algorithms and an educational virtual companion. Front. Artif. Intell. 2022, 5, 654930. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lewandowsky, S.; Ecker, U.K.; Cook, J. Beyond Misinformation: Understanding and Coping with the “Post-Truth” Era. J. Appl. Res. Mem. Cogn. 2017, 6, 353–369. [Google Scholar] [CrossRef] [Scilit]
- Zollo, F.; Novak, P.K.; Vicario, M.D.; Bessi, A.; Mozetič, I.; Scala, A.; Caldarelli, G.; Quattrociocchi, W. Emotional Dynamics in the Age of Misinformation. PLoS ONE 2015, 10, e0138740. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Donkers, T.; Ziegler, J. De-sounding echo chambers: Simulation-based analysis of polarization dynamics in social networks. Online Soc. Netw. Media 2023, 37, 100275. [Google Scholar] [CrossRef] [Scilit]
- Caled, D.; Silva, M.J. Digital media and misinformation: An outlook on multidisciplinary strategies against manipulation. J. Comput. Soc. Sci. 2021, 5, 123–159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shah, S.M.; Gillani, S.A.; Baig, M.S.A.; Saleem, M.A.; Siddiqui, M.H. Advancing depression detection on social media platforms through fine-tuned large language models. Online Soc. Netw. Media 2025, 46, 100311. [Google Scholar] [CrossRef] [Scilit]
- Bukar, U.A.; Sayeed, M.S.; Amodu, O.A.; Razak, S.F.A.; Yogarayan, S.; Othman, M. Leveraging VOSviewer approach for mapping, visualisation, and interpretation of crisis data for disaster management and decision-making. Int. J. Inf. Manag. Data Insights 2025, 5, 100314. [Google Scholar] [CrossRef] [Scilit]
- Arora, S.D.; Singh, G.P.; Chakraborty, A.; Maity, M. Polarization and social media: A systematic review and research agenda. Technol. Forecast. Soc. Change 2022, 183, 121942. [Google Scholar] [CrossRef] [Scilit]
- Yurkova, O.; Smola, L.; Kasianchuk, V. Informal communications in the conditions of the Russian-Ukrainian war. Soc. Sci. Humanit. Open 2025, 12, 101702. [Google Scholar] [CrossRef] [Scilit]
- Erokhin, D.; Komendantova, N. Social media data for disaster risk management and research. Int. J. Disaster Risk Reduct. 2024, 114, 104980. [Google Scholar] [CrossRef] [Scilit]
- Garg, M. WellXplain: Wellness concept extraction and classification in Reddit posts for mental health analysis. Knowl.-Based Syst. 2024, 284, 111228. [Google Scholar] [CrossRef] [Scilit]
- Gupta, U.; Jain, V. Social Neuroscience: Inferring Mental States in Social Media. In Emotional AI and Human-AI Interactions in Social Networking; Academic Press: London, UK, 2024. [Google Scholar] [CrossRef] [Scilit]
- Hossain, M.A.; Quaddus, M.; Akter, S.; Mikalef, P.; Warren, M. Trolling in social media: A deindividuation and contagion perspective. Inf. Manag. 2025, 62, 104211. [Google Scholar] [CrossRef] [Scilit]
- Sufi, F. Social Media Analytics on Russia—Ukraine Cyber War with Natural Language Processing: Perspectives and Challenges. Information 2023, 14, 485. [Google Scholar] [CrossRef] [Scilit]
- Vyas, P.; Reisslein, M.; Rimal, B.P.; Vyas, G.; Basyal, G.P.; Muzumdar, P. Automated classification of societal sentiments on Twitter with machine learning. IEEE Trans. Technol. Soc. 2021, 3, 100–110. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Jin, L.; Li, X.; Sun, X.; Wang, X.; Zhang, Z.; Liu, J.; Lu, Z.; Xu, G. Flexible Optimal Transport with Contrastive Graphical Modeling For Multimodal Hate Detection. IEEE Trans. Multimed. 2025, 27, 6397–6409. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Zhao, J.; Chi, C.; Wang, C.; Zhao, X.; Jain, A. Federated learning-inspired user personality prediction using sentiment analysis and topic preference. IEEE Trans. Consum. Electron. 2023, 70, 2729–2737. [Google Scholar] [CrossRef] [Scilit]
- Sufi, F.K. A New Computational Method for Quantification and Analysis of Media Bias in Cybersecurity Reporting. IEEE Trans. Comput. Soc. Syst. 2025, 12, 4561–4570. [Google Scholar] [CrossRef] [Scilit]
Figure 1.
Infographic summary of the proposed probabilistic framework for modeling latent social psychological theories in Russia and Ukraine war discourse on social media. The figure shows the flow from raw tweets to feature extraction, latent construct inference, theory-specific detection, and validation diagnostics.
Figure 2.
Eight-step operationalization pipeline.
Figure 3.
Temporal dynamics of the top four constructs (January–June 2022). Weekly counts are shown for Deindividuation, Cognitive Distortion, Threat Appraisal, and lAggression/Deterrence. Peaks in late January and February correspond to the escalation period preceding and surrounding the full-scale invasion of Ukraine on 24 February 2022, while later stabilization indicates rhetorical normalization.
Figure 4.
Normalized radar plots comparing Pro-Ukraine and Pro-Russia discourse across Care, Fairness, Loyalty, Authority, and Sanctity. Axis values represent normalized moral foundation proportions within each stance-based subgroup.
Figure 5.
Full hashtag co-occurrence network (unpruned). The dense visualization (modularity
) is provided for completeness; the main text uses a pruned view (
Figure 6) to aid legibility and interpretation.
Figure 6.
Hashtag co-occurrence network (top-50 hashtags). Nodes are hashtags; edges indicate within-tweet co-occurrence (weighted). Colors denote Louvain communities; modularity evidences moderate topical polarization. Special characters are offensive & degradory words (i.e., not suitable for display).
Figure 7.
Cascade dynamics of diffusion. (a) Retweet cascade sizes follow a heavy-tailed distribution, indicating strong amplification by a minority of items. (b) Conversation thread sizes are comparatively shallow, indicating limited deliberative depth relative to broadcast-style spread.
Table 1.
Notation and meaning of symbols used in the methodology.
| Symbol | Meaning |
|---|
| Number of posts, users, and communities |
| Text and timestamp of post t |
| Author index of post t |
| Retweet and reply graphs |
| Community label of user u |
| Contextual embedding of |
| Feature vector combining text, parse, lexicon, temporal, and graph features |
| Latent continuous constructs |
| Moral Foundations proportions |
| Frame role sequence |
| Binary indicator that theory k is activated in post t |
| Modularity of stance-signed graph |
| Jensen–Shannon divergence of moral or topical distributions |
Table 2.
Corpus statistics and preprocessing diagnostics for the Russia–Ukraine tweet set (January–June 2022).
| Metric | Value | Notes / Remarks |
|---|
| Total tweet records | 10,815 | Full analytical corpus after record-level cleaning |
| Unique tweet identifiers | 10,815 | No duplicate Tweet_ID values in the v2 dataset |
| Unique tweet texts | 10,229 | 94.58% unique textual records |
| Repeated tweet texts | 586 | 5.42% repeated textual records |
| Average tweet length (characters) | 155.32 | Mean length of tweet_content |
| Median tweet length (characters) | 143 | Median as robustness check |
| Minimum tweet length | 8 | Shortest post observed |
| Maximum tweet length | 320 | Longest post observed |
| Tweets receiving at least one retweet | 3303 | 30.54% of corpus |
| Total retweet engagements | 32,260 | Sum of retweets_count across the corpus |
| Maximum retweet count | 1916 | Largest observed retweet count for a single tweet |
Table 3.
Language distribution in the Russia–Ukraine tweet set (January–June 2022).
| Language Code/Category | Count | Percentage (%) |
|---|
| en (English) | 6646 | 61.45 |
| Non-linguistic (hashtags/media) | 1462 | 13.52 |
| ja (Japanese) | 465 | 4.30 |
| de (German) | 423 | 3.91 |
| fr (French) | 320 | 2.96 |
| und (Undefined) | 210 | 1.94 |
| es (Spanish) | 204 | 1.89 |
| uk (Ukrainian) | 142 | 1.31 |
| ru (Russian) | 90 | 0.83 |
| th (Thai) | 75 | 0.69 |
| it (Italian) | 59 | 0.55 |
| zh (Chinese) | 57 | 0.53 |
| pt (Portuguese) | 51 | 0.47 |
| fi (Finnish) | 49 | 0.45 |
| in (Indonesian) | 48 | 0.44 |
Table 4.
Top-5 most frequent theoretical constructs detected in the Russia–Ukraine tweet set (January–June 2022).
| Theory/Construct | Count | Percentage (%) |
|---|
| Deindividuation | 1654 | 15.29 |
| Cognitive Distortion | 525 | 4.85 |
| Threat Appraisal | 503 | 4.65 |
| General Aggression/Deterrence | 193 | 1.78 |
| Rumor Transmission | 79 | 0.73 |
Table 5.
Computational accuracy metrics on the manually validated subset N = 250.
| Latent Construct | TP | FP | FN | Precision | Recall | F1-Score |
|---|
| Deindividuation | 38 | 2 | 3 | 95.00% | 92.68% | 93.83% |
| Cognitive Distortion | 82 | 4 | 7 | 95.35% | 92.13% | 93.71% |
| Threat Appraisal | 25 | 1 | 1 | 96.15% | 96.15% | 96.15% |
| Aggression/Deterrence | 21 | 0 | 2 | 100.00% | 91.30% | 95.45% |
| Rumor Transmission | 18 | 1 | 1 | 94.74% | 94.74% | 94.74% |
| Macro Average | – | – | – | 96.25% | 93.40% | 94.78% |
Table 6.
Ablation based robustness check for deindividuation detection.
| Model Variant | Feature Configuration | Detected Prevalence | Jaccard with Full Model | Cascade Model LL |
|---|
| Full model | Contextual, lexical, temporal, hashtag, engagement, and network features | 15.29% | 1.00 | −865.00 |
| Lexicon excluded model | Full model excluding deindividuation specific dictionary terms | 14.82% | 0.91 | −878.45 |
| Non-lexical model | Transformer embeddings, synchrony, hashtag uniformity, engagement, and network features only | 14.15% | 0.86 | −891.20 |
| Dictionary only model | Deindividuation dictionary indicators only | 8.42% | 0.44 | −1012.15 |
Table 7.
Co-occurrence matrix and Jaccard indices among theoretical constructs in the Russia–Ukraine tweet set (Jan–Jun 2022). Each cell reports “Count (Jaccard index)”.
| Theory | SIT/ RCT | MFT | Deindivi-Duation | Aggression | Threat | Distortion | TPB | Rumor |
|---|
| SIT/RCT (Hostility) | 2815 (1.00) | 335 (0.10) | 559 (0.14) | 90 (0.03) | 219 (0.07) | 203 (0.06) | 81 (0.03) | 24 (0.01) |
| MFT (Moral Foundations) | 335 (0.10) | 724 (1.00) | 260 (0.12) | 21 (0.02) | 79 (0.07) | 72 (0.06) | 50 (0.05) | 9 (0.01) |
| Deindividuation | 559 (0.14) | 260 (0.12) | 1654 (1.00) | 41 (0.02) | 143 (0.07) | 445 (0.26) | 208 (0.12) | 17 (0.01) |
| General Aggression/Det. | 90 (0.03) | 21 (0.02) | 41 (0.02) | 193 (1.00) | 21 (0.03) | 16 (0.02) | 1 (0.00) | 3 (0.01) |
| Threat Appraisal | 219 (0.07) | 79 (0.07) | 143 (0.07) | 21 (0.03) | 503 (1.00) | 44 (0.04) | 20 (0.03) | 4 (0.01) |
| Cognitive Distortion | 203 (0.06) | 72 (0.06) | 445 (0.26) | 16 (0.02) | 44 (0.04) | 525 (1.00) | 33 (0.04) | 3 (0.00) |
| TPB (Collective Action) | 81 (0.03) | 50 (0.05) | 208 (0.12) | 1 (0.00) | 20 (0.03) | 33 (0.04) | 314 (1.00) | 1 (0.00) |
| Rumor Transmission | 24 (0.01) | 9 (0.01) | 17 (0.01) | 3 (0.01) | 4 (0.01) | 3 (0.00) | 1 (0.00) | 79 (1.00) |
Table 8.
Internal consistency, calibration, and robustness diagnostics of theoretical construct detection.
| Metric | Value | Notes/Interpretation |
|---|
| Internal lexical consistency (Hostility) | r = 0.82 | Association between SIT/RCT detections and derogatory lexical indicators; interpreted as internal consistency rather than independent construct validation |
| Internal lexical consistency (Threat) | r = 0.76 | Association between Threat Appraisal detections and conditional or threat cue terms |
| Internal lexical consistency (Deindividuation) | r = 0.88 | Association between Deindividuation detections and collective pronoun or crowd-language indicators |
| Cross-construct consistency | 14.6% | Tweets often combine Distortion + Deindividuation, consistent with theory |
| Internal calibration (Rumor) | 93% | Majority of “BREAKING/rumor” tweets flagged correctly |
| Reliability under resampling | 91% | Construct detection robust under bootstrap sampling |
Table 9.
Incremental nested model comparison for retweet cascade explanation.
| Model | Predictors Included | LL | AIC | BIC | | p |
|---|
| M0 | Temporal baseline metrics | −1142.50 | 2289.00 | 2303.40 | – | – |
| M1 | Temporal and user engagement metrics | −1012.15 | 2036.30 | 2080.03 | 260.70 | <0.001 |
| M2 | Temporal, engagement, and network controls | −941.80 | 1901.60 | 1967.19 | 140.70 | <0.001 |
| M3 | M2 plus latent psychological constructs | −865.00 | 1758.00 | 1860.03 | 153.60 | <0.001 |
Table 10.
Comparison between the early-onset corpus and the external temporal validation corpus.
| Dimension | Primary Corpus | External Validation Corpus [22] | Interpretation |
|---|
| Temporal window | January–June 2022 | October 2022–April 2023 | The primary corpus captures the onset and early escalation phase, while the validation corpus captures a later and more routinized phase of conflict discourse. |
| Corpus size | 10,815 tweets | 37,386 tweets | The validation corpus substantially expands the empirical basis and addresses concerns about reliance on a single 10 K-tweet dataset. |
| User base | Not used as a primary unit of inference | 30,706 users | The external corpus provides broader user-level coverage, although the present analysis remains focused on discourse-level rather than individual-level inference. |
| Language diversity | English-dominant corpus; 61.45% English | 54 languages | The validation corpus provides wider multilingual coverage, while also reinforcing the need for cautious interpretation of cross-language construct equivalence. |
| Conflict phase | Initial escalation and full-scale invasion period | Later conflict normalization period | Psychological signals are expected to be more concentrated during crisis onset and more diffuse during later phases. |
| Expected construct pattern | Stronger activation of threat appraisal, deindividuation, hostility, and cognitive distortion | More dispersed and attenuated construct activation | The comparison supports the interpretation that acute geopolitical escalation intensifies collective psychological signals. |
| Analytical role | Main model development and primary empirical analysis | External temporal validation only | The validation corpus is used to test temporal robustness, not to retrain or redefine the model. |
Table 11.
Research gaps addressed by this study.
| Gap in Literature | Evidence in Prior Work | Contribution of this Study |
|---|
| Focus on sentiment/stance rather than psychological constructs | Studies emphasize sentiment and stance classification without grounding in theory [1,2,3] | Operationalizes theories such as SIT, MFT, and Threat Appraisal into probabilistic detection models |
| Difficulty in translating constructs like deindividuation or distortions into computational form | Prior work notes challenges in bridging social and computational sciences [12,13] | Introduces latent construct models with partial pooling and theory-specific logistic heads |
| Limited integration of diffusion/polarization studies with psychology | Empirical studies highlight polarization but neglect theory-driven explanations [14,15,16,17,18] | Links latent constructs to cascade dynamics and hashtag polarization |
| Lack of robustness and calibration in construct detection | Many models face issues of bias and generalization [19,20,21] | Achieves 93% internal calibration for rumor detection and 91% resampling reliability |
Table 12.
Comparison of baseline models with the proposed theory-driven framework. Metrics are reported on annotated subsets () and cascade prediction tasks. The downward arrow for Brier score indicates that lower values represent better calibration performance.
| Model | F1 (Post Hoc Construct-Label Alignment) | Calibration (Brier ↓) | Cascade Log-Likelihood |
|---|
| Sentiment (VADER) | 0.41 | 0.23 | −1243 |
| Sentiment (BERTweet) | 0.55 | 0.19 | −1084 |
| Topic Modeling (LDA) | 0.47 | 0.21 | −1167 |
| Stance Detection | 0.58 | 0.18 | −1026 |
| Proposed Framework | 0.72 | 0.12 | −865 |
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |