DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text
Abstract
1. Introduction
- 1.
- A group-sparse attention regularizer ( elastic net) on the query and key projection matrices that pushes whole input dimensions toward zero, yielding more concentrated and inspectable attention patterns—a structural property we distinguish from full faithfulness, which would require additional perturbation-based evaluation (Section 4.3).
- 2.
- A supervised contrastive loss [9] on the [CLS] projection that organizes the latent space by class, improving generalization and visual separability.
- 3.
- A lightweight design (six layers, eight heads, ∼9.5 M parameters) that we train and evaluate end-to-end on a single NVIDIA GTX 1660 (6 GB).
- 4.
- An empirical study on Dreaddit [10] with five seeded runs, bootstrap confidence intervals, McNemar significance tests (reported both per seed and pooled), and ablations isolating each regularizer.
2. Related Work
2.1. Mental Health Detection from Social Media
2.2. Sparse and Efficient Attention
2.3. Contrastive Learning for Text
3. Materials and Methods
3.1. Overall Architecture
3.2. Input Representation
3.3. Transformer Encoder with Sparse Attention
3.4. Supervised Contrastive Learning
3.5. Total Objective
3.6. Training Procedure
- Inputs: dataset ; hyperparameters ; architecture parameters ; number of epochs E; learning rate .
- Output: trained model parameters .
- 1.
- Initialize the encoder, classifier head, projection head, and AdamW optimizer.
- 2.
- For each epoch :
- (a)
- Shuffle .
- (b)
- For each mini-batch :
- (i)
- Tokenize and embed sequences to obtain .
- (ii)
- Compute encoder output: .
- (iii)
- Extract [CLS] embeddings: .
- (iv)
- Compute contrastive projections: .
- (v)
- Compute class logits: .
- (vi)
- Compute , , and per Equations (5)–(7).
- (vii)
- Form total loss: .
- (viii)
- Update parameters: .
- (c)
- Evaluate on validation set; save checkpoint if improves.
- 3.
- Return the checkpoint with the highest validation .
3.7. Dataset
3.8. Experimental Setup
3.8.1. Implementation and Hardware
Hardware
3.8.2. Baselines
- 1.
- Logistic Regression (BoW): TF-IDF weighted unigrams and bigrams, .
- 2.
- BiLSTM: 2 layers, 128 hidden units per direction, 200-dim learned embeddings.
- 3.
- Standard Transformer: identical architecture to DR-Transformer but with (no regularization beyond cross-entropy).
- 4.
- MentalBERT [7]: mental/mental-bert-base-uncased fine-tuned for 4 epochs at .
- 5.
- DR-Transformer (Sparse Only): ablation with .
- 6.
- DR-Transformer (SupCon Only): ablation with .
3.8.3. Evaluation Metrics and Statistical Testing
- 1.
- Per-seed McNemar tests: for each of the five seeds, a McNemar test is computed independently on that seed’s test set predictions. The reported per seed p-value is the median across seeds after Bonferroni correction for 6 pairwise comparisons. This approach respects the independence assumption of the McNemar test.
- 2.
- Pooled McNemar tests: predictions from all five seeds are pooled, effectively multiplying the sample size by five. This pooled test is reported as a secondary reference; because the same 715 test instances are reused across seeds, pooled predictions are not fully independent, and the pooled p-values should be interpreted with this caveat in mind.
- In practice, per seed and pooled results agree in direction and significance for all comparisons reported in Table 2. The bootstrap 95% CIs are similarly computed from pooled predictions; they reflect stability across seeds rather than strictly the most frequent coverage and should be interpreted accordingly.
4. Results
4.1. Classification Performance
4.2. Training Dynamics
4.3. Interpretability
Token Deletion Analysis
4.4. Computational Efficiency
5. Discussion
5.1. Principal Findings
5.2. Comparison with Existing Approaches
5.3. Ethical Considerations
- False positives and stigmatization: Our model is far from perfect, and predictions must never trigger interventions unilaterally; a human-in-the-loop review is essential.
- Privacy: Passive monitoring raises consent issues; any deployment should align with platform terms of service and applicable data-protection law (e.g., GDPR).
- Bias: Dreaddit annotations are predominantly in the English language and skew demographically; fairness audits across language, age, and gender are required before generalization.
- Crisis routing: Stress is not equivalent to a mental health crisis. Predictions should not be used in place of validated clinical screening.
5.4. Limitations
- 1.
- Dataset: Dreaddit is a relatively small dataset (3553 labeled segments) and English-language-only. Validation on multilingual or cross-platform data is needed before broad claims about generalizability can be made.
- 2.
- Binary task: We focus on binary stress; multi-class or continuous severity prediction would be a natural extension.
- 3.
- No temporal modeling: The model treats each segment independently; many practical applications involve longitudinal user data, which is outside the scope of this work and would require a different dataset (e.g., CLPsych under DUA).
- 4.
- Limited interpretability evaluation: We report quantitative sparsity, attention entropy, token deletion confidence drop, and qualitative attention maps covering four prediction types, but do not conduct a clinical-expert evaluation or a full faithfulness evaluation (comprehensiveness/sufficiency tests); both are left for future work.
- 5.
- Baseline scope: The current comparison does not include compact pretrained encoders such as DistilBERT or RoBERTa-base. As discussed in Section 3.8.2, including a fine-tuned pretrained encoder would be methodologically asymmetric given that DR-Transformer is trained from scratch. A controlled comparison with matching initialization conditions is left for future work. Consequently, the present experiments do not determine whether DR-Transformer is preferable to compact pretrained encoders such as DistilBERT or RoBERTa-base under realistic fine-tuning conditions; this question is left to the controlled comparison proposed above and in future work (Section 6).
5.5. Implications
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AUC | Area Under the Curve |
| BiLSTM | Bidirectional Long Short-Term Memory |
| CE | Cross-Entropy |
| CI | Confidence Interval |
| CLS | Classification token |
| DR-Transformer | Dual-Regularized Transformer |
| DSM-5 | Diagnostic and Statistical Manual of Mental Disorders, 5th ed. |
| GDPR | General Data Protection Regulation |
| IRB | Institutional Review Board |
| LLM | Large Language Model |
| NLP | Natural Language Processing |
| NT-Xent | Normalized Temperature-scaled Cross-Entropy |
| ROC | Receiver Operating Characteristic |
| SupCon | Supervised Contrastive (loss) |
| TF-IDF | Term Frequency–Inverse Document Frequency |
References
- Naslund, J.A.; Aschbrenner, K.A.; Marsch, L.A.; Bartels, S.J. The Future of Mental Health Care: Peer-to-Peer Support and Social Media. Epidemiol. Psychiatr. Sci. 2016, 25, 113–122. [Google Scholar] [CrossRef] [PubMed]
- Coppersmith, G.; Dredze, M.; Harman, C. Quantifying Mental Health Signals in Twitter. In Proceedings of the Workshop on Computational Linguistics and Clinical Psychology; Association for Computational Linguistics: Stroudsburg, PA, USA, 2014. [Google Scholar]
- De Choudhury, M.; Gamon, M.; Counts, S.; Horvitz, E. Predicting Depression via Social Media. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Washington, DC, USA, 2013. [Google Scholar]
- Guntuku, S.C.; Yaden, D.B.; Kern, M.L.; Ungar, L.H.; Eichstaedt, J.C. Detecting Depression and Mental Illness on Social Media: An Integrative Review. Curr. Opin. Behav. Sci. 2017, 18, 43–49. [Google Scholar] [CrossRef]
- Caruana, R.; Lou, Y.; Gehrke, J.; Koch, P.; Sturm, M.; Elhadad, N. Intelligible Models for Healthcare. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2015; pp. 1721–1730. [Google Scholar]
- Doshi-Velez, F.; Kim, B. Towards a Rigorous Science of Interpretable Machine Learning. arXiv 2017, arXiv:1702.08608. [Google Scholar]
- Ji, S.; Zhang, T.; Ansari, L.; Fu, J.; Tiwari, P.; Cambria, E. MentalBERT: Publicly Available Pretrained Language Models for Mental Health. arXiv 2022, arXiv:2203.06785. [Google Scholar]
- Yang, K.; Zhang, T.; Kuang, Z.; Xie, Q.; Ananiadou, S. MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models. In Proceedings of the ACM Web Conference 2024; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised Contrastive Learning. In Proceedings of the 34th International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020. [Google Scholar]
- Turcan, E.; McKeown, K. Dreaddit: A Reddit Dataset for Stress Analysis in Social Media. In Proceedings of the Tenth International Workshop on Health Text Mining and Information Analysis (LOUHI 2019); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 97–107. [Google Scholar]
- Coppersmith, G.; Dredze, M.; Harman, C.; Hollingshead, K. From ADHD to SAD: Analyzing the Language of Mental Health on Twitter through Self-Reported Diagnoses. In Proceedings of the 2nd Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality; Association for Computational Linguistics: Stroudsburg, PA, USA, 2015; pp. 1–10. [Google Scholar]
- Zirikly, A.; Resnik, P.; Uzuner, O.; Hollingshead, K. CLPsych 2019 Shared Task: Predicting the Degree of Suicide Risk in Reddit Posts. In Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 24–35. [Google Scholar]
- Sawhney, R.; Joshi, H.; Gandhi, S.; Shah, R. A Time-Aware Transformer Based Model for Suicide Ideation Detection on Social Media. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021. [Google Scholar]
- Matero, M.; Idnani, A.; Son, Y.; Giorgi, S.; Vu, H.; Zamani, M.; Limbachiya, P.; Guntuku, S.C.; Schwartz, H.A. Suicide Risk Assessment with Multi-level Dual-Context Language and BERT. In Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019. [Google Scholar]
- Cao, L.; Zhang, H.; Feng, L.; Wei, Z.; Wang, X.; Li, N.; He, X. Latent Suicide Risk Detection on Microblog via Suicide-Oriented Word Embeddings and Layered Attention. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019. [Google Scholar]
- Kyrou, M.; Kompatsiaris, I.; Petrantonakis, P. Deep Learning Approaches for Stress Detection: A Survey. IEEE Trans. Affect. 2025, 16, 499–517. [Google Scholar] [CrossRef]
- Soufleri, E.; Ananiadou, S. Enhancing Stress Detection on Social Media Through Multi-Modal Fusion of Text and Synthesized Visuals. In Proceedings of the 24th Workshop on Biomedical Language Processing (BioNLP 2025); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 34–43. [Google Scholar]
- Zhuang, M.; Cheng, D.; Lu, X.; Tan, X. Postgraduate Psychological Stress Detection from Social Media Using BERT-Fused Model. PLoS ONE 2024, 19, e0312264. [Google Scholar] [CrossRef] [PubMed]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The Long-Document Transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar]
- Zaheer, M.; Guruganesh, G.; Dubey, K.A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big Bird: Transformers for Longer Sequences. In Proceedings of the 34th International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020. [Google Scholar]
- Kitaev, N.; Kaiser, L.; Levskaya, A. Reformer: The Efficient Transformer. In Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, Online, 13–18 July 2020; pp. 1597–1607. [Google Scholar]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020. [Google Scholar]
- Gao, T.; Yao, X.; Chen, D. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021. [Google Scholar]
- Gunel, B.; Du, J.; Conneau, A.; Stoyanov, V. Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning. In Proceedings of the 9th International Conference on Learning Representations, Virtual Event, 3–7 May 2021. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the 31st International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019. [Google Scholar]
- Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 38–45. [Google Scholar]






| Parameter | Symbol | Value |
|---|---|---|
| Transformer layers | L | 6 |
| Attention heads | H | 8 |
| Embedding dimension | 256 | |
| Feed-forward dimension | 1024 | |
| Maximum sequence length | N | 128 |
| Sparsity coefficient () | ||
| Sparsity coefficient () | ||
| Sparse-loss weight | 0.10 | |
| Contrastive-loss weight | 0.05 | |
| Contrastive temperature | 0.07 | |
| Batch size | B | 16 |
| Peak learning rate | ||
| Warm-up steps | — | 500 |
| Total epochs | E | 20 |
| Run seeds | — | 42, 123, 456, 789, 101,112 |
| Split seed | — | 7 |
| Model | (95% CI) | Acc. | Prec. | Rec. | AUC | p |
|---|---|---|---|---|---|---|
| Logistic Reg. (BoW) | (–) | <0.001 | ||||
| BiLSTM | (–) | |||||
| Standard Transformer | (–) | ref. | ||||
| MentalBERT | (–) | |||||
| DR-Transformer (Sparse Only) | (–) | |||||
| DR-Transformer (SupCon Only) | (–) | |||||
| DR-Transformer (Full) | (–) | <0.001 |
| Variant | Attn. Sparsity | Silhouette | Attn. Entropy (Nats) |
|---|---|---|---|
| Standard Transformer (Vanilla) | |||
| DR-Transformer (Sparse Only) | |||
| DR-Transformer (SupCon Only) | |||
| DR-Transformer (Full) |
| Model | Train (min) | Peak GPU (GB) | Inference (ms/Sample) | # Params (M) |
|---|---|---|---|---|
| Logistic Regression | — | |||
| BiLSTM | ||||
| Standard Transformer | ||||
| MentalBERT † | ||||
| DR-Transformer |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chrifi Alaoui, M.; Joudar, N.-E.; Ettaouil, M. DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text. AI 2026, 7, 300. https://doi.org/10.3390/ai7080300
Chrifi Alaoui M, Joudar N-E, Ettaouil M. DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text. AI. 2026; 7(8):300. https://doi.org/10.3390/ai7080300
Chicago/Turabian StyleChrifi Alaoui, Mehdi, Nour-Eddine Joudar, and Mohamed Ettaouil. 2026. "DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text" AI 7, no. 8: 300. https://doi.org/10.3390/ai7080300
APA StyleChrifi Alaoui, M., Joudar, N.-E., & Ettaouil, M. (2026). DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text. AI, 7(8), 300. https://doi.org/10.3390/ai7080300

