MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability
Abstract
1. Introduction
- Evaluation through a heterogeneous dataset composed of different languages, different recording conditions, etc.
- Calibration/reliability: Classification accuracy is not an indicator of trustworthiness of a system.
- Explainability: The lack of quantitatively verified modules for explaining which part of speech influences the prediction.
- Deployment: The ack of frameworks that are available on the web.
- Speech based depression-severity framework development and evaluation through a heterogeneous dataset.
- Domain handling through domain-specific feature normalization.
- Calibration for confidence estimation.
- Proposing the Pipeline-Aware Grad-CAM explanation module.
- Integration of all the above in a web-based app, “MindVoice”.
2. Related Work
2.1. Speech-Based Depression Detection
2.2. Multimodal Depression Detection
2.3. Multi-Dataset Depression Detection
2.4. Confidence Calibration
2.5. Explainability in Speech-Based Depression Detection
2.6. Mobile Mental Health Application
3. Proposed Framework
3.1. Multi-Dataset Preparation
3.2. CNN Feature Extraction
- and are fully connected layer weights;
- is ReLU activation function;
- is Sigmoid activation.
Enhancement of Feature Extractor
- and = spatial dimensions;
- = channels;
- = one element of features.
3.3. Domain Adaptation
- is domain;
- is mean of domain ;
- is standard deviation of domain .
3.4. Feature Selection
3.5. Multi-Layer Perceptron (MLP) Classification
- ;
- bias;
- = activation function.
3.6. Probability Calibration
- If predicted probability > actual probability, then the model is said to be over-confident.
- If predicted probability < actual probability, then the model is said to be under-confident.
3.6.1. Temperature Scaling
- is predicted logic;
- is number of classes;
- is logit for class .
- If , then original probabilities are not changed.
- If , then over-confident predictions are softened.
- If , then under-confident predictions are sharpened.
3.6.2. Platt Scaling
- is classifier output before calibration.
- and are calculated by minimizing the NLL over the calibration data.
3.6.3. Isotonic Regression
- (.) is constant monotonic function.
3.6.4. Vector Scaling
- is diagonal scaling matrix;
- is bias vector.
3.6.5. Matrix Scaling
3.6.6. Dirichlet Calibration
- is original probability;
- is transformation matrix;
- is bias vector.
3.6.7. Beta Calibration
3.7. Reliability Estimation
- = number of samples;
- = bin accuracy;
- = prior accuracy = 0.5;
- = prior weight = 10.
- = samples present in the bin;
- = prior weight = 10.
- Trained CNN backbone.
- Scaler.
- ANOVA indices.
- ABC indices.
- Raw model.
- Best calibration method.
- Reliability Lookup Table.
4. Web-Based Screening System for Depression and PTSD
- Confidence score that is basically the calibrated confidence reported after calibration.
- Reliability estimation of each prediction for increasing the trustworthiness of the application.
- Recommendation module that provides personalized guidance as per the severity category of depression.
- Explainability module that encourages decision transparency. The explainability module is built on the proposed Pipeline-Aware GradCAM (PA_GradCAM).
4.1. Confidence Score
4.2. Reliability Estimation
4.3. Recommendation Module
4.4. Explainability Module
4.4.1. PA-Grad-CAM Generation
| Algorithm 1. Pipeline-Aware Grad-CAM (PA-GradCAM) |
Phase 1: Forward Prediction Inference
|
- Standard Scaler Reconstruction:
- 2.
- ANOVA and ABC Reconstruction:
- 3.
- MLP Reconstruction:
- ;
- bias;
- = activation function.
- = probability of class ;
- = logit of class ;
- = logits of all classes.
4.4.2. Heatmap-to-Text Interpretation
- Calculate temporal and frequency attributions.
- Locate their maximum positions.
- Normalize peak positions to [0, 1].
- Calculate the heatmap concentration and active-region fraction.
- Categorize attribution as localized or distributed attributes.
- Map with pre-written text templates.
- is number of frequency;
- is number of time.
- Let be the heatmap at frequency f and time t.
- The time-wise importance profile is given by (30):
- 2.
- The frequency-wise importance profile is given by (31):
- 3.
- Temporal peak:
- 4.
- Frequency peak:
- 5.
- Concentration of the heatmap is given by (35),Large shows that attribution values vary more strongly across the spectrogram.Smaller shows that attribution values are homogeneous across the spectrogram.
- 6.
- For checking localization, the Active fraction is utilized as (37):
- 7.
- According to the predefined heuristic threshold,
- 8.
- Lastly, predefined text templets are mapped.
4.4.3. PA-Grad-CAM Validation Protocol
- Prediction Equivalence;
- Gradient Correctness;
- Perturbation-Based Faithfulness.
Prediction Equivalence
Gradient Correctness
Perturbation-Based Faithfulness
5. Results and Discussion
5.1. Performance of the Proposed Framework
5.1.1. Combined Dataset Analysis
5.1.2. Individual Dataset Performance
- is set of predictions of confidence bin;
- is number of samples in bin;
- is total test sample;
- = classification accuracy of samples;
- is average predicted confidence.
- () is predicted probability;
- is total test sample.
- is total number of samples;
- is number of classes;
- is predicted probability;
- is true label.
5.1.3. Leave-One-Dataset-Out (LODO) Evaluation
5.2. Comparison Among the Calibration Techniques
5.3. Validation of PA-Grad-CAM Explanations
5.3.1. Prediction Equivalence
5.3.2. Gradient Correctness
5.3.3. Perturbation-Based Faithfulness
5.4. Ablation Study and Baseline Comparison with the Proposed Framework
5.5. Reliability-Weight Ablation Study
- Correctness = 1 for correct prediction;
- Correctness = 0 for incorrect prediction.
- = reliability score;
- = correctness.
5.6. Signal Independence Analysis
- = bin accuracy;
- = entropy.
5.7. Web Application
- Level of depression classified as Low, Moderate and High Depression.
- Confidence score.
- Reliability score.
- Next screening day recommendation.
- Why this prediction? Using our proposed PA-GradCAM.
- Recommendations.
5.8. Comparison with Existing Work
6. Conclusions
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- World Health Organization. Depression. Available online: https://www.who.int/news-room/fact-sheets/detail/depression (accessed on 29 August 2026).
- Cummins, N.; Scherer, S.; Krajewski, J.; Schnieder, S.; Epps, J.; Quatieri, T.F. A Review of Depression and Suicide Risk Assessment Using Speech Analysis. Speech Commun. 2015, 71, 10–49. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.-Y.; Han, D.-K.; Park, S.-H.; Jang, G.-D.; Lee, S.-W. Improving Generalization of Drowsiness State Classification by Domain-Specific Normalization. In Proceedings of the International Winter Workshop on Brain-Computer Interface (BCI), Gangwon, Republic of Korea, 26–28 February 2024. [Google Scholar] [CrossRef] [Scilit]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 1321–1330. [Google Scholar]
- Darby, J.K.; Simmons, N.; Berger, P.A. Speech and Voice Parameters of Depression: A Pilot Study. J. Commun. Disord. 1984, 17, 75–85. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jiang, H.; Hu, B.; Liu, Z.; Wang, G.; Zhang, L.; Li, X.; Kang, H. Detecting Depression Using an Ensemble Logistic Regression Model Based on Multiple Speech Features. Comput. Math. Methods Med. 2018, 2018, 6508319. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ooi, K.E.B.; Lech, M.; Allen, N.B. Prediction of Major Depression in Adolescents Using an Optimized Multi-Channel Weighted Speech Classification System. Biomed. Signal Process. Control 2014, 14, 228–239. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Wang, F.; Gao, Y.; Liao, Y.; Zhang, W.; Zhang, L.; Xu, Z. Depression Recognition Using Voice-Based Pre-Training Model. Sci. Rep. 2024, 14, 12734. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kalatzis, I.; Piliouras, N.; Ventouras, E.; Papageorgiou, C.C.; Rabavilas, A.D.; Cavouras, D. Design and Implementation of an SVM-Based Computer Classification System for Discriminating Depressive Patients from Healthy Controls Using the P600 Component of ERP Signals. Comput. Methods Programs Biomed. 2004, 75, 11–22. [Google Scholar] [CrossRef] [PubMed]
- Kumar, P.; Garg, S.; Garg, A. Assessment of Anxiety, Depression and Stress Using Machine Learning Models. Procedia Comput. Sci. 2020, 171, 1989–1998. [Google Scholar] [CrossRef] [Scilit]
- Atila, O.; Şengür, A. Attention Guided 3D CNN-LSTM Model for Accurate Speech Based Emotion Recognition. Appl. Acoust. 2021, 182, 108260. [Google Scholar] [CrossRef] [Scilit]
- Rezaee, K. Depression Detection from Speech Data Using Deep Learning-Based Optimized Temporal-Frequency-Channel Attention with Interpretable Acoustic-Prosodic Mapping. J. Affect. Disord. 2026, 399, 121077. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, L.; Li, Z.; Tiwari, P.; Cao, C.; Xue, J.; Zhu, F.; Wu, D. Depressformer: Leveraging Video Swin Transformer and fine-grained local features for depression scale estimation. Biomed. Signal Process. Control 2024, 96, 900507. [Google Scholar] [CrossRef] [Scilit]
- Ye, J.; Yu, Y.; Lu, L.; Wang, H.; Zheng, Y.; Liu, Y. DEP-Former: Multimodal Depression Recognition Based on Facial Expressions and Audio Features via Emotional Changes. IEEE Trans. Circuits Syst. Video Technol. 2024, 35, 2087–2100. [Google Scholar] [CrossRef] [Scilit]
- Ding, H.; Du, Z.; Wang, Z.; Xue, J.; Wei, Z.; Yang, K.; Jin, S.; Zhang, Z.; Wang, J. IntervoxNet: A novel dual-modal audio-text fusion network for automatic and efficient depression detection from interviews. Front. Phys. 2024, 12, 1430035. [Google Scholar] [CrossRef] [Scilit]
- Ryumina, E.; Axyonov, A.; Dolgushin, M.; Ryumin, D.; Karpov, A. DEPART: Multi-Task Interpretable Depression and Parkinson ’s Disease Detection from In-the-Wild Video Data. Big Data Cogn. Comput. 2026, 10, 89. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Yang, X.; Zhao, M.; Qi, S. Predicting depression by using a novel deep learning model and video-audio-text multimodal data. Front. Psychiatry 2025, 16, 1602650. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, X.; He, R.; Wen, S.; Zhou, R.; Wang, J. BCMA-MBF: Research on Depression Prediction Based on Bidirectional Cross-Modal Attention with Multi-Task Linear Fusion. J. Affect. Disord. 2026, 393, 120385. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gupta, A.K.; Dhamaniya, A.; Gupta, P. RADIANCE: Reliable and Interpretable Depression Detection from Speech Using Transformer. Comput. Biol. Med. 2024, 183, 109325. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sun, B.; Feng, J.; Saenko, K. Return of Frustratingly Easy Domain Adaptation. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016; Volume 30, pp. 2058–2065. [Google Scholar] [CrossRef] [Scilit]
- Sun, B.; Saenko, K. Deep CORAL: Correlation Alignment for Deep Domain Adaptation. In Computer Vision—ECCV 2016 Workshops, Amsterdam, The Netherlands, 8–10 October 2016; Springer: Cham, Switzerland, 2016; pp. 443–450. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wang, N.; Shi, J.; Hou, X.; Liu, J. Adaptive Batch Normalization for Practical Domain Adaptation. Pattern Recognit. 2018, 80, 109–117. [Google Scholar] [CrossRef] [Scilit]
- Chang, W.-G.; You, T.; Seo, S.; Kwak, S.; Han, B. Domain-Specific Batch Normalization for Unsupervised Domain Adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 7354–7362. [Google Scholar] [CrossRef] [Scilit]
- Platt, J.C. Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods. In Advances in Large Margin Classifiers; Smola, A.J., Bartlett, P.L., Schölkopf, B., Schuurmans, D., Eds.; MIT Press: Cambridge, MA, USA, 1999; pp. 61–74. [Google Scholar]
- Zadrozny, B.; Elkan, C. Transforming Classifier Scores into Accurate Multiclass Probability Estimates. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Edmonton, AB, Canada, 23–26 July 2002; pp. 694–699. [Google Scholar] [CrossRef] [Scilit]
- Kull, M.; Silva Filho, T.; Flach, P. Beta Calibration: A Well-Founded and Easily Implemented Improvement on Logistic Calibration for Binary Classifiers. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA, 20–22 April 2017; Volume 54, pp. 623–631. [Google Scholar]
- Kull, M.; Perello-Nieto, M.; Kängsepp, M.; Silva Filho, T.; Song, H.; Flach, P. Beyond Temperature Scaling: Obtaining Well-Calibrated Multiclass Probabilities with Dirichlet Calibration. In Proceedings of the Advances in Neural Information Processing Systems 32, Vancouver, BC, Canada, 8–14 December 2019; pp. 12316–12326. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]
- Chattopadhyay, A.; Sarkar, A.; Howlader, P.; Balasubramanian, V.N. Grad-CAM++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Lake Tahoe, NV, USA, 12–15 March 2018; pp. 839–847. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Wang, Z.; Du, M.; Yang, F.; Zhang, Z.; Ding, S.; Mardziel, P.; Hu, X. Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 14–19 June 2020; pp. 24–25. [Google Scholar] [CrossRef] [Scilit]
- Desai, S.; Ramaswamy, H.G. Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-Free Localization. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Snowmass, CO, USA, 1–5 March 2020; pp. 972–980. [Google Scholar] [CrossRef] [Scilit]
- Muhammad, M.B.; Yeasin, M. Eigen-CAM: Class Activation Map Using Principal Components. In Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN), Glasgow, UK, 19–24 July 2020; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Fu, R.; Hu, Q.; Dong, X.; Guo, Y.; Gao, Y.; Li, B. Axiom-Based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs. In Proceedings of the 31st British Machine Vision Conference (BMVC), Manchester, UK, 7–10 September 2020. Paper 631. [Google Scholar]
- Jiang, P.-T.; Zhang, C.-B.; Hou, Q.; Cheng, M.-M.; Wei, Y. LayerCAM: Exploring Hierarchical Class Activation Maps for Localization. IEEE Trans. Image Process. 2021, 30, 5875–5888. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Draelos, R.L.; Carin, L. Use HiResCAM Instead of Grad-CAM for Faithful Explanations of Convolutional Neural Networks. arXiv 2020, arXiv:2011.08891. [Google Scholar] [CrossRef] [Scilit]
- Torous, J.; Staples, P.; Shanahan, M.; Lin, C.; Peck, P.; Keshavan, M.; Onnela, J.-P. Utilizing a Personal Smartphone Custom App to Assess the Patient Health Questionnaire-9 (PHQ-9) Depressive Symptoms in Patients with Major Depressive Disorder. JMIR Ment. Health 2015, 2, e8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- BinDhim, N.F.; Shaman, A.M.; Trevena, L.; Basyouni, M.H.; Pont, L.G.; Alhawassi, T.M. Depression Screening via a Smartphone App: Cross-Country User Characteristics and Feasibility. J. Am. Med. Inform. Assoc. 2015, 22, 29–34. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jacobson, N.C.; Chung, Y.J. Passive Sensing of Prediction of Moment-to-Moment Depressed Mood among Undergraduates with Clinical Levels of Depression Sample Using Smartphones. Sensors 2020, 20, 3572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, X.; Triantafyllopoulos, A.; Kathan, A.; Milling, M.; Yan, T.; Rajamani, S.T.; Kuster, L.; Harrer, M.; Heber, E.; Grossmann, I.; et al. Depression Diagnosis and Forecast Based on Mobile Phone Sensor Data. In Proceedings of the 44th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Glasgow, UK, 11–15 July 2022; pp. 4679–4682. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Stamatis, C.A.; Meyerhoff, J.; Meng, Y.; Lin, Z.C.C.; Cho, Y.M.; Liu, T.; Karr, C.J.; Curtis, B.L.; Ungar, L.H.; Mohr, D.C. Differential Temporal Utility of Passively Sensed Smartphone Features for Depression and Anxiety Symptom Prediction: A Longitudinal Cohort Study. npj Ment. Health Res. 2024, 3, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gratch, J.; Artstein, R.; Lucas, G.; Stratou, G.; Scherer, S.; Nazarian, A.; Wood, R.; Boberg, J.; DeVault, D.; Marsella, S.; et al. The Distress Analysis Interview Corpus of Human and Computer Interviews. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), Reykjavik, Iceland, 26–31 May 2014; pp. 3123–3128. [Google Scholar]
- Cai, H.; Yuan, Z.; Gao, Y.; Sun, S.; Li, N.; Tian, F.; Xiao, H.; Li, J.; Yang, Z.; Li, X.; et al. A Multi-Modal Open Dataset for Mental-Disorder Analysis. Sci. Data 2022, 9, 178. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, Y.; Yang, H.; Lin, L. Automatic Depression Detection: An Emotional Audio-Textual Corpus and a GRU/BiLSTM-Based Model. In Proceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 6247–6251. [Google Scholar] [CrossRef] [Scilit]
- Choubey, A.; Pandey, M.K.; Dubey, A.K.; Rocha, A. Speech-Based Depression Severity Estimation Using a Multi-Scale CNN with Channel Attention Mechanism. Speech Commun. 2026; submitted.
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017. [Google Scholar]
- Ben-Nun, T.; Besta, M.; Huber, S.; Ziogas, A.N.; Peter, D.; Hoefler, T. A Modular Benchmarking Infrastructure for High-Performance and Reproducible Deep Learning. In Proceedings of the2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS), Rio de Janeiro, Brazil, 20–24 May 2019. [Google Scholar] [CrossRef] [Scilit]
- Petsiuk, V.; Das, A.; Saenko, K. RISE: Randomized Input Sampling for Explanation of Black-box Models. In Proceedings of the British Machine Vision Conference (BMVC), Newcastle, UK, 3–4 September 2018; BMVA Press: Durham, UK, 2018; p. 151. [Google Scholar]
- Mdhaffar, A.; Cherif, F.; Kessentini, Y.; Maalej, M.; Ben Thabet, J.; Maalej, M.; Jmaiel, M.; Freisleben, B. DL4DED: Deep Learning for Depressive Episode Detection on Mobile Devices. In How AI Impacts Urban Living and Public Health; Pagán, J., Mokhtari, M., Aloulou, H., Abdulrazak, B., Cabrera, M., Eds.; Springer: Cham, Switzerland, 2019; Volume 11862, pp. 109–121. [Google Scholar] [CrossRef] [Scilit]
- Tonn, P.; Seule, L.; Degani, Y.; Herzinger, S.; Klein, A.; Schulze, N. Digital Content Free Speech Analysis Tool to Measure Affective Distress in Mental Health: Evaluation Study. JMIR Form. Res. 2022, 6, e37061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, L.; Tydeman, F.; Xie, W.; Wang, Y. Development of an AI-Based Mobile App for Automatic Depression Screening Using Speech in English and Chinese. In Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 14–18 July 2025; pp. 1–5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, A.Y.; Jeon, M.; Cho, C.-H.; Shin, M.-S.; Byun, S. Screening for Depression Risk via Smartphone Narratives with Fully Fine-Tuned WavLM. Sci. Rep. 2026, 16, 22741. [Google Scholar] [CrossRef] [Scilit] [PubMed]







| Year | Study | Modality | Task | Method | Advantage | Limitation |
|---|---|---|---|---|---|---|
| 2018 | Jiang et al. [6] | Audio | Depression detection | Ensemble logistic regression using multiple speech features | Combines various speech attributes with an interpretable ML model | Binary classification only, dependent on handcrafted features |
| 2020 | Kumar et al. [10] | Questionnaire responses | Depression severity classification | Eight ML algorithms and hybrid models | Multiple severity levels | Prone to user bias |
| 2024 | Huang et al. [8] | Audio | Depression severity classification | wav2vec 2.0 and fine-tuning network | Extract features directly from raw speech | Evaluated on DAIC-WOZ only; cross-dataset validation insufficient |
| 2024 | Depressformer [13] | Video | Depression severity classification | Video Swin Transformer, fine-grained local features, channel-attention fusion | Captures fine-grained features of video | Greater computational burden than speech-only systems |
| 2024 | IntervoxNet [15] | Audio + text | Depression detection | Mel-Spectrogram Transformer for audio; BERT/CNN for text and attention fusion | Combines acoustic and linguistic information | Binary classification only, requires two synchronized modalities, single dataset evaluation |
| 2025 | DEP-Former [14] | Audio + facial expressions | Depression detection | Modality adapter, shared attention index, multimodal cross-attention | Uses facial and acoustic information | Binary classification only, requires two synchronized modalities |
| 2025 | IMDD-Net [17] | Audio + video + text | Depression severity estimation | TimeSformer for video, MFCC and eGeMAPS for audio and BERT for text | Integrates local and global information across three modalities | High computational complexity due to three modalities, single dataset evaluation |
| 2026 | Rezaee [12] | Audio | Depression detection | Optimized temporal–frequency–channel attention, interpretable acoustic–prosodic mapping | Combines temporal–frequency–channel attention and interpretable acoustic–prosodic mapping | Binary classification only |
| 2026 | DEPART [16] | Video | Depression detection | CLIP-based visual representation, temporal Transformer, prototype-aware modeling | Integrate interpretability into video-based depression detection | High computational load, binary classification only |
| Harmonized Classes | PHQ-8/PHQ-9 Range | Original Categories Combined |
|---|---|---|
| Low | 0–9 | None/Minimal and Mild |
| Moderate | 10–19 | Moderate and Moderately Severe |
| High | ≥20 | Severe |
| SDS Raw Score | Severity Category |
|---|---|
| 20–39 | Normal/No depression |
| 40–47 | Mild depression |
| 48–55 | Moderate depression |
| 56–80 | Severe depression |
| Class | Precision | Recall | F1 Score |
|---|---|---|---|
| Low | 0.695 | 0.680 | 0.688 |
| Moderate | 0.772 | 0.723 | 0.747 |
| High | 0.823 | 0.893 | 0.857 |
| Dataset | N | Accuracy | AUC | ECE | Brier Score | NLL |
|---|---|---|---|---|---|---|
| DAIC | 54 | 0.814 | 0.948 | 0.159 | 0.317 | 0.971 |
| E-DAIC | 72 | 0.722 | 0.893 | 0.229 | 0.474 | 1.348 |
| MODMA | 15 | 0.80 | 0.84 | 0.178 | 0.370 | 1.911 |
| Combined | 141 | 0.765 | 0.906 | 0.190 | 0.403 | 1.263 |
| Dataset | N | Accuracy | AUC | ECE | Brier Score | NLL |
|---|---|---|---|---|---|---|
| EATD | 39 | 0.785 | 0.972 | 0.159 | 0.116 | 0.518 |
| Held Out Dataset | Training Datasets | Accuracy | AUC | ECE | Brier Score | NLL |
|---|---|---|---|---|---|---|
| DAIC | E-DAIC + MODMA | 0.777 | 0.906 | 0.147 | 0.330 | 0.914 |
| E-DAIC | DAIC + MODMA | 0.513 | 0.700 | 0.351 | 0.787 | 2.51 |
| MODMA | DAIC + E-DAIC | 0.4 | 0.706 | 0.527 | 1.108 | 5.77 |
| Method | Accuracy | ECE | Brier Score | NLL |
|---|---|---|---|---|
| Platt Sigmoid | 0.750 | 0.030 | 0.409 | 0.702 |
| Dirichlet Calibration | 0.715 | 0.059 | 0.400 | 0.669 |
| Isotonic | 0.770 | 0.061 | 0.362 | 0.608 |
| Matrix Scaling | 0.715 | 0.061 | 0.400 | 0.669 |
| Vector Scaling | 0.680 | 0.091 | 0.406 | 0.688 |
| Beta Calibration | 0.722 | 0.094 | 0.414 | 0.704 |
| Temperature Scaling | 0.743 | 0.104 | 0.417 | 0.744 |
| Uncalibrated | 0.743 | 0.223 | 0.478 | 1.794 |
| Model | Dataset(s) | Final Probability Max abs. diff. | Prediction Reproducibility | Result |
|---|---|---|---|---|
| Combined 3-class PHQ model | DAIC + EDAIC + MODMA | 3.46 × 10−6 | 100% | PASS |
| EATD 4-class SDS model | EATD | 1.50 × 10−6 | 100% | PASS |
| Overall | All datasets | 3.46 × 10−6 | 100% | PASS |
| Metric | Result |
|---|---|
| Test samples | 20 |
| Feature-map locations per sample | 20 |
| Total gradient locations evaluated | 400 |
| Gradient target | Original predicted-class logit |
| Analytical gradient | Backpropagation gradient |
| Numerical gradient | Central finite-difference approximation |
| Step size, ε | 0.01 |
| Mean Pearson correlation | 0.999989 |
| Median Pearson correlation | 1.000000 |
| Minimum Pearson correlation | 0.999812 |
| Mean absolute error | 4.34 × 10−5 |
| Median sample-level relative error | 1.59 × 10−4 |
| Verification result | PASS |
| Dataset | Deletion AUC Grad-CAM ↓ | Deletion AUC Random ↓ | Deletion Advantage ↑ | p-Value | Insertion AUC Grad-CAM ↑ | Insertion Advantage ↑ | p-Value |
|---|---|---|---|---|---|---|---|
| DAIC | 0.4666 | 0.6116 | 0.1450 | <0.001 | 0.7437 | 0.0838 | <0.001 |
| EDAIC | 0.4056 | 0.5749 | 0.1693 | <0.001 | 0.6694 | 0.0704 | 0.0014 |
| MODMA | 0.3117 | 0.5734 | 0.2617 | 0.0038 | 0.7004 | 0.0936 | 0.0945 |
| EATD | 0.3641 | 0.4782 | 0.1141 | <0.001 | 0.6727 | 0.1468 | <0.001 |
| All | 0.4099 | 0.5705 | 0.1606 | <0.001 | 0.6964 | 0.0894 | <0.001 |
| Configuration | Features Generated | Features After ANOVA | Features After ABC | Accuracy | ECE | Brier Score | NLL |
|---|---|---|---|---|---|---|---|
| Untrained CNN + Flatten | 12,288 | 5335 | 2671 | 0.794 | 0.091 | 0.309 | 0.722 |
| Untrained CNN + GAP | 192 | 145 | 67 | 0.737 | 0.136 | 0.382 | 0.797 |
| Trained CNN + Flatten | 12,288 | 7222 | 3600 | 0.765 | 0.184 | 0.410 | 1.263 |
| Trained CNN + GAP (Main Pipeline) | 192 | 191 | 99 | 0.765 | 0.190 | 0.403 | 1.263 |
| Formulation | Alpha Bin Accuracy | Beta Entropy | Pearson Correlation with Correctness | AUC | Brier Vs Correctness |
|---|---|---|---|---|---|
| 0.7 bin/0.3 entropy | 0.7 | 0.3 | 0.450 | 0.804 | 0.145 |
| 0.6 bin/0.4 entropy | 0.6 | 0.4 | 0.447 | 0.800 | 0.146 |
| 0.5 bin/0.5 entropy | 0.5 | 0.5 | 0.443 | 0.798 | 0.148 |
| Adaptive Reliability Score (ARS) | NaN | NaN | 0.408 | 0.797 | 0.149 |
| Entropy only | NaN | NaN | 0.414 | 0.656751 | 0.1659 |
| Bin accuracy only | NaN | NaN | 0.445 | 0.674 | 0.147 |
| Metrics | Results |
|---|---|
| Pearson r (bin accuracy, entropy) | 0.841 |
| AUC—bin-accuracy only | 0.674 |
| AUC—entropy only | 0.796 |
| AUC—both combined | 0.797 |
| Logistic coefficient—bin accuracy | 2.315 |
| Logistic coefficient—entropy | 2.231 |
| System | Task | Datasets Used | Accuracy | F1 | AUC |
|---|---|---|---|---|---|
| DL4DED [48] | Binary Classification | DAIC-WOZ | 0.50 | - | - |
| VoiceSense [49] | Binary Classification | DAIC-WOZ + Androids Corpus | 0.68 | - | - |
| MoodEcho [50] | Binary Classification | DAIC-WOZ + Chinese clinical-interview dataset | - | English—
0.86 Chinese—0.75 | - |
| Smartphone narratives + WavLM [51] | Binary Classification | DAIC-WOZ + MODMA | 5-fold cross val-
0.68 OOF- 0.79 | - | 5-fold cross val-0.90
OOF-0.86 |
| MindVoice (Proposed) | Multi-class Classification (3 Classes for DAIC-WOZ, E-DAIC, MODMA and 4 Classes for EATD) | DAIC-WOZ, E-DAIC, MODMA, EATD | DAIC-0.814 E-DAIC-0.722 MODMA-0.8 EATD-0.785 Combined-0.765 | - | DAIC-0.948 E-DAIC-0.893 MODMA-0.84 EATD-0.972 Combined-0.862 |
| System | Confidence | Reliability | Explanation | Generalization |
|---|---|---|---|---|
| DL4DED | No | No | No | No |
| VoiceSense | Yes | No | Partial (Vocal characteristics/parameters) | Yes (2 datasets) |
| MoodEcho | Yes | No | No | Yes (2 datasets) |
| Smartphone narratives + WavLM | Yes | No | No | Yes (2 datasets) |
| MindVoice (Proposed) | Yes | Yes | Yes (PA-Grad-CAM) | Yes (4 datasets and combined data evaluation) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Choubey, A.; Pandey, M.K.; Dubey, A.K.; Rocha, A. MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability. Computers 2026, 15, 646. https://doi.org/10.3390/computers15100646
Choubey A, Pandey MK, Dubey AK, Rocha A. MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability. Computers. 2026; 15(10):646. https://doi.org/10.3390/computers15100646
Chicago/Turabian StyleChoubey, Arjita, Manoj Kumar Pandey, Ashwani Kumar Dubey, and Alvaro Rocha. 2026. "MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability" Computers 15, no. 10: 646. https://doi.org/10.3390/computers15100646
APA StyleChoubey, A., Pandey, M. K., Dubey, A. K., & Rocha, A. (2026). MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability. Computers, 15(10), 646. https://doi.org/10.3390/computers15100646

