Divergence Shepherd Feature Optimization-Based Stochastic-Tuned Deep Multilayer Perceptron for Emotional Footprint Identification
Abstract
1. Introduction
1.1. Motivation
1.2. Major Contributions
- ➢
- To enhance the accuracy of emotion footprint identification, a novel DSFO-STDMP model has been developed by including two processes: feature selection and classification.
- ➢
- To minimize the emotion footprint identification time, the DSFO-STDMP model executes a feature selection process by applying a shuffling shepherd optimization. To enhance the performance of optimization, the Sokal–Sneath similarity index and Jensen Jensen–Shannon divergence function are applied. This helps to minimize the dimensionality by reducing the redundant features.
- ➢
- To improve the precision and recall, a Rosenthal correlative stochastic-tuned deep multilayer perceptron classifier is used to analyze the correlation between data samples. Based on this analysis, different emotional footprints are correctly classified. In the fine-tuning phase, the stochastic gradient method is employed to minimize error.
- ➢
- Extensive experimentation is carried out to estimate the performance of the PRGKDFC model and other related works with different metrics.
1.3. Organization of Paper
2. Related Works
3. Proposal Methodology
3.1. Data Acquisition
3.2. Sokal–Sneath Divergence Shuffling Shepherd Optimization-Based Feature Reduction
| Algorithm 1: Sokal Sneath Divergence Shuffling Shepherd Optimization-based feature selection |
| Input: Number of features , data samples ‘ |
| Output: select the optimal features |
| Begin 1. Initialize the population of features 2. Initialize the position of features using (3) 3. While () 4. For each features 5. Calculate the fitness using (4) 6. End for 7. Sort the features based on fitness 8. Divide features into herds using (5) 9. For each features in herd ‘’ 10. update the position using (6) 11. End for 12. Perform shuffling process using (7)and (8) 13. For each features in merged herd ‘’ 14. find maximum fitness using (9) 15. End for 16. Replace global best solution 17. t = t + 1 18. Go to step 3 19. End while 20. Return (optimal features) End |
3.3. Rosenthal Correlative Stochastic-Tuned Deep Multilayer Perceptron Classifier-Based Emotional Footprint Identification
| Algorithm 2: Rosenthal Correlative Stochastic-Tuned Deep Multilayer Perceptron Classifier-Based Emotional Footprint Classification |
| Input: Dataset, selected optimal features , data samples ‘ Output: Increase the emotional footprint classification accuracy |
Begin
|
4. Experimental Setup
5. Performance Comparison Analysis
5.1. Performance Analysis of Accuracy
5.2. Performance Analysis of Precision
5.3. Performance Analysis of Recall
5.4. Performance Analysis of F1 Scores
5.5. Performance Analysis of Emotional Footprint Identification Time
5.6. Confusion Matrix
5.7. Ablation Study
5.8. Pairwise Paired T-Test
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Zhou, K.; Sisman, B.; Rana, R.; Schuller, B.W.; Li, H. Emotion Intensity and its Control for Emotional Voice Conversion. IEEE Trans. Affect. Comput. 2023, 14, 31–48. [Google Scholar] [CrossRef] [Scilit]
- Fu, F.; Ai, W.; Yang, F.; Shou, Y.; Meng, T.; Li, K. SDR-GNN: Spectral Domain Reconstruction Graph Neural Network for incomplete multimodal learning in conversational emotion recognition. Knowl.-Based Syst. 2025, 309, 112825. [Google Scholar] [CrossRef] [Scilit]
- Ai, W.; Zhang, F.; Shou, Y.; Meng, T.; Chen, H.; Li, K. Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph Spectrum. Proc. AAAI Conf. Artif. Intell. 2025, 39, 11418–11426. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Li, Y.; Pandelea, V.; Ge, M.; Zhu, L.; Cambria, E. ECPEC: Emotion-Cause Pair Extraction in Conversations. IEEE Trans. Affect. Comput. 2023, 14, 1754–1765. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.; Qin, Y.; Zhang, W.; Wu, J.; Li, M.; Gai, Y.; Jiang, F.; Lee, T. iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for Speech Synthesis based on Disentanglement between Prosody and Timbre. IEEE/ACM Trans. Audio Speech Lang. Process. 2023, 31, 1693–1705. [Google Scholar] [CrossRef] [Scilit]
- Alsaadawı, H.F.T.; Daş, R. Multimodal Emotion Recognition Using Bi-LG-GCN for MELD Dataset. Balk. J. Electr. Comput. Eng. 2024, 12, 36–46. [Google Scholar] [CrossRef] [Scilit]
- Nassif, A.B.; Shahin, I.; Lataifeh, M.; Elnagar, A.; Nemmour, N. Empirical Comparison between Deep and Classical Classifiers for Speaker Verification in Emotional Talking Environments. Information 2022, 13, 456. [Google Scholar] [CrossRef] [Scilit]
- Triantafyllopoulos, A.; Reichel, U.; Liu, S.; Huber, S.; Eyben, F.; Schuller, B.W. Multistage linguistic conditioning of convolutional layers for speech emotion recognition. Front. Comput. Sci. 2023, 5, 1072479. [Google Scholar] [CrossRef] [Scilit]
- Meng, T.; Zhang, F.; Shou, Y.; Shao, H.; Ai, W.; Li, K. Masked Graph Learning With Recurrent Alignment for Multimodal Emotion Recognition in Conversation. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 4298–4312. [Google Scholar] [CrossRef] [Scilit]
- Meng, T.; Shou, Y.; Ai, W.; Yin, N.; Li, K. Deep Imbalanced Learning for Multimodal Emotion Recognition in Conversations. IEEE Trans. Artif. Intell. 2024, 5, 6472–6487. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wang, Y.; Yang, X.; Im, S.-K. Speech emotion recognition based on Graph-LSTM neural network. EURASIP J. Audio, Speech, Music. Process. 2023, 2023, 40. [Google Scholar] [CrossRef] [Scilit]
- Bhangale, K.; Kothandaraman, M. Speech Emotion Recognition Based on Multiple Acoustic Features and Deep Convolutional Neural Network. Electronics 2023, 12, 839. [Google Scholar] [CrossRef] [Scilit]
- Chowdhury, J.H.; Ramanna, S.; Kotecha, K. Speech emotion recognition with light weight deep neural ensemble model using hand crafted features. Sci. Rep. 2025, 15, 11824. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Zhang, S.; Li, P. Multi-modal emotion recognition in conversation based on prompt learning with text-audio fusion features. Sci. Rep. 2025, 15, 8855. [Google Scholar] [CrossRef] [Scilit]
- Pallewela, N.; Alahakoon, D.; Adikari, A.; Pierce, J.E.; Rose, M.L. Optimizing Speech Emotion Recognition with Machine Learning Based Advanced Audio Cue Analysis. Technologies 2024, 12, 111. [Google Scholar] [CrossRef] [Scilit]
- Filali, H.; Boulealam, C.; El Fazazy, K.; Mahraz, A.M.; Tairi, H.; Riffi, J. Meaningful Multimodal Emotion Recognition Based on Capsule Graph Transformer Architecture. Information 2025, 16, 40. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wang, X.; Lv, G.; Zeng, Z. GraphCFC: A Directed Graph Based Cross-Modal Feature Complementation Approach for Multimodal Conversational Emotion Recognition. IEEE Trans. Multimed. 2023, 26, 77–89. [Google Scholar] [CrossRef] [Scilit]
- Duong, A.-Q.; Ho, N.-H.; Pant, S.; Kim, S.; Kim, S.-H.; Yang, H.-J. Residual Relation-Aware Attention Deep Graph-Recurrent Model for Emotion Recognition in Conversation. IEEE Access 2024, 12, 2349–2360. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wang, X.; Lv, G.; Zeng, Z. GraphMFT: A graph network based multimodal fusion technique for emotion recognition in conversation. Neurocomputing 2023, 550, 126427. [Google Scholar] [CrossRef] [Scilit]
- Fan, C.; Lin, J.; Mao, R.; Cambria, E. Fusing pairwise modalities for emotion recognition in conversations. Inf. Fusion 2024, 106, 102306. [Google Scholar] [CrossRef] [Scilit]
- Shixin, P.; Kai, C.; Tian, T.; Jingying, C. An autoencoder-based feature level fusion for speech emotion recognition. Digit. Commun. Networks 2024, 10, 1341–1351. [Google Scholar] [CrossRef] [Scilit]
- Shou, Y.; Liu, H.; Cao, X.; Meng, D.; Dong, B. A Low-Rank Matching Attention Based Cross-Modal Feature Fusion Method for Conversational Emotion Recognition. IEEE Trans. Affect. Comput. 2024, 16, 1177–1189. [Google Scholar] [CrossRef] [Scilit]
- Chen, W.; Xing, X.; Chen, P.; Xu, X. Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition. IEEE Trans. Affect. Comput. 2024, 15, 1711–1724. [Google Scholar] [CrossRef] [Scilit]
- Lian, Z.; Sun, L.; Sun, H.; Chen, K.; Wen, Z.; Gu, H.; Liu, B.; Tao, J. GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition. Inf. Fusion 2024, 108, 102367. [Google Scholar] [CrossRef] [Scilit]
- Pentari, A.; Kafentzis, G.; Tsiknakis, M. Speech emotion recognition via graph-based representations. Sci. Rep. 2024, 14, 4484. [Google Scholar] [CrossRef] [Scilit]
- Jiang, D.; Liu, H.; Tu, G.; Wei, R.; Cambria, E. Self-supervised utterance order prediction for emotion recognition in conversations. Neurocomputing 2024, 577, 127370. [Google Scholar] [CrossRef] [Scilit]
- Makhmudov, F.; Kultimuratov, A.; Cho, Y.-I. Enhancing Multimodal Emotion Recognition through Attention Mechanisms in BERT and CNN Architectures. Appl. Sci. 2024, 14, 4199. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Mei, H.; Jia, L.; Zhang, X. Multimodal Emotion Recognition in Conversation Based on Hypergraphs. Electronics 2023, 12, 4703. [Google Scholar] [CrossRef] [Scilit]
- Wang, B.; Dong, G.; Zhao, Y.; Li, R.; Cao, Q.; Hu, K.; Jiang, D. Hierarchically stacked graph convolution for emotion recognition in conversation. Knowl.-Based Syst. 2023, 263, 110285. [Google Scholar] [CrossRef] [Scilit]
- Lu, N.; Han, Z.; Han, M.; Qian, J. Bi-stream graph learning based multimodal fusion for emotion recognition in conversation. Inf. Fusion 2024, 106, 102272. [Google Scholar] [CrossRef] [Scilit]
- Chouhayebi, H.; Mahraz, M.A.; Riffi, J.; Tairi, H.; Alioua, N. Human Emotion Recognition Based on Spatio-Temporal Facial Features Using HOG-HOF and VGG-LSTM. Computers 2024, 13, 101. [Google Scholar] [CrossRef] [Scilit]
- Azizian, P.; Honarmand, M.; Jaiswal, A.; Kline, A.; Dunlap, K.; Washington, P.; Wall, D.P. Multimodal LLM vs. Human-Measured Features for AI Predictions of Autism in Home Videos. Algorithms 2025, 18, 687. [Google Scholar] [CrossRef] [Scilit]
- Abuqaddom, I.; Mahafzah, B.A.; Faris, H. Oriented stochastic loss descent algorithm to train very deep multi-layer neural networks without vanishing gradients. Knowl.-Based Syst. 2021, 230, 107391. [Google Scholar] [CrossRef] [Scilit]









| S. No | Column Headers | Features | Description |
|---|---|---|---|
| 1 | “0”–“34” | Emotional categories [0–6] | 0—no emotion 1—angry 2—disgust 3—fear 4—happiness 5—sadness 6—surprise Length of conversation |
| 2 | “35” | Conv_length | Length of conversation |
| 3 | “36”–“39” | Inform, question, directive, commissive | Number of utterances within the conversation that fall into their respective categories |
| 4 | “40” | Emotion_footprint_first_person | Emotional footprint of the first person in conversation [No emotion, Anger, Disgust, Trust, Happy, Sadness, Surprise Anticipation] |
| 5 | “41” | Emotion_footprint_second_person | Emotional footprint of the second person in conversation [No emotion, Anger, Disgust, Trust, Happy, Sadness, Surprise Anticipation] |
| 6 | “42” | Conv_intensity_footprint_first_person | Emotional footprint and intensity of the first person in conversation |
| 7 | “43” | Conv_intensity_footprint_second_person | Emotional footprint and intensity of the second person in conversation |
| 8 | “44” | Emotion_intensity _first_person | Intensity of emotion experienced by first person in conversation |
| 9 | “45” | Emotion_intensity_second_person | Intensity of emotion experienced by second person in conversation |
| 10 | “46” | Conversation | Complete and unchanged conversation between two parties |
| 11 | “47” | Speaker 1 | Utterances made by first speaker within the conversation |
| 12 | “48” | Speaker 2 | Utterances made by second speaker within the conversation |
| Number of Data Samples | Accuracy | ||
|---|---|---|---|
| Proposed DSFO-STDMP | EMOVOX [1] | SDR-GNN [2] | |
| 180 | 0.95 | 0.87 | 0.92 |
| 360 | 0.94 | 0.8 | 0.91 |
| 540 | 0.95 | 0.75 | 0.9 |
| 720 | 0.94 | 0.77 | 0.89 |
| 900 | 0.94 | 0.8 | 0.88 |
| 1080 | 0.95 | 0.75 | 0.87 |
| 1260 | 0.96 | 0.71 | 0.88 |
| 1440 | 0.96 | 0.72 | 0.9 |
| 1620 | 0.95 | 0.75 | 0.91 |
| 1800 | 0.95 | 0.77 | 0.9 |
| Number of Data Samples | Precision | ||
|---|---|---|---|
| Proposed DSFO-STDMP | EMOVOX [1] | SDR-GNN [2] | |
| 180 | 0.96 | 0.9 | 0.92 |
| 360 | 0.94 | 0.86 | 0.91 |
| 540 | 0.92 | 0.83 | 0.89 |
| 720 | 0.91 | 0.78 | 0.88 |
| 900 | 0.93 | 0.77 | 0.89 |
| 1080 | 0.94 | 0.78 | 0.86 |
| 1260 | 0.92 | 0.81 | 0.87 |
| 1440 | 0.93 | 0.84 | 0.9 |
| 1620 | 0.92 | 0.84 | 0.88 |
| 1800 | 0.93 | 0.85 | 0.89 |
| Number of Data Samples | Recall | ||
|---|---|---|---|
| Proposed DSFO-STDMP | EMOVOX [1] | SDR-GNN [2] | |
| 180 | 0.98 | 0.95 | 0.96 |
| 360 | 0.97 | 0.92 | 0.94 |
| 540 | 0.96 | 0.91 | 0.93 |
| 720 | 0.98 | 0.91 | 0.93 |
| 900 | 0.97 | 0.92 | 0.94 |
| 1080 | 0.96 | 0.91 | 0.93 |
| 1260 | 0.98 | 0.92 | 0.94 |
| 1440 | 0.97 | 0.91 | 0.93 |
| 1620 | 0.98 | 0.92 | 0.94 |
| 1800 | 0.97 | 0.91 | 0.93 |
| Number of Data Samples | F1 Score | ||
|---|---|---|---|
| Proposed DSFO-STDMP | EMOVOX [1] | SDR-GNN [2] | |
| 180 | 0.97 | 0.92 | 0.94 |
| 360 | 0.97 | 0.88 | 0.92 |
| 540 | 0.96 | 0.86 | 0.90 |
| 720 | 0.98 | 0.84 | 0.90 |
| 900 | 0.97 | 0.83 | 0.91 |
| 1080 | 0.96 | 0.84 | 0.89 |
| 1260 | 0.98 | 0.86 | 0.90 |
| 1440 | 0.97 | 0.87 | 0.91 |
| 1620 | 0.98 | 0.88 | 0.90 |
| 1800 | 0.97 | 0.87 | 0.90 |
| Number of Data Samples | Emotional Footprint Identification Time (Sec) | ||
|---|---|---|---|
| Proposed DSFO-STDMP | EMOVOX [1] | SDR-GNN [2] | |
| 180 | 57.6 | 77.4 | 66.6 |
| 360 | 65.8 | 115 | 82.5 |
| 540 | 75.6 | 135 | 92.6 |
| 720 | 88.3 | 165 | 123.5 |
| 900 | 105.7 | 172 | 141.6 |
| 1080 | 110.5 | 185 | 153.8 |
| 1260 | 114.6 | 215 | 164.4 |
| 1440 | 121.3 | 230 | 182.3 |
| 1620 | 127.5 | 245 | 192.3 |
| 1800 | 134.9 | 275 | 213.2 |
| Methods | Two-Party Conversation with Emotional Footprint and Emotional Intensity Dataset | ||||
|---|---|---|---|---|---|
| Accuracy (%) | Precision | Recall | F1 Score | Emotional Footprint Identification Time (Sec.) | |
| SSDSSO | 0.81 | 0.78 | 0.85 | 0.82 | 145 |
| DMP | 0.85 | 0.80 | 0.89 | 0.84 | 128 |
| RCG-FT-DMLP | 0.88 | 0.84 | 0.92 | 0.88 | 112 |
| DSFO-STDMP | 0.92 | 0.91 | 0.96 | 0.93 | 95 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Jagadeesan, K.; Kumarappan, A. Divergence Shepherd Feature Optimization-Based Stochastic-Tuned Deep Multilayer Perceptron for Emotional Footprint Identification. Algorithms 2025, 18, 801. https://doi.org/10.3390/a18120801
Jagadeesan K, Kumarappan A. Divergence Shepherd Feature Optimization-Based Stochastic-Tuned Deep Multilayer Perceptron for Emotional Footprint Identification. Algorithms. 2025; 18(12):801. https://doi.org/10.3390/a18120801
Chicago/Turabian StyleJagadeesan, Karthikeyan, and Annapurani Kumarappan. 2025. "Divergence Shepherd Feature Optimization-Based Stochastic-Tuned Deep Multilayer Perceptron for Emotional Footprint Identification" Algorithms 18, no. 12: 801. https://doi.org/10.3390/a18120801
APA StyleJagadeesan, K., & Kumarappan, A. (2025). Divergence Shepherd Feature Optimization-Based Stochastic-Tuned Deep Multilayer Perceptron for Emotional Footprint Identification. Algorithms, 18(12), 801. https://doi.org/10.3390/a18120801

