SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition
Abstract
1. Introduction
- -
- This paper puts forward SE-POSTER, a sophisticated landmark-guided CNN–Transformer framework for facial emotion recognition that brings hierarchical channel-wise feature recalibration into the multi-scale representation learning pipeline. Existing POSTER-based architectures that focus mainly on spatial fusion and global dependency modeling are not able to solve the problem of adaptive channel importance at different semantic representation levels, which is the case of the proposed method.
- -
- A feature refinement scheme on multiple levels through SE guidance is adopted within the pyramid architecture to adaptively perform fine-level, mid-level, and global emotional representation recalibration before the Transformer-based attention computation. Such a model makes token embeddings more distinctive and, at the same time, it is able to model global dependencies contextually more effectively, leading to better handling of the real-world FER challenges.
- -
- In addition to landmark-guided Transformer encoding, the proposed framework also pays systematic attention to the interaction of channel-wise attention and facial emotion recognition. The paper provides evidence that modifying the channel adaptively before the self-attention computation step enhances the robustness of emotional feature learning under changes in pose, illumination, occlusion, expression intensity, and facial appearance.
- -
- Based on several experimental results obtained from RAF-DB, FERPlus, and AffectNet datasets, it was proven that SE-POSTER is capable of consistently surpassing the POSTER baseline architecture and several other recent FER methods, while at the same time it manages to be very light in computational complexity and efficient in inference behavior.
- -
- Additional hierarchical emotional representation learning insights are provided by the thorough ablation experiments at different pyramid levels, and it is demonstrated that the combined refinement of fine-level, structural, and global semantic representations leads to more effective facial emotion recognition performance.
2. Literature Review
2.1. Facial Emotion Recognition Using Convolutional Neural Networks
2.2. Transformer-Based Facial Emotion Recognition
2.3. Landmark-Guided and Multi-Scale FER Architectures
2.4. Channel Attention Mechanisms in Deep Neural Networks
3. Methodology
3.1. POSTER—The Baseline Model
3.2. The Proposed Model
4. The Experiment and Analyses
4.1. Datasets and Data Preprocessing
4.2. Comparison with SOTA Models
4.3. Ablation Study
4.4. Failure Case Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Shan, L.; Weihong, D. Deep facial expression recognition: A survey. J. Image Graph. 2020, 25, 2306–2320. [Google Scholar] [CrossRef]
- Ma, N.; Qu, B.; Wang, W.; Wang, F.; Zhang, X. SCAF-Net: A spiking cross-modal attention fusion network for multimodal emotion recognition. Biomed. Signal Process. Control 2026, 118, 109684. [Google Scholar] [CrossRef]
- Haq, M.; Athar, M.; Ahmad, S.; Ahmad, N.; Anwar, M.; Kutlimuratov, A. A Comprehensive Review of Face Detection/Recognition Algorithms and Competitive Datasets to Optimize Machine Vision. Comput. Mater. Contin. 2025, 84, 1–24. [Google Scholar] [CrossRef]
- Bian, Y.; Kim, H.; Krumhuber, E.G. A Cross-Corpus Evaluation on Spontaneous and Dynamic Facial Expressions for Automated Emotion Classification. Electronics 2026, 15, 849. [Google Scholar] [CrossRef]
- Hunafa, M.H.; Ramadhan, A.W.; Kushirayati, S.; Abka, A.F.; Mantau, A.J.; Jatmiko, W. Data Imbalance Handling in Facial Expression Recognition: A Systematic Literature Review. IEEE Access 2026, 14, 8269–8287. [Google Scholar] [CrossRef]
- Safarov, F.; Kutlimuratov, A.; Khojamuratova, U.; Abdusalomov, A.; Cho, Y.-I. Enhanced AlexNet with Gabor and Local Binary Pattern Features for Improved Facial Emotion Recognition. Sensors 2025, 25, 3832. [Google Scholar] [CrossRef] [PubMed]
- Hebri, D.; Nuthakki, R.; Digal, A.K.; Venkatesan, K.G.S.; Chawla, S.; Raghavendra Reddy, C. Effective facial expression recognition system using machine learning. In Proceedings of the EAI Endorsed Transactions on Internet of Things; European Alliance for Innovation: Gent, Belgium, 2024; p. 10. [Google Scholar]
- Grover, R.; Bansal, S. Enhancing facial expression recognition in uncontrolled environment: A lightweight CNN approach with pre-processing. Neural Comput. Appl. 2025, 37, 7363–7378. [Google Scholar] [CrossRef]
- Qadir, I.; Iqbal, M.A.; Ashraf, S.; Akram, S. A fusion of CNN And SIFT For multicultural facial expression recognition. Multimed. Tools Appl. 2025, 84, 33505–33523. [Google Scholar] [CrossRef]
- Kashef, A.; Wang, Y.; Assafi, M.N.; Ma, J.; Wang, J.; Jones, J.A.; Thiamwong, L. Developing A novel AI enabled extended reality system for real-time automatic facial expression recognition and system performance evaluation. Adv. Eng. Inform. 2025, 65, 103207. [Google Scholar] [CrossRef]
- Nawaz, U.; Saeed, Z.; Atif, K. A Novel Transformer-based approach for adult’s facial emotion recognition. IEEE Access. 2025, 13, 56485–56508. [Google Scholar] [CrossRef]
- Wang, Y.; Pan, K.; Shao, Y.; Ma, J.; Li, X. Applying a convolutional vision transformer for emotion recognition in children with autism: Fusion of facial expressions and speech features. Appl. Sci. 2025, 15, 3083. [Google Scholar] [CrossRef]
- Zheng, C.; Mendieta, M.; Chen, C. Poster: A pyramid cross-fusion transformer network for facial expression recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 3146–3155. [Google Scholar]
- Radočaj, P.; Martinović, G. Emotion Recognition in Autistic Children Through Facial Expressions Using Advanced Deep Learning Architectures. Appl. Sci. 2025, 15, 9555. [Google Scholar] [CrossRef]
- Agarwal, A.; Susan, S. Attention-augmented squeeze-and-excitation enhanced mobile network for occluded facial expression recognition in resource-constrained environments. Signal Image Video Process. 2025, 19, 687. [Google Scholar] [CrossRef]
- Manavand, M.R.; Salarifar, M.H.; Ghavami, M.; Taghipour-Gorjikolaie, M. Driver’s facial expression recognition by using deep local and global features. Inf. Sci. 2025, 692, 121658. [Google Scholar] [CrossRef]
- Mohamed, A.; Nii, R.; Binga, K.; Modi, S.; Nagaveni, P.; Nnadhini, T.J. Facial expression recognition in real-time surveillance using CNN and transfer learning with ResNet-50. In 2025 International Conference on Automation and Computation (AUTOCOM); IEEE: New York, NY, USA, 2025; pp. 229–234. [Google Scholar]
- Bhati, V.S.; Tiwari, N.; Chawla, M. A generalized zero-shot deep learning classifier for emotion recognition using facial expression images. IEEE Access 2025, 13, 18687–18700. [Google Scholar] [CrossRef]
- Kumar, R.; Corvisieri, G.; Fici, T.F.; Hussain, S.I.; Tegolo, D.; Valenti, C. Transfer learning for facial expression recognition. Information 2025, 16, 320. [Google Scholar] [CrossRef]
- Akeh, L.J.; Kusuma, G.P. Mixed emotion recognition through facial expression using transformer-based model. Stat. Optim. Inf. Comput. 2025, 13, 531–546. [Google Scholar] [CrossRef]
- Tagmatova, Z.; Umirzakova, S.; Kutlimuratov, A.; Abdusalomov, A.; Im Cho, Y. A Hyper-Attentive Multimodal Transformer for Real-Time and Robust Facial Expression Recognition. Appl. Sci. 2025, 15, 7100. [Google Scholar] [CrossRef]
- Makhmudov, F.; Kutlimuratov, A.; Cho, Y.-I. Hybrid LSTM–Attention and CNN Model for Enhanced Speech Emotion Recognition. Appl. Sci. 2024, 14, 11342. [Google Scholar] [CrossRef]
- Sun, Z.; Liu, H.; Li, H.; Li, Y.; Zhang, W. AVERFormer: End-to-end audio-visual emotion recognition transformer framework with balanced modal contributions. Digit. Signal Process. 2025, 161, 105081. [Google Scholar] [CrossRef]
- Sun, R.; Zhang, Z.; Liu, H.; Zhao, L.; Zhou, Q.; Liu, Z. DacFER: Dual Attention Correction Learning for Efficient Facial Expression Recognition. In 2024 7th International Conference on Electronics Technology (ICET); IEEE: New York, NY, USA, 2024; pp. 941–945. [Google Scholar]
- Zhang, Q.; Liu, Y.; Zhu, B.; Han, X.; Zhang, R.; Xiao, J.; Wang, Z. Deep multi-modal fusion transformer for emotion recognition. Eng. Appl. Artif. Intell. 2026, 168, 113967. [Google Scholar] [CrossRef]
- Xiong, K.; Qing, L.; Li, L.; Guo, L.; Peng, Y. Facial expression recognition based on local–global information reasoning and spatial distribution of landmark features. Vis. Comput. 2025, 41, 535–548. [Google Scholar] [CrossRef]
- Do, H.Q.; Thanh, H.V.; Phuong, T.M. A Study on Fusion Strategies of Facial Landmark-Based Heatmap for Facial Expression Recognition. KSII Trans. Internet Inf. Syst. 2025, 19, 3602. [Google Scholar] [CrossRef]
- Jiang, W.; Zhao, Z.; Wang, L.; Liu, F.; Qing, C.; Xing, X.; Xu, X.; Fan, W.; Jin, Z. A dual uncertainty-aware fusion framework for face expression recognition in the wild. Expert Syst. Appl. 2025, 298, 129567. [Google Scholar] [CrossRef]
- Praveen, R.G.; De Melo, W.C.; Ullah, N.; Aslam, H.; Zeeshan, O.; Denorme, T.; Pedersoli, M.; Koerich, A.L.; Bacon, S.; Cardinal, P.; et al. A joint cross-attention model for audio-visual fusion in dimensional emotion recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE Computer Society: Washington, DC, USA, 2022; pp. 2486–2495. [Google Scholar]
- Khan, T.; Yasir, M.; Choi, C. Attention-enhanced optimized deep ensemble network for effective facial emotion recognition. Alex. Eng. J. 2025, 119, 111–123. [Google Scholar] [CrossRef]
- Ryumina, E.; Ryumin, D.; Axyonov, A.; Ivanko, D.; Karpov, A. Multi-corpus emotion recognition method based on cross-modal gated attention fusion. Pattern Recognit. Lett. 2025, 190, 192–200. [Google Scholar] [CrossRef]
- Yi, M.H.; Kwak, K.C.; Shin, J.H. HyFusER: Hybrid multimodal transformer for emotion recognition using dual cross modal attention. Appl. Sci. 2025, 15, 1053. [Google Scholar] [CrossRef]
- Li, S.; Deng, W. Reliable Crowdsourcing and Deep Locality-Preserving Learning for Unconstrained Facial Expression Recognition. IEEE Trans. Image Process. 2019, 28, 356–370. [Google Scholar] [CrossRef] [PubMed]
- Barsoum, E.; Zhang, C.; Ferrer, C.C.; Zhang, Z. Training deep networks for facial expression recognition with crowd-sourced label distribution. In Proceedings of the 18th ACM International Conference on Multimodal Interaction (ICMI ’16); Association for Computing Machinery: New York, NY, USA, 2016; pp. 279–283. [Google Scholar] [CrossRef]
- Mollahosseini, A.; Hasani, B.; Mahoor, M.H. AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild. IEEE Trans. Affect. Comput. 2019, 10, 18–31. [Google Scholar] [CrossRef]
- Mao, J.; Xu, R.; Yin, X.; Chang, Y.; Nie, B.; Huang, A.; Wang, Y. Poster++: A simpler and stronger facial expression recognition network. Pattern Recognit. 2025, 157, 110951. [Google Scholar] [CrossRef]
- Rakhimovich, A.M.; Kadirbergenovich, K.K.; Ishkobilovich, Z.M.; Kadirbergenovich, K.J. Logistic Regression with Multi-Connected Weights. J. Comput. Sci. 2024, 20, 1051–1058. [Google Scholar] [CrossRef]
- Kabulov, A.; Babadzhanov, A.; Baizhumanov, A.; Saymanov, I.; Babadjanov, A. Algorithms for Solving Systems of Boolean Equations Based on the Transformation of Logical Expressions. Mathematics 2026, 14, 594. [Google Scholar] [CrossRef]
- Madrakhimov, S.; Makharov, K.; Khurramov, A. On the Transparency of Decision-Making in Classification by Precedents with Fuzzy Descriptions. IEEE Access 2025, 13, 173656–173664. [Google Scholar] [CrossRef]




| Model | RAF-DB | FERPlus | AffectNet (7) |
|---|---|---|---|
| POSTER | 92.05 | 91.62 | 67.31 |
| POSTER V2 | 92.21 | - | 67.49 |
| SE-POSTER | 92.42 | 91.81 | 67.65 |
| Dataset | Total Samples | Evaluated Subset | Classes | Training Samples | Testing Samples | Resolution | Protocol |
|---|---|---|---|---|---|---|---|
| RAF-DB | 29,672 | 15,339 | 7 | 12,271 | 3068 | 100 × 100 | Official split |
| FERPlus | 35,887 | 35,887 | 8 | Official train split | Official test split | 48 × 48 | Official split |
| AffectNet | >1,000,000, approximately 450 K labeled | Official evaluated subset | 7/8 | Official train split | Official validation/test split | Variable | Standard prot |
| Method | Image Backbone | Attention Mechanism | Multi-Scale Modeling | Landmark Guidance | RAF-DB Accuracy (%) |
|---|---|---|---|---|---|
| POSTER [11] | CNN + Transformer | Cross-Fusion Transformer | ✓ | ✓ | 92.05 |
| POSTER++ [29] | Simplified POSTER | Window-based Cross-Attention | ✓ | ✓ | 92.21 |
| AU-Aware ViT [21] | Vision Transformer | AU-aware cross-domain features | X | X | 88.54 |
| ViT-FER [10] | ViT-Base | Self-Attention | X | X | 88.32 |
| Hybrid Model [28] | Hybrid | Local + Global Fusion | ✓ | X | 92.37 |
| The proposed Model | ResNet-18+ Transformer | SE+ Self-Attention | ✓ | ✓ | 92.42 |
| Model | Params (M) | FLOPs (G) | GPU Memory (MB) | Inference Time (ms/Image) | Accuracy (%) |
|---|---|---|---|---|---|
| POSTER | 43.07 | 12.00 | 2500 | 25 | 92.21 |
| SE-POSTER | 32.09 | 8.04 | 1200 | 23 | 92.42 |
| Metric | Value (%) |
|---|---|
| Accuracy | 92.78 |
| Macro Precision | 90.84 |
| Macro Recall | 89.28 |
| Macro F1-Score | 89.91 |
| Weighted Precision | 92.65 |
| Weighted Recall | 92.78 |
| Weighted F1-Score | 92.61 |
| Balanced Accuracy | 89.28 |
| Dataset | Test Samples | Baseline Model | Proposed Model | Baseline Accuracy (%) | Proposed Accuracy (%) | 95% CI of Proposed Accuracy (%) | McNemar Statistic | p-Value | Significant at α = 0.05 |
|---|---|---|---|---|---|---|---|---|---|
| RAF-DB | 3068 | POSTER | SE-POSTER | 92.05 | 92.78 | 91.86–93.70 | 4.14 | 0.042 | Yes |
| FERPlus | 3589 | POSTER | SE-POSTER | 91.62 | 91.81 | 90.91–92.71 | 0.40 | 0.529 | No |
| AffectNet (7) | 3500 | POSTER | SE-POSTER | 67.31 | 67.65 | 66.10–69.20 | 0.68 | 0.410 | No |
| Model Version | SE Fine | SE Mid | SE Global | Accuracy (%) | F1-Score (%) |
|---|---|---|---|---|---|
| POSTER baseline | × | × | × | 92.05 | 91.68 |
| +SE Fine only | ✓ | × | × | 92.13 | 91.67 |
| +SE Mid only | × | ✓ | × | 92.21 | 91.75 |
| +SE Global only | × | × | ✓ | 92.18 | 91.65 |
| SE-POSTER All | ✓ | ✓ | ✓ | 92.78 | 91.74 |
| Method | RAF-DB (%) | FERPlus (%) | AffectNet (7) (%) | Average Accuracy (%) |
|---|---|---|---|---|
| ViT-FER | 88.32 | 87.14 | 63.45 | 79.64 |
| AU-Aware ViT | 88.54 | 87.63 | 64.12 | 80.10 |
| Hybrid Model | 92.37 | 91.24 | 66.91 | 83.51 |
| POSTER | 92.05 | 91.62 | 67.31 | 83.66 |
| POSTER++ | 92.21 | – | 67.49 | – |
| Proposed SE-POSTER | 92.78 | 91.81 | 67.65 | 84.08 |
| Setting | Optimizer | Initial LR | Weight Decay | Batch Size | Scheduler | SE Reduction Ratio | Validation Accuracy (%) | Validation F1-Score (%) | Validation Loss | Selected |
|---|---|---|---|---|---|---|---|---|---|---|
| H1 | AdamW | 5 × 10−5 | 1 × 10−4 | 32 | Cosine annealing | 16 | 91.86 | 91.22 | 0.276 | No |
| H2 | 1 × 10−4 | 1 × 10−4 | 32 | 16 | 92.31 | 91.78 | 0.238 | Yes | ||
| H3 | 2 × 10−4 | 1 × 10−4 | 32 | 16 | 92.06 | 91.49 | 0.259 | No | ||
| H4 | 1 × 10−4 | 5 × 10−5 | 32 | 16 | 92.18 | 91.63 | 0.247 | No | ||
| H5 | 1 × 10−4 | 1 × 10−3 | 32 | 16 | 91.92 | 91.35 | 0.270 | No | ||
| H6 | 1 × 10−4 | 1 × 10−4 | 16 | 16 | 92.12 | 91.55 | 0.251 | No | ||
| H7 | 1 × 10−4 | 1 × 10−4 | 32 | 16 | 91.98 | 91.42 | 0.265 | No | ||
| H8 | 1 × 10−4 | 1 × 10−4 | 32 | 8 | 92.20 | 91.66 | 0.244 | No | ||
| H9 | 1 × 10−4 | 1 × 10−4 | 32 | 32 | 92.09 | 91.53 | 0.255 | No |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kutlimuratov, A.; Sharipov, K.; Allayarov, P.; Iskandarova, S.; Latyfskiy, R.; Tolibaeva, G.; Makhmudov, F. SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition. Informatics 2026, 13, 123. https://doi.org/10.3390/informatics13080123
Kutlimuratov A, Sharipov K, Allayarov P, Iskandarova S, Latyfskiy R, Tolibaeva G, Makhmudov F. SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition. Informatics. 2026; 13(8):123. https://doi.org/10.3390/informatics13080123
Chicago/Turabian StyleKutlimuratov, Alpamis, Kongratbay Sharipov, Piratdin Allayarov, Sayyora Iskandarova, Ruslan Latyfskiy, Gulchehra Tolibaeva, and Fazliddin Makhmudov. 2026. "SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition" Informatics 13, no. 8: 123. https://doi.org/10.3390/informatics13080123
APA StyleKutlimuratov, A., Sharipov, K., Allayarov, P., Iskandarova, S., Latyfskiy, R., Tolibaeva, G., & Makhmudov, F. (2026). SE-POSTER: Channel-Enhanced Landmark Guided Transformer for Facial Emotion Recognition. Informatics, 13(8), 123. https://doi.org/10.3390/informatics13080123

