Head-Movement-Robust Micro-Expression Detection Method via 3D Motion Correction and Transformers
Abstract
1. Introduction
- (1)
- We propose a 3D head-motion-correction-guided deep-weighted optical flow extraction method, which decomposes and suppresses global rigid head motion while leveraging depth information to enhance local optical flow weighting, thereby obtaining robust temporal representations of micro-expressions.
- (2)
- We constructed a Transformer-based sequential modeling and classification module. It leverages the model’s strong long-range dependency modeling capability to efficiently encode the preprocessed temporal optical flow features.
- (3)
- We conducted systematic experiments on real-scene micro-expression datasets with noticeable head movements. The results show that the proposed two-stage hybrid framework effectively overcomes head motion interference. It significantly outperforms existing mainstream methods in detection performance.
2. Related Work
2.1. Methods for Suppressing Head Movement Interference
2.2. Micro-Expression Detection Based on Handcrafted Features
2.3. Micro-Expression Detection Based on Deep Learning
2.4. Micro-Expression Detection Combining Manual Features and Deep Learning
3. Methodology
3.1. Head Pose Estimation Based on Depth Compensation
- (1)
- RGB-D Dual-Branch Feature Fusion: This module constructs a four-channel RGB-D input and adopts a dual-branch network architecture. The RGB branch uses RepVGG as the backbone to extract visual appearance features, while the depth branch employs a lightweight convolutional network to extract geometric features robust to illumination variations. To address multimodal feature alignment, the network incorporates a squeeze-and-excitation channel-attention mechanism in both branches. This mechanism adaptively enhances the feature responses of key anatomical regions that are highly discriminative for pose estimation (e.g., nasal bridge, eye sockets), achieving depth-guided intelligent feature fusion.
- (2)
- 6D Rotation Representation and End-to-End Regression: The model adopts a continuous and unambiguous 6D rotation representation. The fused multimodal features are passed through fully connected layers to directly regress this 6D representation, and a Gram–Schmidt orthogonalization process is used to recover a strictly orthogonal rotation matrix. This design eliminates the reliance on intermediate facial landmark detection common in traditional methods, achieving efficient end-to-end regression from input to pose.
- (3)
- Nasal Bridge Normal Anatomical Constraint Loss: To enhance the geometric plausibility of rotation estimation, an anatomically driven orientation–constraint loss is introduced. This loss term leverages the relatively stable geometric prior of the human nasal bridge region under pose variations. By comparing the angle between the actual nasal bridge normal vector estimated from the depth map and the ideal normal vector obtained by transforming via the predicted rotation matrix, a constraint term is constructed. It is then weighted and combined with the standard geodesic loss to form the final rotation loss . This loss function forces the network’s predicted rotation directions to align with the head’s local anatomical structure, effectively suppressing pose deviations caused by noise or occlusion.
3.2. Video Framing Strategy Based on Head Motion Dynamics
3.3. Depth-Weighted Optical Flow Compensation
3.4. Micro-Expression Detection Based on Transformer
3.4.1. Feature Preprocessing and Sequence Construction
- Rotation component elimination (Section 3.1): Through head pose estimation, the theoretical optical flow components , induced by head rotational motion are calculated and subtracted.
- Translation component elimination (Section 3.3): Through the depth-weighted compensation model, the theoretical optical flow components (, ) caused by head translational motion are adaptively eliminated. It is important to note that the core reference baseline for this compensation model is derived from the depth and optical flow information of the nose tip and its surrounding area. Therefore, in subsequent feature extraction, this nasal region is excluded to avoid re-inputting signals that have already been used in the compensation calculation as recognition features, thereby ensuring the independence and validity of the feature sources.
3.4.2. Reasoning and Post-Processing Flow
4. Experiment
4.1. Dataset
4.2. Experimental Setup
4.3. LOSO Cross-Validation Experiment
4.4. Ablation Experiment
4.4.1. Deep Information Utilization Ablation Experiment
4.4.2. Loss-Function-Type Ablation Experiment
4.4.3. Time-Series Window-Size Ablation Experiment
4.5. Generalization Ability Verification
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ekman, P. Telling Lies: Clues to Deceit in the Marketplace, Politics, and Marriage, Revised ed.; Norton: New York, NY, USA, 2009. [Google Scholar]
- Ben, X.; Ren, Y.; Zhang, J.; Wang, S.-J.; Kpalma, K.; Meng, W.; Liu, Y.-J. Video-based facial micro-expression analysis: A survey of datasets, features and algorithms. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 5826–5846. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, T.; Pu, T.; Wu, H.; Xie, Y.; Liu, L.; Lin, L. Crossdomain facial expression recognition: A unified evaluation benchmark and adversarial graph learning. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 9887–9903. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shao, Z.; Zhu, H.; Zhou, Y.; Xiang, X.; Liu, B.; Yao, R.; Ma, L. Facial action unit detection by adaptively constraining self-attention and causally deconfounding sample. Int. J. Comput. Vis. 2025, 133, 1711–1726. [Google Scholar] [CrossRef] [Scilit]
- Yan, W.-J.; Wu, Q.; Liang, J.; Chen, Y.-H.; Fu, X. How fast are the leaked facial expressions: The duration of microexpressions. J. Nonverbal Behav. 2013, 37, 217–230. [Google Scholar] [CrossRef] [Scilit]
- Chaudhry, R.; Ravichandran, A.; Hager, G.; Vidal, R. Histograms of oriented optical flow and Binet-Cauchy kernels on nonlinear dynamical systems for the recognition of human actions. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2009; pp. 1932–1939. [Google Scholar]
- Zhao, G.; Pietikainen, M. Dynamic texture recognition using local binary patterns with an application to facial expressions. IEEE Trans. Pattern Anal. Mach. Intell. 2007, 29, 915–928. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Davison, A.K.; Yap, M.H.; Lansley, C. Microfacial movement detection using individualised baselines and histogram-based descriptors. In Proceedings of the 2015 IEEE International Conference on Systems, Man, and Cybernetics (SMC); IEEE: New York, NY, USA, 2015; pp. 1864–1869. [Google Scholar]
- Khor, H.-Q.; See, J.; Phan, R.C.W.; Lin, W. Enriched long-term recurrent convolutional network for facial microexpression recognition. In Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018); IEEE: New York, NY, USA, 2018; pp. 667–674. [Google Scholar]
- Reddy, S.P.T.; Karri, S.T.; Dubey, S.R.; Mukherjee, S. Spontaneous facial micro-expression recognition using 3D spatiotemporal convolutional neural networks. In Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary, 14–19 July 2019; pp. 1–8. [Google Scholar]
- Huan, J.; Li, M.; Zhou, H. Emotion-aware Adaptation of CLIP Model for Facial Expression Recognition. Artif. Intell. Rev. 2026, 59, 66. [Google Scholar] [CrossRef] [Scilit]
- Shao, Z.; Cheng, Y.; Li, F.; Zhou, Y.; Lu, X.; Xie, Y.; Ma, L. MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style Convolution. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 8756–8768. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, L.-W.; Li, J.; Wang, S.-J.; Duan, X.-H.; Yan, W.-J.; Xie, H.-Y.; Huang, S.-C. Spatio-temporal fusion for macro-and micro-expression spotting in long video sequences. In Proceedings of the 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020); IEEE: New York, NY, USA, 2020; pp. 734–741. [Google Scholar]
- Yuhong, H. Research on micro-expression spotting method based on optical flow features. In Proceedings of the 29th ACM International Conference on Multimedia, Virtual, 20–24 October 2021. [Google Scholar]
- Yu, W.; Jiang, J.; Li, Y. LGSNet: A two-stream network for micro and macro-expression spotting with background modeling. IEEE Trans. Affect. Comput. 2024, 15, 223–240. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Hong, X.; Moilanen, A.; Huang, X.; Pfister, T.; Zhao, G.; Pietikäinen, M. Towards reading hidden emotions: A comparative study of spontaneous micro-expression spotting and recognition methods. IEEE Trans. Affect. Comput. 2017, 9, 563–577. [Google Scholar] [CrossRef] [Scilit]
- Davison, A.; Merghani, W.; Lansley, C.; Ng, C.-C.; Yap, M.H. Objective micro-facial movement detection using FACS-based regions and baseline evaluation. In Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018); IEEE: New York, NY, USA, 2018; pp. 642–649. [Google Scholar]
- Li, J.; Soladie, C.; Seguier, R. Local temporal pattern and data augmentation for micro-expression spotting. IEEE Trans. Affect. Comput. 2020, 14, 811–822. [Google Scholar] [CrossRef] [Scilit]
- Shreve, M.; Godavarthy, S.; Goldgof, D.; Sarkar, S. Macro-and micro-expression spotting in long videos using spatio-temporal strain. In Proceedings of the IEEE Conference on Automatic Face and Gesture Recognition; IEEE: New York, NY, USA, 2011; pp. 51–56. [Google Scholar]
- Patel, D.; Zhao, G.; Pietikäinen, M. Spatiotemporal integration of optical flow vectors for micro-expression detection. In Proceedings of the 16th International Conference on Advanced Concepts for Intelligent Vision Systems (ACIVS), Catania, Italy, 26–29 October 2015. [Google Scholar]
- Liu, Y.-J.; Zhang, J.-K.; Yan, W.-J.; Wang, S.-J.; Zhao, G.; Fu, X. A main directional mean optical flow feature for spontaneous micro-expression recognition. IEEE Trans. Affect. Comput. 2015, 7, 299–310. [Google Scholar] [CrossRef] [Scilit]
- Lei, L.; Li, J.; Chen, T.; Li, S. A novel graph-TCN with a graph structured representation for micro-expression recognition. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020. [Google Scholar]
- Wei, M.; Zheng, W.; Zong, Y.; Jiang, X.; Lu, C.; Liu, J. A novel micro-expression recognition approach using attention-based magnification-adaptive networks. In Proceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2022; pp. 2420–2424. [Google Scholar]
- Liu, G.; Huang, S.; Wang, G.; Li, M. EMRNet: Enhanced micro-expression recognition network with attention and distance correlation. Artif. Intell. Rev. 2025, 58, 176. [Google Scholar] [CrossRef] [Scilit]
- Verma, M.; Vipparthi, S.K.; Singh, G.; Murala, S. LEARNet: Dynamic imaging network for micro expression recognition. IEEE Trans. Image Process. 2020, 29, 1618–1627. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, L.; Mao, Q.; Xue, L. Dual-inception network for cross-database micro-expression recognition. In Proceedings of the 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019); IEEE: New York, NY, USA, 2019; pp. 1–5. [Google Scholar]
- Yang, X.; Yang, H.; Li, J.; Wang, S.-J. Simple but effective in-the-wild micro-expression spotting based on head pose segmentation. In Proceedings of the 3rd Workshop on Facial Micro-Expression: Advanced Techniques for Multi-Modal Facial Expression Analysis (FME ’23); Association for Computing Machinery: New York, NY, USA, 2023; pp. 9–16. [Google Scholar]
- Jiang, F.; Huang, S.; Li, M. Deep6DHead: A 6D head pose estimation method based on deep feature enhancement. Symmetry 2026, 18, 705. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Huang, S.; Li, M. MSOF: A main and secondary bi-directional optical flow feature method for spotting micro-expression. Neurocomputing 2025, 630, 129676. [Google Scholar] [CrossRef] [Scilit]
- Verkruysse, W.; Svaasand, L.O.; Nelson, J.S. Remote plethysmographic imaging using ambient light. Opt. Express 2008, 16, 21434–21445. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019. [Google Scholar]
- See, J.; Yap, M.H.; Li, J.; Hong, X.; Wang, S. MEGC 2019—The second facial micro-expressions grand challenge. In Proceedings of the 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019); IEEE: New York, NY, USA, 2019; pp. 1–5. [Google Scholar]
- Li, J.; Yap, M.H.; Cheng, W.-H.; See, J.; Hong, X.; Li, X.; Wang, S.-J.; Davison, A.K.; Li, Y.; Dong, Z. MEGC2022: ACM multimedia 2022 micro-expression grand challenge. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022. [Google Scholar]
- Husák, P.; Cech, J.; Matas, J. Spotting facial micro-expressions in the wild. In Proceedings of the 22nd Computer Vision Winter Workshop (CVWW), Retz, Austria, 6–8 February 2017; Kropatsch, W.G., Janusch, I., Artner, N.M., Eds.; PRIP Group, TU Wien: Vienna, Austria, 2017; pp. 1–9. [Google Scholar]
- Li, J.; Dong, Z.; Lu, S.; Wang, S.-J.; Yan, W.-J.; Ma, Y.; Liu, Y.; Huang, C.; Fu, X. CAS(ME)3: A third generation facial spontaneous micro-expression database with depth information and high ecological validity. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 2782–2800. [Google Scholar] [CrossRef] [PubMed]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
- Qu, F.; Wang, S.; Yan, W.; Li, H.; Wu, S.; Fu, X. CAS(ME)2: A database for spontaneous macro-expression and micro-expression spotting and recognition. IEEE Trans. Affect. Comput. 2018, 9, 424–436. [Google Scholar] [CrossRef] [Scilit]
- Yap, C.H.; Kendrick, C.; Yap, M.H. SAMM long videos: A spontaneous facial micro- and macro-expressions dataset. In Proceedings of the 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020); IEEE: New York, NY, USA, 2020; pp. 771–776. [Google Scholar]
- Liong, G.B.; Liong, S.; See, J.; Chee-Seng, C. MTSN: A multi-temporal stream network for spotting facial macro-and micro-expression with hard and soft pseudo-labels. In Proceedings of the 2nd Workshop on Facial Micro-Expression: Advanced Techniques for Multi-Modal Facial Expression Analysis; Association for Computing Machinery: New York, NY, USA, 2022; pp. 3–10. [Google Scholar]
- Zhao, Y.; Tong, X.; Zhu, Z.; Sheng, J.; Dai, L.; Xu, L.; Jiang, Y.; Li, J. Rethinking optical flow methods for micro-expression spotting. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022. [Google Scholar]



| Method | Recall | Precision | F1-Score |
|---|---|---|---|
| Yang et al. [27] | 0.230 | 0.321 | 0.268 |
| Ours | 0.304 | 0.351 | 0.326 |
| Method | Recall | Precision | F1-Score |
|---|---|---|---|
| Yang et al. [27] | 0.250 | 0.080 | 0.121 |
| Ours | 0.861 | 0.083 | 0.151 |
| Method | Recall | Precision | F1-score |
|---|---|---|---|
| MTSN (2022) [40] | 0.342 | 0.385 | 0.362 |
| Zhao et al. (2022) [41] | - | - | 0.403 |
| LGSNet (2023) [15] | 0.367 | 0.630 | 0.464 |
| Ours | 0.405 | 0.563 | 0.471 |
| Method | Recall | Precision | F1-score |
|---|---|---|---|
| MTSN (2022) [40] | 0.260 | 0.319 | 0.287 |
| Zhao et al. (2022) [41] | - | - | 0.386 |
| LGSNet (2023) [15] | 0.355 | 0.429 | 0.388 |
| Ours | 0.450 | 0.354 | 0.396 |
| Method | Recall | Precision | F1-Score |
|---|---|---|---|
| Without depth information | 0.259 | 0.168 | 0.175 |
| With depth information | 0.304 | 0.351 | 0.326 |
| Method | Recall | Precision | F1-Score |
|---|---|---|---|
| Without depth information | 0.277 | 0.006 | 0.087 |
| With depth information | 0.861 | 0.083 | 0.151 |
| Loss Function | Recall | Precision | F1-Score |
|---|---|---|---|
| Cross-Entropy Loss | 0.104 | 0.300 | 0.287 |
| Focal Loss | 0.304 | 0.351 | 0.326 |
| Loss Function | Recall | Precision | F1-Score |
|---|---|---|---|
| Cross-Entropy Loss | 0.651 | 0.035 | 0.111 |
| Focal Loss | 0.861 | 0.083 | 0.151 |
| Window Size | Recall | Precision | F1-Score |
|---|---|---|---|
| 3 frames | 0.197 | 0.314 | 0.279 |
| 5 frames | 0.304 | 0.351 | 0.326 |
| 8 frames | 0.270 | 0.293 | 0.306 |
| Window Size | Recall | Precision | F1-Score |
|---|---|---|---|
| 3 frames | 0.357 | 0.069 | 0.139 |
| 5 frames | 0.861 | 0.083 | 0.151 |
| 8 frames | 0.755 | 0.063 | 0.146 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Feng, K.; Jiang, F.; Huang, S.; Li, M. Head-Movement-Robust Micro-Expression Detection Method via 3D Motion Correction and Transformers. Electronics 2026, 15, 1836. https://doi.org/10.3390/electronics15091836
Feng K, Jiang F, Huang S, Li M. Head-Movement-Robust Micro-Expression Detection Method via 3D Motion Correction and Transformers. Electronics. 2026; 15(9):1836. https://doi.org/10.3390/electronics15091836
Chicago/Turabian StyleFeng, Keyi, Fake Jiang, Shucheng Huang, and Mingxing Li. 2026. "Head-Movement-Robust Micro-Expression Detection Method via 3D Motion Correction and Transformers" Electronics 15, no. 9: 1836. https://doi.org/10.3390/electronics15091836
APA StyleFeng, K., Jiang, F., Huang, S., & Li, M. (2026). Head-Movement-Robust Micro-Expression Detection Method via 3D Motion Correction and Transformers. Electronics, 15(9), 1836. https://doi.org/10.3390/electronics15091836

