A Dual-Branch Spatio-Temporal Feature Differencing Method for Robust rPPG Estimation
Abstract
1. Introduction
2. Related Works
2.1. Video Color Magnification
2.2. VS-Net
3. Methodology
3.1. Problem Statement
3.2. Foreground–Background Dual-Branch Differencing Pipeline
| Algorithm 1: Dual-Branch Foreground–Background Differencing |
![]() |
3.3. Learn-Based Video Color Magnification
3.4. Temporal Shift Module (TSM) Injection for Spatio-Temporal Feature Extraction
4. Experimental Results
4.1. Experimental Setting
4.2. Results
4.3. Ablation Study
5. Conclusions
6. Future Work
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Verkruysse, W.; Svaasand, L.O.; Nelson, J.S. Remote plethysmographic imaging using ambient light. Opt. Express 2008, 16, 21434–21445. [Google Scholar] [CrossRef] [PubMed]
- Poh, M.Z.; McDuff, D.J.; Picard, R.W. Non-contact, automated cardiac pulse measurements using video imaging and blind source separation. Opt. Express 2010, 18, 10762–10774. [Google Scholar] [CrossRef] [PubMed]
- Nowara, E.M.; Marks, T.K.; Mansour, H.; Veeraraghavan, A. Near-Infrared Imaging Photoplethysmography During Driving. IEEE Trans. Intell. Transp. Syst. 2020, 23, 21320–21330. [Google Scholar] [CrossRef]
- Nowara, E.M.; Marks, T.K.; Mansour, H.; Veeraraghavan, A. SparsePPG: Towards Driver Monitoring Using Camera-Based Vital Signs Estimation in Near-Infrared. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, QC, Canada, 10–17 October 2021. [Google Scholar]
- De Haan, G.; Jeanne, V. Robust pulse rate from chrominance-based rPPG. IEEE Trans. Biomed. Eng. 2013, 60, 2878–2886. [Google Scholar] [CrossRef] [PubMed]
- Wang, W.; den Brinker, A.C.; Stuijk, S.; de Haan, G. Algorithmic Principles of Remote PPG. IEEE Trans. Biomed. Eng. 2016, 64, 1479–1491. [Google Scholar] [CrossRef] [PubMed]
- Wu, H.Y.; Rubinstein, M.; Shih, E.; Guttag, J.; Durand, F.; Freeman, W. Eulerian Video Magnification for Revealing Subtle Changes in the World. ACM Trans. Graph. 2012, 31, 1–8. [Google Scholar] [CrossRef]
- Oh, T.H.; Jaroensri, R.; Kim, C.; Elgharib, M.; Durand, F.; Freeman, W.T.; Matusik, W. Learning-based Video Motion Magnification. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 633–648. [Google Scholar]
- Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; Paluri, M. Learning Spatiotemporal Features with 3D Convolutional Networks (C3D). In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 4489–4497. [Google Scholar]
- Carreira, J.; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset (I3D). In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6299–6308. [Google Scholar]
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. SlowFast Networks for Video Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6202–6211. [Google Scholar]
- Niu, X.; Han, H.; Shan, S.; Chen, X. RhythmNet: End-to-End Heart Rate Estimation from Face via Spatial-Temporal Representation. IEEE Trans. Image Process. 2019, 29, 2409–2423. [Google Scholar] [CrossRef] [PubMed]
- Sun, N.; He, P.; Liu, J.; Chai, L.; Wu, C.; Liu, X. Remote heart rate measurement based on video color magnification and spatiotemporal self-attention. Biomed. Signal Process. Control 2025, 106, 107677. [Google Scholar] [CrossRef]
- Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; Hu, H. Video Swin Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; pp. 3202–3211. [Google Scholar]
- Kang, J.; Yang, S.; Zhang, W. TransPPG: Two-stream Transformer for Remote Heart Rate Estimate. CCF Trans. Pervasive Comput. Interact. 2024, 6, 271–280. [Google Scholar] [CrossRef]
- Lin, J.; Gan, C.; Han, S. TSM: Temporal Shift Module for Efficient Video Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7083–7093. [Google Scholar]
- Liu, X.; Fromm, J.; Patel, S.; McDuff, D. Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals Measurement. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 6–12 December 2020. [Google Scholar]
- Bobbia, S.; Macwan, R.; Benezeth, Y.; Mansouri, A.; Dubois, J. Unsupervised skin tissue segmentation for remote photoplethysmography. Pattern Recognit. Lett. 2019, 124, 82–90. [Google Scholar] [CrossRef]
- Qiu, Y.; Liu, Y.; Arteaga-Falconi, J.; Ghasemzadeh, H. EVM-CNN: Real-time contactless heart rate estimation from facial video. IEEE Trans. Multimed. 2018, 21, 1778–1787. [Google Scholar] [CrossRef]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- Gideon, J.; Stent, S. The Way to My Heart is Through Contrastive Learning: Remote Photoplethysmography from Unlabelled Video. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 3995–4004. [Google Scholar]
- Li, Z.; Li, B.; Yin, L. Contactless Pulse Estimation Leveraging Pseudo Labels and Self-Supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 20588–20597. [Google Scholar]
- Zhang, K.; Zhang, Z.; Li, Z.; Qiao, Y. Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks (MTCNN). IEEE Signal Process. Lett. 2016, 23, 1499–1503. [Google Scholar] [CrossRef]
- Chen, W.; McDuff, D. DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 349–365. [Google Scholar]
- Liu, Y.; Wang, C.; Lu, M.; Yang, J.; Gui, J.; Zhang, S. From Simple to Complex Scenes: Learning Robust Feature Representations for Accurate Human Parsing. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5449–5462. [Google Scholar] [CrossRef]
- Sun, G.; Wang, C.; Hua, Y. Spatio-temporal Prompting Network for Robust Video Feature Extraction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023. [Google Scholar]
- Wang, C.; Zhang, Q.; Wang, X.; Zhou, L.; Li, Q.; Xia, Z.; Ma, B.; Shi, Y. Light-Field Image Multiple Reversible Robust Watermarking Against Geometric Attacks. IEEE Trans. Dependable Secur. Comput. 2025, 22, 5861–5875. [Google Scholar] [CrossRef]
- Jiang, Y.; Chiang, H.D. Geometry-Aware Weight Perturbation for Adversarial Training. Electronics 2024, 13, 3508. [Google Scholar] [CrossRef]




| Feature | UBFC-rPPG [18] | MR-NIRP [3,4] |
|---|---|---|
| Environment | Indoor (Controlled) | In-vehicle (Real-world) |
| Subjects | 42 peoples (30 Male, 12 Female) | 19 peoples (17 Male, 2 Female) |
| Resolution | pixels | pixels |
| Frame Rate | 30 fps | 30 fps |
| Conditions | Low-noise studio lighting | Stationary & Driving |
| Noise Factors | Minimal | Natural light, Tunnels, Vibration |
| Parameter | Value |
|---|---|
| Optimizer | Adam |
| Learning Rate (LR) | |
| Batch Size | 4 |
| Num Epochs | 80 |
| VCM Amplification () | 80 |
| TSM Shift Ratio () | 1/8 |
| Loss Function | MSE (UBFC-rPPG), Neg. Pearson (MR-NIRP) |
| Metric | DeepPhys [24] | RhythmNet [12] | VS-Net (Baseline) [13] | Ours |
|---|---|---|---|---|
| MAE (bpm) ↓ | 4.71 | 4.52 | 3.18 | 2.68 |
| RMSE (bpm) ↓ | 6.26 | 6.08 | 4.12 | 3.62 |
| R↑ | 0.693 | 0.680 | 0.818 | 0.845 |
| Metric | DeepPhys [24] | RhythmNet [12] | VS-Net (Baseline) [13] | Ours |
|---|---|---|---|---|
| MAE (bpm) ↓ | 3.18 | 2.87 | 1.55 | 1.78 |
| RMSE (bpm) ↓ | 4.77 | 4.61 | 2.68 | 2.89 |
| R↑ | 0.721 | 0.834 | 0.911 | 0.893 |
| Metric | DeepPhys [24] | RhythmNet [12] | VS-Net (Baseline) [13] | Ours |
|---|---|---|---|---|
| MAE (bpm) ↓ | 13.26 | 11.53 | 4.35 | 3.75 |
| RMSE (bpm) ↓ | 17.92 | 15.03 | 8.12 | 7.02 |
| R↑ | 0.432 | 0.471 | 0.724 | 0.803 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 2.83 | 2.37 | 2.81 | 2.45 |
| RMSE (bpm) ↓ | 3.76 | 3.12 | 4.05 | 3.38 |
| R↑ | 0.851 | 0.912 | 0.869 | 0.895 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 2.36 | 2.51 | 2.45 | 2.79 |
| RMSE (bpm) ↓ | 3.41 | 3.68 | 3.59 | 3.95 |
| R↑ | 0.882 | 0.865 | 0.874 | 0.841 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 3.18 | 2.68 | 2.99 | 2.74 |
| RMSE (bpm) ↓ | 4.12 | 3.62 | 4.07 | 3.83 |
| R↑ | 0.818 | 0.845 | 0.825 | 0.838 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 0.82 | 1.47 | 1.05 | 1.28 |
| RMSE (bpm) ↓ | 1.12 | 1.98 | 1.45 | 1.76 |
| R↑ | 0.998 | 0.990 | 0.995 | 0.993 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 1.55 | 1.78 | 1.62 | 1.70 |
| RMSE (bpm) ↓ | 2.68 | 2.89 | 2.75 | 2.83 |
| R↑ | 0.911 | 0.893 | 0.905 | 0.901 |
| Metric | VS-Net | Ours | Ours-Single Branch | Ours-TSM Module Off |
|---|---|---|---|---|
| MAE (bpm) ↓ | 4.35 | 3.75 | 4.18 | 3.95 |
| RMSE (bpm) ↓ | 8.12 | 7.02 | 7.86 | 7.41 |
| R↑ | 0.724 | 0.803 | 0.751 | 0.782 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Cho, G.; Kim, M.-J.; Ahn, C.W. A Dual-Branch Spatio-Temporal Feature Differencing Method for Robust rPPG Estimation. Mathematics 2025, 13, 3830. https://doi.org/10.3390/math13233830
Cho G, Kim M-J, Ahn CW. A Dual-Branch Spatio-Temporal Feature Differencing Method for Robust rPPG Estimation. Mathematics. 2025; 13(23):3830. https://doi.org/10.3390/math13233830
Chicago/Turabian StyleCho, Gyumin, Man-Je Kim, and Chang Wook Ahn. 2025. "A Dual-Branch Spatio-Temporal Feature Differencing Method for Robust rPPG Estimation" Mathematics 13, no. 23: 3830. https://doi.org/10.3390/math13233830
APA StyleCho, G., Kim, M.-J., & Ahn, C. W. (2025). A Dual-Branch Spatio-Temporal Feature Differencing Method for Robust rPPG Estimation. Mathematics, 13(23), 3830. https://doi.org/10.3390/math13233830


