TDA-Phys: Temporal Difference Adaptation of Video Foundation Model for Remote Photoplethysmography
Featured Application
Abstract
1. Introduction
- We show that a general video foundation model (VideoMAE v2) can be effectively adapted to rPPG signal extraction with only a lightweight adapter and no modification to its pretrained backbone. To the best of our knowledge, this is the first successful application of a generic video foundation model to end-to-end rPPG regression.
- We introduce a Temporal Difference Convolutional Adapter. By freezing the backbone parameters of the VideoMAE v2 foundation model and fine-tuning only the parameters within this adapter, our method effectively adapts to the downstream task of rPPG signal extraction.
- To address the limitation that VideoMAE v2 only processes short 16-frame input sequences, we employ a sliding window strategy. This strategy divides long video sequences into overlapping short clips for segment-wise inference, followed by a temporal aggregation mechanism to reconstruct continuous and stable rPPG signals.
- We conduct comprehensive validation experiments on the COHFACE, UBFC-rPPG, BUAA-MIHR, and VIPL-HR datasets. The results demonstrate that the proposed fine-tuning architecture for video foundation models can effectively adapt to the rPPG signal extraction task and exhibits promising generalization ability. This work robustly validates the potential of large-scale video foundation models for physiological signal perception tasks.
2. Related Work
2.1. Deep Learning-Based rPPG Methods
2.2. Foundation Models and Adapter-Based Transfer Learning
3. Materials and Methods
3.1. Overall Network Framework
3.2. Sliding Window Strategy
| Algorithm 1. Pseudocode for temporal aggregation |
| rPPG_full = zeros(160) # Accumulated signal weight_count = zeros(160) # Coverage counter for i in range(19): # i = 0 to 18 start = i * 8 # Start index in original sequence end = start + 16 # End index (exclusive) pred = model_output [i] # Predicted rPPG segment (length 16) rPPG_full [start:end] += pred weight_count [start:end] += 1.0 rPPG_final = rPPG_full/weight_count # Element-wise division |
3.3. Temporal Difference Convolutional Adapter
4. Results and Discussion
4.1. Model Performance Evaluation Experiment
4.2. Ablation Study
4.2.1. Impact of the Sliding Window Strategy
4.2.2. Impact of the Temporal Difference Convolutional Adapter
5. Conclusions
- Introducing task-aware adapter designs. For example, embedding illumination-invariant modules or motion compensation modules into shallow ViT Blocks to suppress major interference at an early stage.
- Exploring multi-scale spatiotemporal adaptation mechanisms. Integrating local and global contexts through multi-branch adapters to separately model high-frequency noise and low-frequency physiological signals.
- Constructing an rPPG dataset with controlled motion and lighting variations to conduct targeted pre-training for the adapter or the entire model, thereby enhancing its generalization capability in real-world complex scenarios.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Xiao, H.; Liu, T.; Sun, Y.; Li, Y.; Zhao, S.; Avolio, A. Remote photoplethysmography for heart rate measurement: A review. Biomed. Signal Process. Control 2024, 88, 105608. [Google Scholar] [CrossRef]
- Yu, Z.; Li, X.; Zhao, G. Remote photoplethysmograph signal measurement from facial videos using spatio-temporal networks. arXiv 2019, arXiv:1905.02419. [Google Scholar] [CrossRef]
- Yu, Z.; Shen, Y.; Shi, J.; Zhao, H.; Torr, P.; Zhao, G. Physformer: Facial video-based physiological measurement with temporal difference transformer. In Proceedings of the IEEE/CVF Conference On Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 4176–4186. [Google Scholar]
- Chen, W.; McDuff, D. Deepphys: Video-based physiological measurement using convolutional attention networks. arXiv 2018, arXiv:1805.07888. [Google Scholar]
- Zhai, D.; Chen, W.; Ding, Y.; Yu, M.; Li, Q.; Wu, H. Research on Robust Measurement Method of Heart Rate Using Remote Photoplethysmography Based on Adversarial Learning Network with High and Low Frequency Features. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 5208–5222. [Google Scholar] [CrossRef]
- Li, J.; Yu, Z.; Shi, J. Learning motion-robust remote photoplethysmography through arbitrary resolution videos. arXiv 2023, arXiv:2211.16922. [Google Scholar] [CrossRef]
- Wang, L.; Huang, B.; Zhao, Z.; Tong, Z.; He, Y.; Wang, Y.; Wang, Y.; Qiao, Y. Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 14549–14560. [Google Scholar]
- Sun, Y.; Hu, S.; Azorin-Peris, V.; Greenwald, S.; Chambers, J.; Zhu, Y. Motion-compensated noncontact imaging photoplethysmography to monitor cardiorespiratory status during exercise. J. Biomed. Opt. 2011, 16, 077010–077019. [Google Scholar] [CrossRef]
- Guo, Z.; Wang, Z.J.; Shen, Z. Physiological parameter monitoring of drivers based on video data and independent vector analysis. In Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy, 4–9 May 2014; IEEE: Piscataway, NJ, USA, 2014; pp. 4374–4378. [Google Scholar]
- Qi, H.; Guo, Z.; Chen, X.; Shen, Z.; Wang, Z.J. Video-based human heart rate measurement using joint blind source separation. Biomed. Signal Process. Control 2017, 31, 309–320. [Google Scholar] [CrossRef]
- De Haan, G.; Van Leest, A. Improved motion robustness of remote-PPG by using the blood volume pulse signature. Physiol. Meas. 2014, 35, 1913. [Google Scholar] [CrossRef]
- Li, X.; Chen, J.; Zhao, G.; Pietikainen, M. Remote heart rate measurement from face videos under realistic situations. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; IEEE: Piscataway, NJ, USA, 2014; pp. 4264–4271. [Google Scholar]
- Wang, W.J.; Stuijk, S.; De Haan, G. A Novel Algorithm for Remote Photoplethysmography: Spatial Subspace Rotation. IEEE Trans. Biomed. Eng. 2016, 63, 1974–1984. [Google Scholar] [CrossRef]
- Wang, W.; Den Brinker, A.C.; Stuijk, S.; De Haan, G. Algorithmic principles of remote PPG. IEEE Trans. Biomed. Eng. 2016, 64, 1479–1491. [Google Scholar] [CrossRef]
- Špetlík, R.; Franc, V.; Cech, J.; Matas, J. Visual heart rate estimation with convolutional neural network. In Proceedings of the British Machine Vision Conference, Newcastle, UK, 3–6 September 2018; BMVA Press: Durham, UK, 2018. [Google Scholar]
- Liu, X.; Fromm, J.; Patel, S.; McDuff, D. Multi-task temporal shift attention networks for on-device contactless vitals measurement. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Online Conference, 6–12 December 2020; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 19400–19411. [Google Scholar]
- Niu, X.; Han, H.; Shan, S.; Chen, X. Synrhythm: Learning a deep heart rate estimator from general to specific. In Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR), Beijing, China, 20–24 August 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 3580–3585. [Google Scholar]
- Song, R.; Zhang, S.; Li, C.; Zhang, Y.; Cheng, J.; Chen, X. Heart rate estimation from facial videos using a spatiotemporal representation with convolutional neural networks. IEEE Trans. Instrum. Meas. 2020, 69, 7411–7421. [Google Scholar] [CrossRef]
- Bousefsaf, F.; Pruski, A.; Maaoui, C. 3D convolutional neural networks for remote pulse rate measurement and mapping from facial video. Appl. Sci. 2019, 9, 4364. [Google Scholar] [CrossRef]
- Yu, Z.; Li, X.; Wang, P.; Zhao, G. Transrppg: Remote photoplethysmography transformer for 3d mask face presentation attack detection. IEEE Signal Process. Lett. 2021, 28, 1290–1294. [Google Scholar] [CrossRef]
- Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollar, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 16000–16009. [Google Scholar]
- Tong, Z.; Song, Y.; Wang, J.; Wang, L. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Curran Associates, Inc.: Red Hook, NY, USA, 2022; pp. 10078–10093. [Google Scholar]
- Han, Z.; Gao, C.; Liu, J.; Zhang, J. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv 2024, arXiv:2403.14608. [Google Scholar]
- Yu, Y.; Xu, C.; Wang, K. Ts-sam: Fine-tuning segment-anything model for downstream tasks. In Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME), Niagara Falls, ON, Canada, 15–19 July 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar]
- Heusch, G.; Anjos, A.; Marcel, S. A reproducible study on remote heart rate measurement. arXiv 2017, arXiv:1709.00962. [Google Scholar] [CrossRef]
- Bobbia, S.; Macwan, R.; Benezeth, Y.; Mansouri, A.; Dubois, J. Unsupervised skin tissue segmentation for remote photoplethysmography. Pattern Recognit. Lett. 2019, 124, 82–90. [Google Scholar] [CrossRef]
- Xi, L.; Chen, W.; Zhao, C.; Wu, X.; Wang, J. Image enhancement for remote photoplethysmography in a low-light environment. In Proceedings of the 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), Buenos Aires, Argentina, 16–20 November 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 1–7. [Google Scholar]
- Niu, X.; Han, H.; Shan, S.; Chen, X. VIPL-HR: A multi-modal database for pulse estimation from less-constrained face video. In Computer Vision—ACCV 2018; Springer: Cham, Switzerland, 2018; pp. 562–576. [Google Scholar]
- De Haan, G.; Jeanne, V. Robust pulse rate from chrominance-based rPPG. IEEE Trans. Biomed. Eng. 2013, 60, 2878–2886. [Google Scholar] [CrossRef]
- Mehta, A.D.; Sharma, H. CPulse: Heart rate estimation from RGB videos under realistic conditions. IEEE Trans. Instrum. Meas. 2023, 72, 5023312. [Google Scholar] [CrossRef]
- Lokendra, B.; Puneet, G. AND-rPPG: A novel denoising-rPPG network for improving remote heart rate estimation. Comput. Biol. Med. 2022, 141, 105146. [Google Scholar] [CrossRef]
- Joshi, J.; Cho, Y. Imaging blood volume pulse dataset: RGB-thermal remote photoplethysmography dataset with high-resolution signal-quality labels. Electronics 2024, 13, 1334. [Google Scholar] [CrossRef]
- Yu, Z.; Li, X.; Niu, X.; Shi, J.; Zhao, G. Autohr: A strong end-to-end baseline for remote heart rate measurement with neural searching. IEEE Signal Process Lett. 2020, 27, 1245–1249. [Google Scholar] [CrossRef]
- Niu, X.; Shan, S.; Han, H.; Chen, X. Rhythmnet: End-to-end heart rate estimation from face via spatial-temporal representation. IEEE Trans. Image Process. 2019, 29, 2409–2423. [Google Scholar] [CrossRef] [PubMed]







| Dataset | Method | r | MAE | RMSE |
|---|---|---|---|---|
| COHFACE | ICA [9] | - | 8.89 | 14.55 |
| CHROM [30] | - | 7.80 | 12.45 | |
| POS [14] | - | 13.43 | 17.05 | |
| HR-CNN [15] | 0.52 | 8.10 | 10.80 | |
| PhysNet [2] | 0.54 | 8.63 | 9.36 | |
| DeepPhys [4] | 0.62 | 6.60 | 10.79 | |
| Cpulse [31] | 0.99 | 1.99 | 3.83 | |
| PhysFormer [3] | 0.99 | 2.00 | 2.00 | |
| TDA-Phys(ours) | 0.99 | 0.90 ± 0.03 * | 1.22 ± 0.02 * | |
| UBFC-rPPG | ICA [9] | - | 6.43 | 11.43 |
| CHROM [30] | - | 2.98 | 3.80 | |
| POS [14] | - | 3.99 | 6.81 | |
| HR-CNN [15] | 0.64 | 4.90 | 5.89 | |
| DeepPhys [4] | 0.65 | 2.90 | 3.63 | |
| PhysNet [2] | 0.97 | 2.38 | 3.19 | |
| AND-rPPG [32] | 0.99 | 2.67 | 4.07 | |
| PhysFormer [3] | 0.99 | 2.70 | 4.38 | |
| TDA-Phys(ours) | 0.99 | 1.55 ± 0.02 * | 3.35 ± 0.05 * |
| Dataset | Method | r | MAE | RMSE |
|---|---|---|---|---|
| BUAA-MIHR | ICA [9] | - | 11.27 | 16.68 |
| CHROM [30] | - | 12.99 | 17.80 | |
| POS [14] | - | 11.75 | 17.31 | |
| Green [8] | - | 12.76 | 17.89 | |
| iBVPNet [33] | 0.34 | 10.18 | 12.50 | |
| PhysNet [2] | 0.24 | 15.23 | 18.98 | |
| TDA-Phys(ours) | 0.62 | 6.68 | 12.75 | |
| PhysFormer [3] | 0.84 | 4.48 | 8.03 | |
| HLFF [5] | 0.89 | 3.19 | 6.50 | |
| VIPL-HR | POS [14] | - | 11.50 | 17.20 |
| DeepPhys [4] | 0.11 | 11.00 | 13.80 | |
| PhysNet [2] | 0.20 | 10.80 | 14.80 | |
| TDA-Phys(ours) | 0.54 | 8.23 | 16.02 | |
| AutoHR [34] | 0.71 | 5.68 | 8.68 | |
| RhythmNet [35] | 0.74 | 5.30 | 8.14 | |
| PhysFormer [3] | 0.78 | 4.97 | 7.79 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, W.; Ding, Y.; Bu, K.; Yu, M.; Wu, H. TDA-Phys: Temporal Difference Adaptation of Video Foundation Model for Remote Photoplethysmography. Appl. Sci. 2026, 16, 2038. https://doi.org/10.3390/app16042038
Chen W, Ding Y, Bu K, Yu M, Wu H. TDA-Phys: Temporal Difference Adaptation of Video Foundation Model for Remote Photoplethysmography. Applied Sciences. 2026; 16(4):2038. https://doi.org/10.3390/app16042038
Chicago/Turabian StyleChen, Wei, Yinghao Ding, Kunze Bu, Ming Yu, and Hang Wu. 2026. "TDA-Phys: Temporal Difference Adaptation of Video Foundation Model for Remote Photoplethysmography" Applied Sciences 16, no. 4: 2038. https://doi.org/10.3390/app16042038
APA StyleChen, W., Ding, Y., Bu, K., Yu, M., & Wu, H. (2026). TDA-Phys: Temporal Difference Adaptation of Video Foundation Model for Remote Photoplethysmography. Applied Sciences, 16(4), 2038. https://doi.org/10.3390/app16042038

