A Clip-Based Dairy Cow Behavior Recognition Method Integrating Temporal Modeling and Behavioral Priors
Simple Summary
Abstract
1. Introduction
2. Related Work
2.1. Dairy Cow and Video Behavior Recognition
2.2. CLIP-Based Video Adaptation and Parameter-Efficient Fine-Tuning
2.3. Prior Knowledge Constraints in Behavior Recognition
3. Materials and Methods
3.1. Baseline Model
3.2. Proposed Model
3.2.1. Temporal Modeling Modules
3.2.2. Posture–Action Dual-Head Modeling and Behavioral Prior Loss
3.2.3. Training Objective and Inference Mapping
3.3. Dataset and Data Reconstruction
3.4. Experimental Settings and Evaluation Metrics
4. Results
4.1. Comparison with Existing Video Recognition Models
4.2. Ablation Experiments on Main Components
4.3. Analysis of the Behavioral Prior Loss Weight
4.4. Per-Class Five-Class Performance
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Lamanna, M.; Cavallini, D. Climate narratives and global food security. Vet. Rec. 2026, 198, 319–320. [Google Scholar] [CrossRef] [PubMed]
- Berckmans, D. General introduction to precision livestock farming. Anim. Front. 2017, 7, 6–11. [Google Scholar] [CrossRef]
- Wathes, C.M.; Kristensen, H.H.; Aerts, J.M.; Berckmans, D. Is precision livestock farming an engineer’s daydream or nightmare, an animal’s friend or foe, and a farmer’s panacea or pitfall? Comput. Electron. Agric. 2008, 64, 2–10. [Google Scholar] [CrossRef]
- Neethirajan, S. The role of sensors, big data and machine learning in modern animal farming. Sens. Bio-Sens. Res. 2020, 29, 100367. [Google Scholar] [CrossRef]
- Wu, X.; Dong, J.; Bao, W.; Zou, B.; Wang, L.; Wang, H. Augmented intelligence of things for emergency vehicle secure trajectory prediction and task offloading. IEEE Internet Things J. 2024, 11, 36030–36043. [Google Scholar] [CrossRef]
- Borchers, M.R.; Chang, Y.M.; Proudfoot, K.L.; Wadsworth, B.A.; Stone, A.E.; Bewley, J.M. Machine-learning-based calving prediction from activity, lying, and ruminating behaviors in dairy cattle. J. Dairy Sci. 2017, 100, 5664–5674. [Google Scholar] [CrossRef] [PubMed]
- Rushen, J.; Chapinal, N.; de Passille, A. Automated monitoring of behavioural-based animal welfare indicators. Anim. Welf. 2012, 21, 339–350. [Google Scholar] [CrossRef]
- Schillings, J.; Bennett, R.; Rose, D.C. Exploring the potential of precision livestock farming technologies to help address farm animal welfare. Front. Anim. Sci. 2021, 2, 639678. [Google Scholar] [CrossRef]
- Buller, H.; Blokhuis, H.; Jensen, P.; Keeling, L. Towards farm animal welfare and sustainability. Animals 2018, 8, 81. [Google Scholar] [CrossRef] [PubMed]
- Garcia, R.; Aguilar, J.; Toro, M.; Pinto, A.; Rodriguez, P. A systematic literature review on the use of machine learning in precision livestock farming. Comput. Electron. Agric. 2020, 179, 105826. [Google Scholar] [CrossRef]
- Hlimi, A.; El Otmani, S.; Elame, F.; Chentouf, M.; El Halimi, R.; Chebli, Y. Application of precision technologies to characterize animal behavior: A review. Animals 2024, 14, 416. [Google Scholar] [CrossRef] [PubMed]
- Shen, W.; Cheng, F.; Zhang, Y.; Wei, X.; Fu, Q.; Zhang, Y. Automatic recognition of ingestive-related behaviors of dairy cows based on triaxial acceleration. Inf. Process. Agric. 2020, 7, 427–443. [Google Scholar] [CrossRef]
- Tian, F.; Wang, J.; Xiong, B.; Jiang, L.; Song, Z.; Li, F. Real-time behavioral recognition in dairy cows based on geomagnetism and acceleration information. IEEE Access 2021, 9, 109497–109509. [Google Scholar] [CrossRef]
- Liu, M.; Wu, Y.; Li, G.; Liu, M.; Hu, R.; Zou, H.; Wang, Z.; Peng, Y. Classification of cow behavior patterns using inertial measurement units and a fully convolutional network model. J. Dairy Sci. 2023, 106, 1351–1359. [Google Scholar] [CrossRef] [PubMed]
- Wu, D.; Wang, Y.; Han, M.; Song, L.; Shang, Y.; Zhang, X.; Song, H. Using a CNN-LSTM for basic behaviors detection of a single dairy cow in a complex environment. Comput. Electron. Agric. 2021, 182, 106016. [Google Scholar] [CrossRef]
- Nasirahmadi, A.; Edwards, S.A.; Sturm, B. Implementation of machine vision for detecting behaviour of cattle and pigs. Livest. Sci. 2017, 202, 25–38. [Google Scholar] [CrossRef]
- Chen, C.; Zhu, W.; Norton, T. Behaviour recognition of pigs and cattle: Journey from computer vision to deep learning. Comput. Electron. Agric. 2021, 187, 106255. [Google Scholar] [CrossRef]
- Guo, Y.; Zhang, Z.; He, D.; Niu, J.; Tan, Y. Detection of cow mounting behavior using region geometry and optical flow characteristics. Comput. Electron. Agric. 2019, 163, 104828. [Google Scholar] [CrossRef]
- Carreira, J.; Zisserman, A. Quo Vadis, action recognition? A new model and the Kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 6299–6308. [Google Scholar]
- Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; Paluri, M. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 6450–6459. [Google Scholar]
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. SlowFast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 6202–6211. [Google Scholar]
- Tong, Z.; Song, Y.; Wang, J.; Wang, L. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. arXiv 2022, arXiv:2203.12602. [Google Scholar]
- Bertasius, G.; Wang, H.; Torresani, L. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning, Online, 18–24 July 2021; pp. 813–824. [Google Scholar]
- Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; Hu, H. Video Swin Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 3202–3211. [Google Scholar]
- Li, K.; Fan, D.; Wu, H.; Zhao, A. A new dataset for video-based cow behavior recognition. Sci. Rep. 2024, 14, 18702. [Google Scholar] [CrossRef] [PubMed]
- Bai, Q.; Gao, R.; Wang, R.; Li, Q.; Yu, Q.; Zhao, C.; Li, S. X3DFast model for classifying dairy cow behaviors based on a two-pathway architecture. Sci. Rep. 2023, 13, 20519. [Google Scholar] [CrossRef] [PubMed]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, Online, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Ni, B.; Peng, H.; Chen, M.; Zhang, S.; Meng, G.; Fu, J.; Xiang, S.; Ling, H. Expanding language-image pretrained models for general video recognition. In Proceedings of the European Conference on Computer Vision; Springer Nature: Cham, Switzerland, 2022; pp. 1–18. [Google Scholar]
- Wang, M.; Xing, J.; Liu, Y. ActionCLIP: A new paradigm for video action recognition. arXiv 2021, arXiv:2109.08472. [Google Scholar]
- Rasheed, H.; Khattak, M.U.; Maaz, M.; Khan, S.; Khan, F. Fine-tuned CLIP models are efficient video learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 6545–6554. [Google Scholar]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; pp. 2790–2799. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Chen, W. LoRA: Low-rank adaptation of large language models. Int. Conf. Learn. Represent. 2022, 1, 3. [Google Scholar]
- Serafini, L.; Garcez, A.d. Logic tensor networks: Deep learning and logical reasoning from data and knowledge. arXiv 2016, arXiv:1606.04422. [Google Scholar]
- Donadello, I.; Serafini, L.; Garcez, A.d. Logic tensor networks for semantic image interpretation. In Proceedings of the International Joint Conference on Artificial Intelligence; Association for Computing Machinery: New York, NY, USA, 2017; pp. 1596–1602. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. arXiv 2017, arXiv:1711.05101. [Google Scholar]
- Zhang, Y.; Zhang, Y.; Jiang, H.; Du, H.; Xue, A.; Shen, W. New method for modeling digital twin behavior perception of cows: Cow daily behavior recognition based on multimodal data. Comput. Electron. Agric. 2024, 226, 109426. [Google Scholar] [CrossRef]
- Wang, H.; Zhang, X.; Xia, Y.; Wu, X. A differential privacy-preserving deep learning caching framework for heterogeneous communication network systems. Int. J. Intell. Syst. 2022, 37, 11142–11166. [Google Scholar] [CrossRef]








| Behavior Annotation | Posture Branch | Action Branch | Five-Class Label |
|---|---|---|---|
| Standing only | 0: standing | 0: none | standing |
| Lying only | 1: lying | 0: none | lying |
| Standing + feeding | 0: standing | 1: feeding | feeding |
| Standing + drinking | 0: standing | 2: drinking | drinking |
| Standing + rumination | 0: standing | 3: rumination | rumination |
| Lying + rumination | 1: lying | 3: rumination | rumination |
| Subset | Standing | Lying | Feeding | Drinking | Rumination | Total |
|---|---|---|---|---|---|---|
| Training set | 1847 | 1030 | 1784 | 286 | 2444 | 7391 |
| Validation set | 565 | 244 | 501 | 108 | 806 | 2224 |
| Test set | 465 | 332 | 449 | 62 | 761 | 2069 |
| Total | 2877 | 1606 | 2734 | 456 | 4011 | 11,684 |
| Method | 5-Class Acc | Macro Precision | Macro Recall | Macro-F1 | Total | Trainable |
|---|---|---|---|---|---|---|
| R(2+1)D-R34 | 71.15% | 0.6681 | 0.6803 | 0.6711 | 63.557 M | 63.557 M |
| Video Swin-T | 71.34% | 0.6942 | 0.6651 | 0.6732 | 27.854 M | 27.854 M |
| SlowFast-R50 | 73.76% | 0.7043 | 0.7470 | 0.7150 | 33.655 M | 33.655 M |
| Proposed method | 75.45% | 0.7423 | 0.7236 | 0.7246 | 115.088 M | 28.895 M |
| Baseline | TED-Adapter | Frame-Level Module | Posture Macro-F1 | Action Macro-F1 | 5-Class Acc | 5-Class Macro-F1 |
|---|---|---|---|---|---|---|
| ✓ | × | × | 0.9352 | 0.4338 | 58.77% | 0.3733 |
| ✓ | × | ✓ | 0.9414 | 0.6197 | 62.45% | 0.5511 |
| ✓ | ✓ | × | 0.9835 | 0.7263 | 72.45% | 0.6613 |
| ✓ | ✓ | ✓ | 0.9850 | 0.7455 | 73.85% | 0.7019 |
| Posture Macro-F1 | Action Macro-F1 | 5-Class Acc | 5-Class Macro-F1 | |
|---|---|---|---|---|
| 0.00 | 0.9850 | 0.7455 | 73.85% | 0.7019 |
| 0.05 | 0.9869 | 0.7383 | 74.87% | 0.7202 |
| 0.10 | 0.9830 | 0.7605 | 75.45% | 0.7246 |
| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Standing | 0.8854 | 0.7978 | 0.8394 | 465 |
| Lying | 0.6212 | 0.3705 | 0.4642 | 332 |
| Feeding | 0.7726 | 0.8931 | 0.8285 | 449 |
| Drinking | 0.7188 | 0.7419 | 0.7302 | 62 |
| Rumination | 0.7135 | 0.8147 | 0.7607 | 761 |
| Macro average | 0.7423 | 0.7236 | 0.7246 | – |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, X.; Wu, H.; Fan, D.; Bai, J.; Wang, C.; Liu, Y. A Clip-Based Dairy Cow Behavior Recognition Method Integrating Temporal Modeling and Behavioral Priors. Animals 2026, 16, 2087. https://doi.org/10.3390/ani16132087
Li X, Wu H, Fan D, Bai J, Wang C, Liu Y. A Clip-Based Dairy Cow Behavior Recognition Method Integrating Temporal Modeling and Behavioral Priors. Animals. 2026; 16(13):2087. https://doi.org/10.3390/ani16132087
Chicago/Turabian StyleLi, Xiaoying, Huijuan Wu, Daoerji Fan, Jiaqi Bai, Chunyun Wang, and Yan Liu. 2026. "A Clip-Based Dairy Cow Behavior Recognition Method Integrating Temporal Modeling and Behavioral Priors" Animals 16, no. 13: 2087. https://doi.org/10.3390/ani16132087
APA StyleLi, X., Wu, H., Fan, D., Bai, J., Wang, C., & Liu, Y. (2026). A Clip-Based Dairy Cow Behavior Recognition Method Integrating Temporal Modeling and Behavioral Priors. Animals, 16(13), 2087. https://doi.org/10.3390/ani16132087

