Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion
Simple Summary
Abstract
1. Introduction
2. Materials and Methods
2.1. Dataset
2.2. A Novel Network Based on Spatio-Temporal Enhancement and Dynamic Fusion (STED-Net)
2.2.1. Overview
2.2.2. Spatio-Temporal Enhancement Module (STE-Module)
2.2.3. Dynamic Fusion Block (DF-Block)
2.2.4. Joint Optimization Loss for Enhancement and Recognition
2.3. Evaluation Indicators
2.4. Experimental Setup
3. Results and Discussion
3.1. Comparisons with the State-of-the-Art Behavior Recognition Methods
3.2. Visualization and Analysis of Results
3.3. Ablation Study
3.3.1. Effectiveness of the STE-Module and DF-Block
3.3.2. Effectiveness of Different Branches in Dynamic Fusion Block
3.3.3. Effectiveness of the Dynamic Fusion Block
3.4. Robustness and Generalization Analysis
3.4.1. Robustness Analysis Under Different Weather Conditions
3.4.2. Cross-Farm Generalization Analysis
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Zhao, Y.; Feng, L.; Tang, J.; Zhao, W.; Ding, Z.; Li, A.; Zheng, Z. Automatically recognizing four-legged animal behaviors to enhance welfare using spatial temporal graph convolutional networks. Appl. Anim. Behav. Sci. 2022, 249, 105594. [Google Scholar] [CrossRef]
- Deepak, D.; D’Mello, D.A.; Divakarla, U. Advancements in Automated Livestock Monitoring: A Concise Review of Deep Learning-Based Cattle Activity Recognition. In Proceedings of the 2024 10th International Conference on Advanced Computing and Communication Systems (ICACCS); IEEE: Coimbatore, India, 2024; Volume 1, pp. 321–327. [Google Scholar]
- Kim, S.J.; Jin, X.C.; Bharanidharan, R.; Kim, N.Y. Monitoring Multiple Behaviors in Beef Calves Raised in Cow–Calf Contact Systems Using a Machine Learning Approach. Animals 2024, 14, 3278. [Google Scholar] [CrossRef] [PubMed]
- Dhakshinamoorthy, D.; Jha, A.; Majumdar, S.; Ghosh, D.; Chakraborty, R.; Ray, H. Classification of Cattle Behavior and Detection of Heat (Estrus) using Sensor Data. arXiv 2025, arXiv:2506.16380. [Google Scholar]
- Myat Noe, S.; Zin, T.T.; Tin, P.; Kobayashi, I. Comparing state-of-the-art deep learning algorithms for the automated detection and tracking of black cattle. Sensors 2023, 23, 532. [Google Scholar] [PubMed]
- Fuentes, A.; Han, S.; Nasir, M.F.; Park, J.; Yoon, S.; Park, D.S. Multiview monitoring of individual cattle behavior based on action recognition in closed barns using deep learning. Animals 2023, 13, 2020. [Google Scholar] [CrossRef] [PubMed]
- Zheng, Z.; Qin, L. PrunedYOLO-Tracker: An efficient multi-cows basic behavior recognition and tracking technique. Comput. Electron. Agric. 2023, 213, 108172. [Google Scholar]
- Li, G.; Sun, J.; Guan, M.; Sun, S.; Shi, G.; Zhu, C. A New Method for non-destructive identification and Tracking of multi-object behaviors in beef cattle based on deep learning. Animals 2024, 14, 2464. [Google Scholar] [PubMed]
- Giannone, C.; Sahraeibelverdy, M.; Lamanna, M.; Cavallini, D.; Formigoni, A.; Tassinari, P.; Torreggiani, D.; Bovo, M. Automated dairy cow identification and feeding behaviour analysis using a computer vision model based on YOLOv8. Smart Agric. Technol. 2025, 12, 101304. [Google Scholar] [CrossRef]
- Li, X.; Sun, K.; Fan, H.; He, Z. Real-time cattle pose estimation based on improved rtmpose. Agriculture 2023, 13, 1938. [Google Scholar] [CrossRef]
- Wei, Y.; Zhang, H.; Gong, C.; Wang, D.; Ye, M.; Jia, Y. Study of pose estimation based on spatio-temporal characteristics of cow skeleton. Agriculture 2023, 13, 1535. [Google Scholar] [CrossRef]
- Perneel, M.; Adriaens, I.; Verwaeren, J.; Aernouts, B. Dynamic Multi-Behaviour, Orientation-Invariant Re-Identification of Holstein-Friesian Cattle. Sensors 2025, 25, 2971. [Google Scholar] [CrossRef] [PubMed]
- Hua, Z.; Wang, Z.; Xu, X.; Kong, X.; Song, H. An effective PoseC3D model for typical action recognition of dairy cows based on skeleton features. Comput. Electron. Agric. 2023, 212, 108152. [Google Scholar] [CrossRef]
- Wang, Y.; Li, R.; Wang, Z.; Hua, Z.; Jiao, Y.; Duan, Y.; Song, H. E3D: An efficient 3D CNN for the recognition of dairy cow’s basic motion behavior. Comput. Electron. Agric. 2023, 205, 107607. [Google Scholar] [CrossRef]
- Yin, X.; Wu, D.; Shang, Y.; Jiang, B.; Song, H. Using an EfficientNet-LSTM for the recognition of single Cow’s motion behaviours in a complicated environment. Comput. Electron. Agric. 2020, 177, 105707. [Google Scholar]
- Wu, D.; Wang, Y.; Han, M.; Song, L.; Shang, Y.; Zhang, X.; Song, H. Using a CNN-LSTM for basic behaviors detection of a single dairy cow in a complex environment. Comput. Electron. Agric. 2021, 182, 106016. [Google Scholar] [CrossRef]
- Fuentes, A.; Yoon, S.; Park, J.; Park, D.S. Deep learning-based hierarchical cattle behavior recognition with spatio-temporal information. Comput. Electron. Agric. 2020, 177, 105627. [Google Scholar] [CrossRef]
- Tian, F.; Zhang, L.; Zhang, J.; Zhang, S.; Soomro, S.A.; Xiong, B.; Shen, W.; Song, Z.; Yan, Y.; Yu, Z. Cattle-ES3D: A spatiotemporal feature fusion method for detecting tachypnea and salivation behaviors in beef cattle. Comput. Electron. Agric. 2025, 239, 110907. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst. 2015, 28. [Google Scholar] [CrossRef]
- Redmon, J.; Farhadi, A. YOLO9000: Better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 7263–7271. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Tsai, Y.C.; Hsu, J.T.; Ding, S.T.; Rustia, D.J.A.; Lin, T.T. Assessment of dairy cow heat stress by monitoring drinking behaviour using an embedded imaging system. Biosyst. Eng. 2020, 199, 97–108. [Google Scholar] [CrossRef]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Han, Y.; Wu, J.; Zhang, H.; Cai, M.; Sun, Y.; Li, B.; Feng, X.; Hao, J.; Wang, H. Beef cattle abnormal behaviour recognition based on dual-branch frequency channel temporal excitation and aggregation. Biosyst. Eng. 2024, 241, 28–42. [Google Scholar] [CrossRef]
- Xiao, D.; Wang, H.; Liu, Y.; Li, W.; Li, H. DHSW-YOLO: A duck flock daily behavior recognition model adaptable to bright and dark conditions. Comput. Electron. Agric. 2024, 225, 109281. [Google Scholar] [CrossRef]
- Li, D.; Dai, B.; Li, Y.; Song, P.; Dai, X.; He, Y.; Liu, H.; Li, Y.; Shen, W. IATEFF-YOLO: Focus on cow mounting detection during nighttime. Biosyst. Eng. 2024, 246, 54–66. [Google Scholar] [CrossRef]
- Langford, F.; Rutherford, K.; Sherwood, L.; Jack, M.; Lawrence, A.; Haskell, M. Behavior of cows during and after peak feeding time on organic and conventional dairy farms in the United Kingdom. J. Dairy Sci. 2011, 94, 746–753. [Google Scholar] [CrossRef] [PubMed]
- Marumo, J.L.; Lusseau, D.; Speakman, J.R.; Mackie, M.; Byar, A.Y.; Cartwright, W.; Hambly, C. Behavioural variability, physical activity, rumination time, and milk characteristics of dairy cattle in response to regrouping. Animal 2024, 18, 101094. [Google Scholar] [CrossRef] [PubMed]
- Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; Gool, L.V. Temporal Segment Networks for Action Recognition in Videos. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 20–36. [Google Scholar]
- Tran, D.Q.; Aboah, A.; Jeon, Y.; Shoman, M.; Park, M.; Park, S. Low-light image enhancement framework for improved object detection in fisheye lens datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 7056–7065. [Google Scholar]
- Dai, Y.; Gieseke, F.; Oehmcke, S.; Wu, Y.; Barnard, K. Attentional Feature Fusion. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Virtual, 5–9 January 2021; pp. 3560–3569. [Google Scholar]
- Guo, C.; Li, C.; Guo, J.; Loy, C.C.; Hou, J.; Kwong, S.; Cong, R. Zero-reference deep curve estimation for low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 1780–1789. [Google Scholar]
- Zeng, K.; Wang, Z. 3D-SSIM for video quality assessment. In Proceedings of the 2012 19th IEEE International Conference on Image Processing; IEEE: Orlando, FL, USA, 2012; pp. 621–624. [Google Scholar]
- Xu, Y.; Yang, J.; Cao, H.; Mao, K.; Yin, J.; See, S. Arid: A new dataset for recognizing action in the dark. In Deep Learning for Human Activity Recognition, Proceedings of the Second International Workshop, DL-HAR 2020; Springer: Singapore, 2021; pp. 70–84. [Google Scholar]
- Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; Paluri, M. Learning Spatiotemporal Features with 3D Convolutional Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 4489–4497. [Google Scholar]
- Carreira, J.; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6299–6308. [Google Scholar]
- Tran, D.; Wang, H.; Torresani, L. A Closer Look at Spatiotemporal Convolutions for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 6450–6459. [Google Scholar]
- Zhou, B.; Andonian, A.; Oliva, A.; Torralba, A. Temporal relational reasoning in videos. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 803–818. [Google Scholar]
- Feichtenhofer, C. SlowFast Networks for Video Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6202–6211. [Google Scholar]
- Lin, J.; Gan, C.; Han, S. TSM: Temporal Shift Module for Efficient Video Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7083–7093. [Google Scholar]
- Liu, Z.; Wang, L.; Wu, W.; Qian, C.; Lu, T. Tam: Temporal adaptive module for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 13708–13718. [Google Scholar]
- Yang, C.; Xu, Y.; Shi, J.; Dai, B.; Zhou, B. Temporal pyramid network for action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 591–600. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- Chen, R.; Chen, J.; Liang, Z.; Gao, H.; Lin, S. Darklight networks for action recognition in the dark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 846–852. [Google Scholar]
- He, K.; Sun, J.; Tang, X. Single image haze removal using dark channel prior. IEEE Trans. Pattern Anal. Mach. Intell. 2010, 33, 2341–2353. [Google Scholar] [CrossRef] [PubMed]
- Si, Y.; Yang, F.; Guo, Y.; Zhang, W.; Yang, Y. A comprehensive benchmark analysis for sand dust image reconstruction. J. Vis. Commun. Image Represent. 2022, 89, 103638. [Google Scholar] [CrossRef]
- Garg, K.; Nayar, S.K. Vision and rain. Int. J. Comput. Vis. 2007, 75, 3–27. [Google Scholar] [CrossRef]
















| Behavior Category | Behavioral Definition |
|---|---|
| Running | Beef cattle perform rapid and continuous spatial movement at a speed clearly higher than normal walking. |
| Feeding | Beef cattle approach the feed trough and extend their heads into the trough area for feed intake. |
| Drinking | Beef cattle approach the water trough and extend their heads into the trough area for water intake. |
| Grooming | Beef cattle approach the grooming brush and rub their body against the brush repeatedly. |
| Mounting | A beef cattle raises the front leg and mounts another beef cattle. |
| Fighting | Beef cattle show aggressive interactions with other individuals, such as head-butting, pushing, chasing, or physical confrontation. |
| Method | Precision | Recall | Accuracy | F1-Score | Parameters (M) | Inference Time (ms) |
|---|---|---|---|---|---|---|
| C3D | 85.09% | 70.14% | 66.83% | 76.89% | 78.02 | 40.92 |
| TSN | 66.99% | 47.51% | 39.48% | 55.59% | 23.52 | 19.63 |
| I3D | 67.98% | 56.65% | 45.38% | 61.80% | 35.40 | 56.69 |
| R(2 + 1)D | 62.26% | 33.64% | 38.58% | 43.68% | 63.76 | 127.46 |
| TRN | 75.00% | 57.01% | 44.14% | 64.78% | 26.64 | 60.61 |
| SlowFast | 69.07% | 58.47% | 47.57% | 63.33% | 42.10 | 105.61 |
| TSM | 86.75% | 69.23% | 65.96% | 77.01% | 23.86 | 32.10 |
| TAM | 69.35% | 56.54% | 48.48% | 62.29% | 25.59 | 282.46 |
| TPN | 70.38% | 57.92% | 46.87% | 63.54% | 91.50 | 329.27 |
| Swin-T | 73.30% | 58.73% | 74.50% | 65.21% | 88.83 | 123.20 |
| STED-Net (ours) | 88.47% | 80.18% | 83.80% | 84.12% | 178.73 | 292.26 |
| # | Baseline | DF-Block | STE-Module | Precision | Recall | Accuracy | F1-Score |
|---|---|---|---|---|---|---|---|
| 1 | ✓ | 81.82% | 71.82% | 76.59% | 76.49% | ||
| 2 | ✓ | ✓ | 83.87% | 74.65% | 76.57% | 78.99% | |
| 3 | ✓ | ✓ | 85.76% | 75.45% | 78.18% | 80.28% | |
| 4 | ✓ | ✓ | ✓ | 88.47% | 80.18% | 83.80% | 84.12% |
| # | Dark-Branch | STE-Branch | HE-Branch | Precision | Recall | Accuracy | F1-Score |
|---|---|---|---|---|---|---|---|
| 1 | ✓ | 83.61% | 47.51% | 78.35% | 60.59% | ||
| 2 | ✓ | 76.09% | 72.73% | 74.09% | 74.37% | ||
| 3 | ✓ | 84.39% | 75.00% | 70.02% | 79.42% | ||
| 4 | ✓ | ✓ | 85.99% | 75.00% | 80.77% | 80.12% | |
| 5 | ✓ | ✓ | 79.17% | 65.46% | 65.98% | 71.66% | |
| 6 | ✓ | ✓ | 83.86% | 68.64% | 75.70% | 75.49% | |
| 7 | ✓ | ✓ | ✓ | 88.47% | 80.18% | 83.80% | 84.12% |
| Feature Fusion Function | Precision | Recall | Accuracy | F1-Score |
|---|---|---|---|---|
| STED-Net with Concat | 79.61% | 73.64% | 76.82% | 76.51% |
| STED-Net with Cross-Attention | 81.45% | 75.46% | 79.43% | 78.34% |
| STED-Net with AFF | 82.25% | 76.23% | 80.83% | 79.13% |
| STED-Net with DF-Block (ours) | 88.47% | 80.18% | 83.80% | 84.12% |
| Weather Conditions | Precision | Recall | Accuracy | F1-Score |
|---|---|---|---|---|
| Foggy | 84.21% | 76.35% | 79.58% | 80.09% |
| Dusty | 58.10% | 47.30% | 50.20% | 52.10% |
| Rainy | 77.80% | 69.50% | 72.50% | 73.40% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Han, Y.; Zhang, Z.; Zhang, W.; Cao, S.; Sun, Y.; Jia, Z.; Wu, D.; Huang, L.; Zhang, H. Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion. Animals 2026, 16, 1881. https://doi.org/10.3390/ani16121881
Han Y, Zhang Z, Zhang W, Cao S, Sun Y, Jia Z, Wu D, Huang L, Zhang H. Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion. Animals. 2026; 16(12):1881. https://doi.org/10.3390/ani16121881
Chicago/Turabian StyleHan, Yamin, Zhenyu Zhang, Wenchao Zhang, Shichao Cao, Yang Sun, Zixin Jia, Danyang Wu, Lyuwen Huang, and Hongming Zhang. 2026. "Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion" Animals 16, no. 12: 1881. https://doi.org/10.3390/ani16121881
APA StyleHan, Y., Zhang, Z., Zhang, W., Cao, S., Sun, Y., Jia, Z., Wu, D., Huang, L., & Zhang, H. (2026). Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion. Animals, 16(12), 1881. https://doi.org/10.3390/ani16121881

