Validating Foundation Models for Automated Cattle Detection
Abstract
1. Introduction
- Validation of automated annotation: A comprehensive evaluation demonstrating that SAM 3-generated annotations achieve high agreement with manual ground truth, establishing their reliability for cattle detection tasks.
- EMA dataset: A large-scale, manually annotated dataset consisting of 6295 cow images with 25,014 oriented bounding boxes of fully visible cows, and a subset of 1096 images with 13,119 oriented bounding boxes covering both fully and partially visible cows. Manual annotations additionally include head orientations, posture labels, and visibility status—attributes not typically provided by automated tools. These supplementary features were not evaluated against automatic annotations, as the SAM 3 model outputs only object masks (which we converted to oriented bounding boxes) without semantic attributes.
- Large-scale automatically annotated dataset: 6295 cow images annotated using SAM 3, demonstrating the scalability of automated annotation methods for real-world barn environments.
- Training data requirements analysis: A systematic study showing that YOLO11-OBB models trained on SAM 3 annotations achieve comparable performance to those trained on manual annotations, and that only a small fraction of automated annotations is needed to reach high detection accuracy.
2. Related Research
2.1. Datasets
2.2. Automatic Annotation with Foundation Models and Detectors
2.3. Evaluation of Automatic Annotation
3. EMA Dataset
3.1. Image Acquisition
3.2. Manual Annotation Process
- 1.
- The annotator is presented with an image and clicks multiple points (typically 4–8) around each cow’s perimeter.
- 2.
- The tool automatically computes the minimum-area oriented bounding box enclosing these points and displays it for verification.
- 3.
- The annotator selects which side of the bounding box corresponds to the cow’s head orientation.
- 4.
- The annotator assigns a posture label: standing or lying.
- 5.
- For partially visible cows (e.g., at image boundaries or behind obstacles), the annotator sets a visibility flag.
- 6.
- This process repeats for all cows in the image.
- 7.
- If an image contains no cows or is corrupted, the annotator marks it accordingly.
- Image filename.
- Oriented bounding box: four vertex coordinates and center point.
- Head side: coordinates of the two vertices adjacent to the cow’s head.
- Orientation angle (in degrees).
- Posture label: standing or lying.
- Visibility flag: whole (fully visible) or partial (partially visible).
- Validity flag: indicates images without cows or with data corruption.
3.3. Dataset Structure and Composition
Manual Annotations (Ground Truth)
- EMA_6295: Contains manual annotations for 6295 images, focusing exclusively on fully visible cows, resulting in 25,014 oriented bounding boxes. These annotations were initially created to train a YOLO detector before we recognized that SAM 3 automatically annotates both fully and partially visible cows. To preserve this substantial annotation effort (approximately 220 h of expert labor) and provide valuable training data to the research community, we include this subset in the public release. However, for direct comparison with SAM 3 annotations in our experimental evaluation, we use only the EMA_1096 subset described below, which includes both fully and partially visible cows and is therefore directly comparable to SAM 3 output. Additionally, the annotation time invested in EMA_6295 provides important benchmark data for comparing manual versus automated annotation costs—a key contribution of this work.
- EMA_1096: A carefully annotated subset of 1096 images that includes both fully and partially visible cows, yielding 13,119 total annotations. Annotations in this subset, often referred to as ground truth in this paper, captures more challenging scenarios including edge cases, occlusions, and boundary conditions, and serves as the primary evaluation set for all experiments reported in this paper.
Automatic Annotations
- SAM_6295.
- SAM_1096.
Fine-Tuned Models
3.4. Dataset Statistics
3.5. Data Availability
4. Methodology
4.1. Segment Anything Model 3 (SAM 3)
4.2. YOLO11-OBB for Detection
5. Experimental Evaluation
5.1. SAM 3 Annotation Quality Assessment
5.1.1. Parameter Selection: Confidence and Minimal Relative Area
5.1.2. Matching Criterion: Combining IoU and IoM
- IoU ≥ 0.5.
- IoM ≥ threshold (to be determined).
5.1.3. Edge Cases and Limitations
Overlapping Cows
High Occlusion and Ambiguous Poses
Unmatched Pairs
Other Limitations
5.2. Comparing YOLO Models Fine-Tuned on Ground Truth vs. SAM 3 Annotations
5.3. Effect of Training Set Size on Detection Performance
6. Discussion
- Rapid initial improvement: Performance metrics improve substantially even with very small training set fractions. At just 2% of the training data, the model achieves precision of 0.934, recall of 0.912, and F1 of 0.923.
- Strong performance with minimal data: By 5% of the training set, the model reaches precision of 0.957, recall of 0.928, F1 of 0.942, and mAP@0.5–0.95 of 0.832—approaching the performance of models trained on the manual annotations.
- Diminishing returns: Beyond approximately 30% of the training data, metrics plateau, with only marginal improvements as dataset size increases. The mAP@0.5–0.95 stabilizes around 0.86, indicating that additional training data provides limited benefit once the detector has adapted to the visual characteristics of barn environments and cow appearances.
- Full dataset performance: The model trained on 100% of SAM_6295 annotations achieves precision of 0.943, recall of 0.946, F1 of 0.945, and mAP@0.5–0.95 of 0.862, demonstrating robust detection capability.
6.1. Summary of Experimental Findings
- 1.
- SAM 3 produces reliable annotations: With appropriate confidence (0.6) and minimal relative area (0.05) thresholds, with matching criterion IoU ≥ 0.5 or IoM ≥ 0.90, SAM 3 achieves precision of 0.911 and recall of 0.948.
- 2.
- Automated annotations enable competitive detector training: YOLO11-OBB models trained on SAM 3 annotations achieve 0.9391 precision, 0.9372 recall, and 0.8432 mAP@0.5–0.95, compared to 0.9635, 0.9698, and 0.8506 respectively for models trained on manual annotations.
- 3.
- Small training sets are sufficient: Only 5–10% of the SAM 3 training dataset is needed to achieve strong detection performance (F1 > 0.94, mAP@0.5–0.95 > 0.83), with diminishing returns beyond 30%.
- 4.
- Dramatic annotation cost savings: Manual annotation requires approximately 400 h of human labor, while SAM 3 processed the same dataset in 1 h with no human intervention required—completely eliminating the manual annotation bottleneck.
6.2. Limitations and Future Work
- Limited behavioral annotation: Manual annotations provided within our dataset include only posture labels (standing/lying) and head orientation. Comprehensive behavior monitoring requires additional annotations such as feeding, drinking, walking, and lameness indicators. However, labeling these behaviors on existing automated bounding box annotations is substantially faster than annotating from scratch, as annotators can focus solely on behavioral classification rather than object localization. This two-stage approach—automated detection followed by targeted behavioral labeling—represents a practical compromise between full automation and annotation cost.
- Single-farm dataset: Evaluation is limited to Mitrovac farm during two weeks. Generalization to other farms, breeds, and seasons requires empirical validation.
- Edge cases: Severe occlusion, lighting extremes, and small cow fragments create genuine annotation ambiguity for both manual and automated methods.
- Validating on multiple farms with different architectures, lighting, and cattle breeds.
- Extending SAM 3 with pose estimation and action recognition for comprehensive behavior annotation.
- Integrating automated detection with tracking and re-identification to validate end-to-end pipeline performance.
- Investigating whether head orientation and posture can be automatically inferred from segmentation masks or pose estimates.
7. Conclusions
- 1.
- Foundation models produce trustworthy annotations: SAM 3 achieves 0.911 precision and 0.948 recall, validating its use as an automated annotation tool.
- 2.
- Automated annotations enable competitive detector training: YOLO11-OBB models trained on SAM 3 annotations achieve 0.941 precision and 0.847 mAP@0.5–0.95, comparable to manually trained models, while significantly reducing annotation time.
- 3.
- Minimal training data is sufficient: Only 5–10% of the SAM_6295 dataset is needed to achieve strong detection performance, enabling rapid adaptation to new farm environments.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Peng, W.; Liu, Z.; Cai, J.; Zhao, Y. Research and application progress of electronic ear tags as infrastructure for precision livestock industry: A review. Intell. Robot. 2025, 5, 433–449. [Google Scholar] [CrossRef]
- Antognoli, V.; Presutti, L.; Bovo, M.; Torreggiani, D.; Tassinari, P. Computer Vision in Dairy Farm Management: A Literature Review of Current Applications and Future Perspectives. Animals 2025, 15, 2508. [Google Scholar] [CrossRef] [PubMed]
- Guarnido-Lopez, P.; Pi, Y.; Tao, J.; Mendes, E.D.; Tedeschi, L.O. Computer vision algorithms to help decision-making in cattle production. Anim. Front. 2024, 14, 11–22. [Google Scholar] [CrossRef] [PubMed]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. arXiv 2016, arXiv:1506.02640. [Google Scholar]
- Carion, N.; Gustafson, L.; Hu, Y.T.; Debnath, S.; Hu, R.; Suris, D.; Ryali, C.; Alwala, K.V.; Khedr, H.; Huang, A.; et al. SAM 3: Segment Anything with Concepts. arXiv 2025, arXiv:2511.16719. [Google Scholar]
- Andrew, W.; Gao, J.; Mullan, S.; Campbell, N.; Dowsey, A.W.; Burghardt, T. Visual identification of individual Holstein-Friesian cattle via deep metric learning. Comput. Electron. Agric. 2021, 185, 106133. [Google Scholar] [CrossRef]
- Gao, J.; Burghardt, T.; Andrew, W.; Dowsey, A.W.; Campbell, N.W. Towards Self-Supervision for Video Identification of Individual Holstein-Friesian Cattle: The Cows2021 Dataset. arXiv 2021, arXiv:2105.01938. [Google Scholar] [CrossRef]
- Yu, P.; Burghardt, T.; Dowsey, A.W.; Campbell, N.W. Holstein-Friesian re-identification using multiple cameras and self-supervision on a working farm. Comput. Electron. Agric. 2025, 237, 110568. [Google Scholar] [CrossRef]
- Zia, A.; Sharma, R.; Arablouei, R.; Bishop-Hurley, G.; McNally, J.; Bagnall, N.; Rolland, V.; Kusy, B.; Petersson, L.; Ingham, A. CVB: A Video Dataset of Cattle Visual Behaviors. arXiv 2023, arXiv:2305.16555. [Google Scholar] [CrossRef]
- Li, K.; Fan, D.; Wu, H.; Zhao, A. A new dataset for video-based cow behavior recognition. Sci. Rep. 2024, 14, 18702. [Google Scholar] [CrossRef] [PubMed]
- Koskela, O.; Benitez Pereira, L.S.; Pölönen, I.; Aronen, I.; Kunttu, I. Deep learning image recognition of cow behavior and an open data set acquired near an automatic milking robot. Agric. Food Sci. 2022, 31, 89–103. [Google Scholar] [CrossRef]
- Bošnjak, A.; Pejić, P.; Cupec, R.; Job, J.; Nyarko, E.; Lukić, B. Computer Vision for Automated Cattle Monitoring: A Review of Detection, Tracking, and Re-Identification. In Proceedings of the MIPRO Proceedings, Rijeka, Croatia, 25–29 May 2026. [Google Scholar]
- Das, M.; Ferreira, G.; Chen, C.J. Evaluating model generalization for cow detection in free-stall barn settings: Insights from the COw LOcalization (COLO) dataset. Smart Agric. Technol. 2025, 11, 101054. [Google Scholar] [CrossRef]
- Cao, Z.; Li, C.; Yang, X.; Zhang, S.; Luo, L.; Wang, H.; Zhao, H. Semi-automated annotation for video-based beef cattle behavior recognition. Sci. Rep. 2025, 15, 17131. [Google Scholar] [CrossRef] [PubMed]
- Zhao, L.; Olivier, K.; Chen, L. An Automated Image Segmentation, Annotation, and Training Framework of Plant Leaves by Joining the SAM and the YOLOv8 Models. Agronomy 2025, 15, 1081. [Google Scholar] [CrossRef]
- Yao, L.; Liu, J.; Hong, W.; Kong, F.; Fan, Z.; Lei, L.; Li, X. SideCow-VSS: A Video Semantic Segmentation Dataset and Benchmark for Intelligent Monitoring of Dairy Cows Health in Smart Ranch Environments. Vet. Sci. 2025, 12, 1104. [Google Scholar] [CrossRef] [PubMed]
- Araújo, V.M.; Rili, I.; Gisiger, T.; Gambs, S.; Vasseur, E.; Cellier, M.; Diallo, A.B. AI-powered cow detection in complex farm environments. Smart Agric. Technol. 2025, 10, 100770. [Google Scholar] [CrossRef]
- Patel, S.; Neethirajan, S. CowPain Check: AI-Based Facial Expression Analysis for Dairy Cow Welfare. J. Anim. Sci. Technol. 2025. [Google Scholar] [CrossRef]
- Li, G.; Li, X.; Zhang, S.; Yang, J. Towards more reliable evaluation in pedestrian detection by rethinking “ignore regions”. Vis. Intell. 2024, 2, 4. [Google Scholar] [CrossRef]
- Vogel, F.W.; Alipek, S.; Eppler, J.B.; Osuna-Vargas, P.; Triesch, J.; Bissen, D.; Acker-Palmer, A.; Rumpel, S.; Kaschube, M. Utilizing 2D-region-based CNNs for automatic dendritic spine detection in 3D live cell imaging. Sci. Rep. 2023, 13, 20497. [Google Scholar] [CrossRef] [PubMed]
- Velasco-Mata, A.; Ruiz-Santaquiteria, J.; Vallez, N.; Deniz, O. Using human pose information for handgun detection. Neural Comput. Appl. 2021, 33, 17273–17286. [Google Scholar] [CrossRef]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. arXiv 2023, arXiv:2304.02643. [Google Scholar]
- Ravi, N.; Gabeur, V.; Hu, Y.T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. SAM 2: Segment Anything in Images and Videos. arXiv 2024, arXiv:2408.00714. [Google Scholar]
- Ultralytics. Ultralytics YOLO11. 2024. Available online: https://docs.ultralytics.com/models/yolo11/ (accessed on 24 June 2026).







| Dataset | Images/Frames | Cows | Box Type | Setting | Identity | Behavior |
|---|---|---|---|---|---|---|
| OpenCows2020 [6] | 7043 | 46 | AABB | Barn + UAV | Yes | No |
| Cows2021 [7] | 10,402 | 186 | OBB (torso only) | Barn | Yes | No |
| MultiCamCows2024 [8] | 101,329 | 90 | Tracklets | Barn (multi-cam) | Yes | No |
| CVB [9] | 502 clips | 8 | AABB | Outdoor field | Yes | 11 classes |
| CBVD-5 [10] | 206,100 | 107 | AABB | Barn/ranch | No | 5 classes |
| Koskela et al. [11] | 1.7 M | - | Frame-level | Milking station | No | 10 classes |
| COLO [13] | 1254 | - | AABB | Free-stall barn | No | No |
| EMA (ours) | 6295 | - | OBB (full body) | Barn | No | Posture + head |
| Subset | Source Images | Images | Annotations | Annotation Type | Visibility | Experimental Purpose |
|---|---|---|---|---|---|---|
| EMA_6295 | All images | 6295 | 25,014 | Manual | Whole only | Reference for the manual annotation effort required for the complete image set |
| EMA_1096 | Selected common subset | 1096 | 13,119 | Manual | Whole + partial | Ground truth for annotation comparison and evaluation of all trained YOLO models |
| SAM_6295 | All images | 6295 | 77,381 | SAM 3 | Whole + partial | Evaluation of the effect of using a larger automatically annotated training set |
| SAM_1096 | Same images as EMA_1096 | 1096 | 13,638 | SAM 3 | Whole + partial | Direct comparison with EMA_1096 and controlled comparison of models trained using different annotation sources |
| Parameter | Value |
|---|---|
| Ultralytics version | 8.4.15 |
| Python version | 3.12.12 |
| PyTorch version | 2.7.0 + cu126 |
| CUDA version | 12.6 |
| OpenCV version | 4.13.0.92 |
| NumPy version | 2.4.2 |
| GPU | NVIDIA RTX 4000 Ada Generation (NVIDIA Corporation, Santa Clara, CA, USA) |
| Initial weights | yolo11n-obb.pt |
| Input image size (imgsz) | 1024 |
| Maximum epochs | 100 |
| Batch size | 4 |
| Workers | 8 |
| Early-stopping patience | 100 |
| Optimizer | AdamW, selected automatically |
| Initial learning rate | 0.002 |
| Momentum | 0.9 |
| Weight decay | 0.0005 |
| Pretrained weights | Enabled |
| Automatic mixed precision | Enabled |
| Training seeds | 0–9 |
| Dataset split seed | 123 |
| Deterministic option | Enabled |
| HSV hue (hsv_h) | 0.015 |
| HSV saturation (hsv_s) | 0.7 |
| HSV value (hsv_v) | 0.4 |
| Translation (translate) | 0.1 |
| Scaling (scale) | 0.5 |
| Horizontal flip (fliplr) | 0.5 |
| Mosaic (mosaic) | 1.0 |
| Final epochs without mosaic | 10 |
| Disabled augmentations | Rotation, shear, perspective, vertical flip, |
| MixUp, CutMix, and copy-paste |
| Confidence | MRA | Precision | Recall | F1 | AP |
|---|---|---|---|---|---|
| 0.6 | / | 0.889 | 0.954 | 0.920 | 0.938 |
| 0.05 | 0.911 | 0.948 | 0.929 | 0.930 | |
| 0.1 | 0.939 | 0.918 | 0.928 | 0.902 | |
| 0.15 | 0.957 | 0.877 | 0.915 | 0.865 | |
| 0.2 | 0.968 | 0.836 | 0.897 | 0.826 | |
| 0.75 | / | 0.951 | 0.908 | 0.929 | 0.892 |
| 0.05 | 0.956 | 0.904 | 0.929 | 0.892 | |
| 0.1 | 0.968 | 0.883 | 0.923 | 0.874 | |
| 0.15 | 0.977 | 0.849 | 0.908 | 0.836 | |
| 0.2 | 0.981 | 0.816 | 0.891 | 0.807 | |
| 0.9 | / | 0.992 | 0.595 | 0.744 | 0.591 |
| 0.05 | 0.992 | 0.595 | 0.744 | 0.591 | |
| 0.1 | 0.992 | 0.594 | 0.743 | 0.591 | |
| 0.15 | 0.993 | 0.590 | 0.740 | 0.591 | |
| 0.2 | 0.993 | 0.583 | 0.735 | 0.582 |
| Training Data | Precision | Recall | F1 | AP | mAP@0.5–0.95 |
|---|---|---|---|---|---|
| EMA_1096 (Manual) | 0.9635 ± 0.0035 | 0.9698 ± 0.0047 | 0.9666 ± 0.0025 | 0.9811 ± 0.0015 | 0.8506 ± 0.0059 |
| SAM_1096 (Automated) | 0.9391 ± 0.0060 | 0.9372 ± 0.0068 | 0.9381 ± 0.0026 | 0.9619 ± 0.0014 | 0.8432 ± 0.0019 |
| SAM_6295 (Automated) | 0.9485 ± 0.0046 | 0.9393 ± 0.0072 | 0.9438 ± 0.0021 | 0.9678 ± 0.0017 | 0.8551 ± 0.0036 |
| Training % | Precision | Recall | F1 | AP | mAP@0.5–0.95 |
|---|---|---|---|---|---|
| 2 | 0.934 | 0.912 | 0.923 | 0.954 | 0.794 |
| 3 | 0.944 | 0.924 | 0.934 | 0.958 | 0.799 |
| 5 | 0.957 | 0.928 | 0.942 | 0.967 | 0.832 |
| 10 | 0.946 | 0.938 | 0.942 | 0.963 | 0.841 |
| 15 | 0.958 | 0.935 | 0.946 | 0.967 | 0.848 |
| 20 | 0.956 | 0.932 | 0.944 | 0.967 | 0.850 |
| 30 | 0.954 | 0.941 | 0.948 | 0.970 | 0.859 |
| 50 | 0.950 | 0.944 | 0.947 | 0.970 | 0.860 |
| 75 | 0.957 | 0.938 | 0.948 | 0.969 | 0.857 |
| 100 | 0.943 | 0.946 | 0.945 | 0.969 | 0.862 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Pejić, P.; Bošnjak, A.; Cupec, R.; Nyarko, E.K.; Job, J.; Lukić, B. Validating Foundation Models for Automated Cattle Detection. Sensors 2026, 26, 5074. https://doi.org/10.3390/s26165074
Pejić P, Bošnjak A, Cupec R, Nyarko EK, Job J, Lukić B. Validating Foundation Models for Automated Cattle Detection. Sensors. 2026; 26(16):5074. https://doi.org/10.3390/s26165074
Chicago/Turabian StylePejić, Petra, Andrej Bošnjak, Robert Cupec, Emmanuel Karlo Nyarko, Josip Job, and Boris Lukić. 2026. "Validating Foundation Models for Automated Cattle Detection" Sensors 26, no. 16: 5074. https://doi.org/10.3390/s26165074
APA StylePejić, P., Bošnjak, A., Cupec, R., Nyarko, E. K., Job, J., & Lukić, B. (2026). Validating Foundation Models for Automated Cattle Detection. Sensors, 26(16), 5074. https://doi.org/10.3390/s26165074

