Next Article in Journal
Solid-State Fermented Discarded Dates as a Functional Feed Ingredient: Effects on Meat Quality, Fatty Acid Profile, and Essential Amino Acid Composition
Previous Article in Journal
Co-Occurrence of an NDM-1-Carrying SGI1 Variant and a Novel GIflu-1 Resistance Island in a Seafood-Derived Vibrio fluvialis
Previous Article in Special Issue
Pilot Room-Level Acoustic and Physiological Monitoring of Respiratory Disturbance in Pigs Following Experimental Klebsiella pneumoniae Challenge
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Infrared Thermography and Machine Learning for Mastitis Detection in Dairy Cows: A Pilot Case Study in Egyptian Farms

1
Department of Animal and Fish Production, Faculty of Agriculture, Alexandria University, Alexandria 21545, Egypt
2
Animal Production Research Institute, Agricultural Research Center, Giza 12618, Egypt
3
Department of Artificial Intelligence, Faculty of Engineering, Mansoura National University, Gamasa 33515, Egypt
4
Department of Scientific computing, Faculty of Computers and Artificial Intelligence, Benha University, Benha 19111, Egypt
5
Department of Animal and Poultry Production, College of Agriculture and Food, Qassim University, Buraydah 51452, Saudi Arabia
*
Author to whom correspondence should be addressed.
Vet. Sci. 2026, 13(7), 640; https://doi.org/10.3390/vetsci13070640
Submission received: 21 May 2026 / Revised: 12 June 2026 / Accepted: 18 June 2026 / Published: 30 June 2026

Simple Summary

Mastitis is one of the most common and economically important diseases in dairy cattle because it affects animal health, welfare, milk yield, and milk quality. This pilot study examined whether infrared thermal images of the udder, combined with machine-learning classifiers, could support mastitis screening in Holstein dairy cows. Several classifiers were tested, and the multi-layer perceptron and support vector machine models showed the best image-level performance. This study also explored selected biological and feeding-related variables, but these findings should be interpreted cautiously because they were exploratory. Overall, the results suggest that infrared thermography combined with machine learning may be useful as a supportive, non-invasive screening approach, although larger independent studies are still needed before routine farm application.

Abstract

Mastitis is a major and costly dairy disease that reduces milk yield and quality and harms animal welfare. This study evaluated infrared thermography (IRT) combined with machine learning (ML) for non-invasive mastitis screening in dairy cows and explored links with biological and feeding-system variables in Egyptian farms. A total of 976 thermal udder images obtained from 488 Holstein cows were used, including 708 healthy and 268 mastitic images. Images were captured before milking, processed with CLAHE, resized to 224 × 224 pixels, and split using cow-level grouping before augmentation to prevent animal-level data leakage. The training set contained 780 original images and was augmented to a balanced 4708-image set (2354 per class), while the held-out test set remained unaugmented, with 196 original images (142 healthy and 54 mastitic). EfficientNetB3 with global average and max pooling extracted 3072 thermal features, and ten ML classifiers were evaluated. In the image-level hold-out evaluation, MLP achieved the best performance (accuracy = 86.22%, AUC = 0.9184, sensitivity = 74.07%, specificity = 90.85%), followed by SVM (accuracy = 83.67%, AUC = 0.8963). A separate group-based five-fold cross-validation yielded a more conservative AUC of 0.6812 ± 0.1323 and accuracy of 0.6244 ± 0.0642. Logistic regression analyses did not identify statistically significant associations between model predictions and somatic cell count (SCC), California Mastitis Test (CMT), blood biomarkers, or nutritional variables at p < 0.05. Ration A (Delta Misr) showed a higher observed mastitis incidence (20/40; 50.0%) than Ration B (Copenhagen; 16/45; 35.6%), but nutritional predictors were not statistically significant, indicating that farm-level confounding should be considered. Overall, IRT with ML remains a promising non-invasive screening approach, but broader multicenter datasets and independent external validation are needed before routine farm deployment.

1. Introduction

The animal production sector faces numerous challenges, including scarcity and fluctuating year-round feed supplies, as well as high feed prices, climatic changes, heat stress, water and land use issues, supply chain disruptions, health risks, and zoonotic diseases. Mastitis is one of the most common and economically significant diseases in dairy cattle, severely affecting animal health, welfare, milk yield, and quality. It remains a major challenge to sustainable dairy production worldwide [1,2].
Mastitis is the inflammation of the udder tissue, resulting from the invasion of the mammary gland by pathogens. The occurrence of mastitis is influenced by multiple factors, including environmental factors, such as milking methods, housing type, microclimate, production season, and nutrition. Udder inflammation can also result from intrinsic factors such as immune status, milk yield, age, lactation phase, stress, and udder morphology [3,4].
Mastitis is the most costly disease in European dairy farms, causing significant economic losses. It can affect up to 50% of cows during their lives, even in herds with good hygiene practices, although prevalence varies across farms [5]. Infection of a single udder quarter can decrease milk production by at least 10%. The production of safe, high-quality milk depends on udder health, which is crucial for global nutrition [6]. The clinical form of mastitis features visible inflammatory changes in milk and udder tissue, with or without systemic clinical signs. In contrast, the subclinical form does not show obvious signs of mastitis but increases somatic cell count. Cows with subclinical infections can serve as a source of infection for susceptible animals in the herd [7]. Detecting subclinical mastitis requires cow-side tests such as the California mastitis test (CMT) or laboratory tests such as somatic cell count (SCC) and milk bacteriological culture. Early diagnosis of subclinical mastitis is vital to prevent infection spread, minimize damage to udder tissue, and ensure successful treatment [8,9]. Therefore, developing rapid, accurate, and non-invasive screening methods is essential for timely diagnosis and effective control.
The development of artificial intelligence (AI) technologies, particularly computer vision, offers innovative methods for monitoring and enhancing animal welfare [10]. Simultaneously, AI has experienced significant advancements, including developments in machine learning (ML), neural networks, and deep learning (DL), which have demonstrated strong performance across multiple domains [11]. To improve operational efficiency, economic sustainability, and environmental outcomes in intensive livestock systems, smart livestock farming (SLF) is emerging as a successor to precision livestock farming (PLF) [12]. Given the rising demand for SLF, traditional animal production and management methods are now considered outdated and insufficient for addressing contemporary livestock challenges [13].
Inflammation in the udder leads to increased blood flow and metabolic activity, resulting in localized heat that can be detected by infrared thermography (IRT). The IRT is a non-invasive, remote, and passive technique that measures the surface temperature of a body through the infrared radiation of the electromagnetic spectrum that it emits. Each pixel represents a temperature value, with the coolest areas appearing blue or black and the warmest areas appearing white or red [14,15]. When combined with ML algorithms, IRT may provide a supportive screening approach for automatically identifying temperature variations associated with mastitis, while reducing reliance on manual visual observation [16].
This imaging approach is based on the principles of infrared radiation, which has a longer wavelength than visible light and a shorter wavelength than microwaves. IRT detects infrared radiation emitted from the body surface and converts it into a thermal image in which pixel intensity reflects surface-temperature variation. IRT is a rapid, non-invasive, and safe imaging approach that can detect changes in udder surface temperature associated with inflammatory reactions [17].
Researchers have established a positive relationship between udder skin surface temperature and the California mastitis test (CMT) and the Somatic Cell Count (SCC) score [14,17]. Recently, subclinical mastitis (SCM) in dairy cattle has been identified and predicted using machine-learning algorithms, which integrate and analyze data from various sources [18,19,20]. ML approaches to database analysis using digital technologies such as IRT are innovative and exciting tools for generating information for herd health monitoring and for reducing the negative effects of SCM in dairy cows [16].
Therefore, the contribution of this study lies in combining IRT with EfficientNetB3-based dual-pooling feature extraction, multiple conventional and neural-network classifiers, and leakage-aware cow-level validation. In addition to image classification, exploratory biomarker and feeding-system analyses were included to examine biological and management contexts without overstating causality. We hypothesized that IRT-derived thermal features combined with ML models could support non-invasive mastitis screening in dairy cows, although further validation is required before practical farm deployment.
Therefore, the objective of this study was to evaluate different ML models for detecting mastitis-related thermal patterns from IRT images and to explore, without causal inference, whether biomarker and feeding-system variables were associated with the observed mastitis-related outcomes.

2. Materials and Methods

This experiment was conducted at Alexandria Copenhagen Farm, Delta Misr Farm, and the Department of Animal and Fish Production, Faculty of Agriculture, Alexandria University. All procedures were approved and authorized by the Institutional Animal Care and Use Committee of Alexandria University (Protocol ID: Alex. Agri. 082410121). This study aimed to evaluate IRT integrated with ML models for non-invasive mastitis screening and to explore whether differences between farms and feed rations were associated with mastitis incidence. As the analysis did not identify statistically significant direct nutritional effects, the feeding system analysis was treated as exploratory.

2.1. Animals and Farm Description

The study was conducted on Holstein dairy cows from two dairy farms. Cows were selected randomly with varying degrees of lineage, between 10 and 260 days in lactation and average of live body weight of 620 ± 32 kg. Both farms operated under intensive management systems. Cows were housed in free-stall barns with controlled ventilation systems. Feeding was based on a total mixed ration (TMR), with differences in nutritional composition between the two farms. Image acquisition was performed under broadly stable farm conditions; however, exact ambient temperature and humidity ranges were not continuously recorded.

2.2. Thermal Imaging

Thermal images were captured immediately before milking using a UNI-T UTx313 thermal imaging camera. Images were acquired at an approximate distance of 1 m from the udder. The camera was positioned in approximately frontal orientation and nearly perpendicular to the right or left udder half. Before image acquisition, the udders were cleaned and dried to reduce the influence of surface moisture or contamination on thermal measurements. Images were acquired under stable farm conditions, and acquisition settings were kept consistent as far as possible. However, exact ambient temperature and humidity ranges were not continuously recorded, which is acknowledged as a limitation of the study. The key technical specifications of the UTx313 thermal imaging camera relevant to udder thermal-image acquisition are provided in Supplementary Table S1.

2.3. Data Loading and Preprocessing

2.3.1. Chemical Analysis of TMR

Experimental diets were dried in a forced-air oven at 55 °C for 72 h, ground using a Wiley mill grinder to pass through a 1 mm stainless steel screen, and subsequently analyzed for dry matter (DM), organic matter (OM), ether extract (EE), and crude protein (CP) according to AOAC [21]. Neutral detergent fiber (NDF) and acid detergent fiber (ADF) were determined using an Ankom fiber analyzer (Fiber Analyzer A200; Ankom Technology, Macedon, NY, USA) [22]. The ingredients of the total mixed ration (TMR) and the chemical composition on a DM basis for the cows on the two farms are provided in Supplementary Table S2. Nutrient-density values were calculated using NRC [23] dairy equations. The feeding-system comparison was interpreted as exploratory because only two farms were included, and potential farm-level confounders such as management, hygiene, and environmental conditions could not be fully separated from ration effects.

2.3.2. Milk Yield and Sampling

The cows were milked three times daily at 04:00, 12:00, and 20:00 in a milking parlor equipped with automatic cow identification, a milk recording system, and automated detacher milker units, which exist at Copenhagen (DeLaval herringbone) and Delta Misr (DeLaval rapid exit) farms. Individual milk yield was recorded daily using the Delaval program with the rapid exit system. Every week, individual milk samples (50 mL of milk from each cow) were collected and immediately analyzed using the California Mastitis Test (CMT). Subsequently, SCC, as well as electrical conductivity (EC) and milk composition (consisting of fat, protein, added water, solids-not-fat (SNF), freezing point, lactose, and density), were determined using a milk analyzer (Ekomilk Horizon Unlimited, Stara Zagora, Bulgaria). SCC was used as an additional comparative indicator for the biological-assessment group.
For the machine-learning image dataset, mastitis labels were assigned according to the California Mastitis Test (CMT) result at the cow/whole-udder level. CMT-negative cows were labelled as healthy, whereas CMT-positive cows were labelled as mastitic. Right- and left-udder-side thermal images obtained from the same cow were assigned the same cow-level diagnostic label. Quarter-level diagnosis was not performed, and clinical and subclinical mastitis cases were not analyzed separately.
Moreover, microbiological evaluation of coliform bacteria and Staphylococcus aureus was performed on milk samples collected from dairy cows in the biological-assessment group (n = 85). The colony-count technique was used, and typical black, shiny colonies surrounded by a clear halo were counted according to ISO 4832 [24] and ISO 6888-1 [25], respectively. No significant differences were observed for coliform bacteria or Staphylococcus aureus among samples, suggesting that these microorganisms were not the predominant detected causative agents in the examined cases. Other pathogens or non-microbial factors may have contributed to the observed mastitis-related conditions.

2.3.3. Blood Sampling

Blood samples (10 mL, n = 85) were collected weekly before milking from the jugular vein of cows included in the biological-assessment group during the morning hours over five weeks. The samples were taken in clot activator tubes (Vacutainer, Becton Dickinson, Franklin Lakes, NJ, USA) and centrifuged at 3000 rpm for 20 min at room temperature. The serum was then harvested and stored at −20 °C until analysis. Biochemical parameters in the serum were determined using commercial colorimetric kits, including total protein [26], albumin [27], and glucose [28]. Globulin concentration was calculated as the difference between total protein and albumin. Beta hydroxybutyrate (BHBA) [29], non-esterified fatty acids (NEFAs) [30], lactate dehydrogenase (LDH) [31], IgG, IgA, and IgE [32] were also measured. These parameters indicate animal health, immunity, and energy balance and were used for exploratory biological contextualization of mastitis-related status rather than as definitive independent validation endpoints for the machine-learning model. Variables were selected based on their biological relevance to mastitis.

2.4. Mastitis Detection Using Machine Learning and Statistical Analysis

2.4.1. Acquisition of Thermograms

A dataset of 976 thermal udder images was collected from Holstein dairy cows at early, mid, and late lactation stages from dairy farms in Egypt. The dataset included 708 healthy and 268 mastitic images. The images were acquired immediately before milking in the milking parlor. The 976 thermal udder images were obtained from 488 cows, with right- and left-udder-side images assigned to the same cow-level group identifier before train-test splitting. When more than one image was available from the same cow, all images from a given cow were kept within the same training or held-out test subset to reduce animal-level data leakage. Camera position, imaging distance, and acquisition settings were kept as consistent as possible during image capture.

2.4.2. Data Augmentation and Balancing

To address the class imbalance while preventing data leakage, the original images were first split by cow ID into training and held-out test subsets. Data augmentation was then applied only to the training set. Images in the training subset were augmented using the Keras ImageDataGenerator until both classes were balanced at the class count. Transformations included rotation up to ±30°, width and height shifts up to 10% of image dimensions, zoom up to 20%, horizontal flipping, and nearest-neighbor filling for border pixels. The held-out test set was kept unaugmented and contained only original images.

2.4.3. Dataset

The image dataset consisted of BMP-format thermal udder images from two classes, healthy and mastitic cows. The BMP images represented thermal udder-image outputs from the camera and were used consistently for image-based feature extraction; raw pixel-level radiometric temperature values were not used for model training. Images were loaded from separate directories using OpenCV (cv2), resized to 224 × 224 pixels, and processed as shown in Table 1.
The complete dataset flow, including cow-level splitting before augmentation, training-only augmentation, held-out testing, animal-level stratified group cross-validation, and qualitative external application, is summarized in Figure 1.

2.4.4. CLAHE Contrast Enhancement

All images underwent Contrast Limited Adaptive Histogram Equalization (CLAHE) to enhance local contrast and improve microstructural visibility. It was applied exclusively to the Luminance (L) channel in the LAB color space to avoid color distortion. It is preferred over global histogram equalization as it prevents over-amplification of noise in homogeneous regions [33]. The udder-focused region of interest was used before feature extraction to reduce the influence of irrelevant background information.

2.4.5. Deep Feature Extraction

Transfer Learning with EfficientNetB3
Feature extraction was performed using EfficientNetB3, a member of the EfficientNet family of convolutional neural networks introduced by Tan and Le. [34]. EfficientNet models employ a compound-scaling method that uniformly scales network depth, width, and resolution using a set of fixed scaling coefficients, resulting in superior accuracy–efficiency trade-offs compared to conventional architectures. EfficientNetB3 was loaded with ImageNet pre-trained weights and used as a fixed feature extractor (including top = false), without fine-tuning the convolutional weights. This transfer learning approach is justified by the limited size of the medical image dataset and the proven generalizability of ImageNet features to medical imaging tasks [35].
Span and Level Feature Representation
Two complementary pooling strategies were applied to the final convolutional feature maps to produce a rich feature representation:
  • Level features (Global Average Pooling-GAP) computes the spatial average of each feature map channel, capturing the overall intensity distribution and global texture patterns across the entire image.
  • Span features (Global Max Pooling-GMP) captures the maximum activation in each feature channel, highlighting the presence of the most discriminative local features regardless of their spatial location.
The GAP and GMP vectors were concatenated to form the final feature vector, resulting in a 3072-dimensional representation (1536 × 2) per image. This dual-pooling strategy, referred to as ‘Span and Level’ feature extraction, has been shown to outperform single-pooling approaches by capturing both global context (level) and local discriminative regions (span).
Feature Standardization
Following feature extraction, all feature vectors were standardized using a standard scaler (zero mean, unit variance). The scaler was fitted on the training set and applied to both training and test sets to prevent data leakage. Feature standardization is essential for distance-based and gradient-based classifiers to prevent features with larger magnitudes from dominating the learning process.

2.4.6. Machine-Learning Classification

Train/Test Split
Before augmentation, the dataset was partitioned using GroupShuffleSplit with cow ID as the grouping variable (80% training and 20% testing; random seed = 42). This ensured that no individual cow appeared in both training and testing subsets. The resulting training subset contained 780 original images (566 healthy and 214 mastitic), and the held-out test subset contained 196 original images (142 healthy and 54 mastitic). After training-only augmentation, the training set was balanced at 2354 images per class (4708 images total).
Classifiers (Model Training and Evaluation)
Ten ML models (decision tree, Gaussian NB, AdaBoost, linear discriminant analysis (LDA), random forest, extra trees, logistic regression, K-nearest neighbors (KNNs), support vector machine (SVM), and multi-layer perceptron (MLP) were used and evaluated on the extracted features (Table 2).
Evaluation Metrics
For each classifier, predictions were evaluated on the held-out test set using the confusion matrix and standard diagnostic-performance metrics. The positive class was mastitis, whereas the negative class was healthy. The confusion matrix was used to derive true positives, true negatives, false positives, and false negatives, from which accuracy, sensitivity, specificity, precision, F1-score, and AUC were calculated.
Each classifier model’s performance was evaluated on the held-out test set using the following predictive metrics [20].
Precision: proportion of positive predictions that are truly mastitis. It measures the total number of true positives (TPs) divided by the total number of predicted positives and was calculated as:
Precision = TP/(TP + FP).
Accuracy (Acc): proportion of correctly classified samples; it measures the total number of correct classifications divided by the total number of cases and was calculated as:
Acc = (TP + TN)/(TP + TN + FP + FN).
Sensitivity (Se): also known as recall, sensitivity is the proportion of true mastitis cases correctly identified, is critical for minimizing missed diagnoses, and was calculated as:
Se = TP/(TP + FN).
Specificity (Sp): specificity is the proportion of true healthy cases correctly identified, as described in the equation:
Sp = TN/(TN + FP).
F1-Score: a single metric that is a harmonic mean of precision and recall, as described in the equation:
F1 Score = (2 × Precision × Se)/(Precision + Se).
Area Under the Receiver Operating Characteristic Curve (AUC-ROC): It measures the discriminative ability across all classification thresholds. The best-performing model was selected based on the highest AUC-ROC score, as AUC is more informative than Acc for imbalanced or threshold-sensitive medical classification tasks.

2.5. Statistical Analysis—Logistic Regression

2.5.1. Rationale for Logistic Regression

Binary logistic regression was selected as the statistical method for investigating the relationship between continuous biomarker variables and binary mastitis outcomes (0 = healthy, 1 = mastitis). Unlike linear regression, logistic regression constrains predicted values to the 0–1 range and models the log-odds of the outcome as a linear function of the predictor, making it the standard method for binary outcome analysis in veterinary and clinical research [36].

2.5.2. SCC vs. CMT and ML Prediction: Simple Logistic Regression

In the first regression analysis, somatic cell count (SCC) was used as the sole predictor variable to model two binary outcomes: (1) CMT result (positive/negative) and (2) ML model prediction (mastitis/healthy). This analysis was intended to assess whether model predictions followed the expected SCC–mastitis relationship rather than establish definitive clinical validation.

2.5.3. SCC Data Cleaning

The SCC values contained non-standard string entries. Values recorded as ‘<90,000’ were replaced with the midpoint estimate of 45,000 cells/mL. Values recorded as ‘>9,000,000’ were assigned the boundary value of 9,000,000 cells/mL. All comma separators were removed before numeric conversion, and remaining non-convertible entries were treated as missing and excluded.

2.5.4. Outcome Encoding

The CMT results containing the keyword ‘positive’ were binary encoded as 1 (mastitis), and all other results as 0 (healthy). ML model predictions containing ‘mastitis’ were encoded as 1 and ‘healthy’ predictions as 0. Encoding was performed using case-insensitive string matching to ensure robustness.

2.5.5. Model Fitting and Visualization

L1-penalized logistic regression was fitted independently for each outcome using statsmodels for p-value estimation and scikit-learn LogisticRegression with the liblinear solver for plotting predicted probabilities. Predicted probabilities were computed over 300 evenly spaced SCC values spanning the observed range to generate smooth curves. p-values were reported as returned by the penalized logistic models without post-hoc forcing or adjustment.

2.5.6. L1-Penalized Logistic Regression: Multi-Variable Analysis

For the multi-variable analysis across blood and milk parameters, L1-penalized (Lasso) logistic regression was applied using the statsmodels library. L1 regularization was chosen over L2 because it performs automatic feature selection by shrinking irrelevant coefficients to zero, which is particularly valuable when analyzing a large panel of potentially correlated biomarkers.

2.5.7. Variables Analyzed

The following independent variables were analyzed against the ML model prediction outcome (Table 3).

2.5.8. Handling Complete Separation

Complete separation is a common problem in logistic regression with small veterinary datasets, occurring when a predictor perfectly separates the two classes, causing maximum likelihood estimates to be infinite and standard errors to be undefined. L1 regularization was employed to constrain coefficient magnitudes and produce finite estimates when possible [37]. The regularization parameter alpha was set to 0.01 for most variables and increased to 0.1 for APP (%) and total protein, which were more prone to separation. Maximum iterations were set at 5000 to support convergence.

2.5.9. Statistical Significance

Statistical significance of each predictor was assessed using the Wald test p-value extracted from the regularized logistic regression results when a stable p-value was returned. A significance threshold of p < 0.05 was applied. The results were reported as significant or not significant without manual p-value adjustment.

2.5.10. Visualization

Logistic regression results were visualized using a p-value summary plot and a separate SCC-based logistic regression comparison between CMT and ML predictions. These visualizations were used to summarize biomarker significance and to assess whether ML predictions followed the expected SCC-related mastitis pattern.

2.6. Model Evaluation Strategy

2.6.1. Internal Validation—Held-Out Test Set

Formal image-level model evaluation was performed on the held-out 20% cow-level test set, which included 196 original images from the total dataset of 976 images. This test set was not augmented and was not used during model training, feature scaling, parameter selection, or classifier optimization. Because the training set was augmented only after cow-level splitting, the held-out test results are interpreted as within-dataset image-level performance.

2.6.2. Animal-Level Stratified Group Cross-Validation

To obtain a more conservative estimate of animal-level generalization, a separate 5-fold stratified group cross-validation was conducted using cow-level grouping. This procedure ensured that all images from the same cow were assigned exclusively to either the training or testing subset within each fold. The cross-validation results were used to evaluate the robustness of model performance under stricter animal-level separation.

2.6.3. Qualitative External Application—Serial Images

The best-performing model was additionally applied to 170 serial thermal udder images as a qualitative external application. The final prediction distribution was summarized as the number and percentage of images classified as mastitis-positive or healthy. Because no independent ground-truth labels were available for these serial images, this analysis was not treated as independent external validation.

3. Results

3.1. Performance of Machine Learning Models

The analysis evaluated 10 machine-learning classifiers using thermal udder-image features extracted from EfficientNetB3. The evaluation was performed on the unaugmented cow-level held-out test set and was supplemented by animal-level stratified group cross-validation.
Table 4 summarizes the confusion-matrix counts for each classifier on the unaugmented held-out test set, which included 196 original images: 142 healthy and 54 mastitic images. The positive class was mastitis. Among the 10 classifiers, the MLP model showed the best overall image-level performance.
On the unaugmented cow-level held-out test set, the MLP achieved the highest image-level accuracy and AUC among the ten classifiers.
Figure 2 shows the confusion matrix of the MLP classifier evaluated on the 20% held-out cow-level test dataset. The confusion matrix showed 129 true negatives, 40 true positives, 13 false positives, and 14 false negatives. For the selected MLP model, 14 mastitic images were misclassified as healthy.
Table 5 summarizes the comparative accuracy and AUC values of the ten machine-learning models on the unaugmented held-out test set. MLP achieved the highest image-level accuracy (0.8622) and AUC (0.9184; bootstrap 95% CI: 0.8740–0.9557), followed by SVM, with an accuracy of 0.8367 and AUC of 0.8963. For the selected MLP model, the bootstrap 95% confidence intervals were 0.8163–0.9082 for accuracy, 0.8740–0.9557 for AUC, 0.6250–0.8519 for sensitivity, 0.8603–0.9524 for specificity, and 0.6457–0.8320 for F1-score.
The ROC analysis confirmed that MLP and SVM were the leading classifiers, with MLP showing the highest AUC; the corresponding AUC values and bootstrap 95% confidence intervals are summarized in Table 5, while the full ROC curves are provided in Supplementary Figure S1. In the additional animal-level 5-fold stratified group cross-validation, the mean AUC was 0.6812 ± 0.1323, with mean precision of 0.6255 ± 0.0799, recall of 0.6725 ± 0.1733, F1-score of 0.6344 ± 0.0726, and accuracy of 0.6244 ± 0.0642.
Since MLP was the best-performing model, it was applied to 170 serial thermal udder images as a qualitative external application. The model classified 90 images (52.9%) as mastitis-positive and 80 images (47.1%) as healthy, as shown in Figure 3. These results are presented as prediction distributions and should not be interpreted as independent external validation.

3.2. Exploratory Biomarker Association Analysis

The exploratory L1-penalized logistic regression analyses did not identify statistically significant associations between the evaluated blood/milk parameters and ML-predicted mastitis status. As summarized in Figure 4 and Table 6, none of the evaluated variables reached the significance threshold of p < 0.05. The smallest p-values were observed for APP (%) (p = 0.16788), LDH (p = 0.19468), BHBA (p = 0.21476), and SCC (p = 0.27087), but none reached statistical significance. IgA and NEFA did not return stable p-values in the final model output and were therefore interpreted conservatively as not significant.
Figure 5 compares the logistic regression relationships between SCC and the traditional California Mastitis Test (CMT) and between SCC and ML predictions. In this analysis, neither association reached statistical significance (CMT: p = 0.88908; ML prediction: p = 0.67885).

3.3. The Impacts of Different Dietary Factors

When the dataset was analyzed by ration/farm grouping, Ration A (Delta Misr) showed a higher observed mastitis incidence than Ration B (Copenhagen). In the final dataset, Ration A included 20 mastitic and 20 healthy cows (20/40; 50.0% incidence), whereas Ration B included 16 mastitic and 29 healthy cows (16/45; 35.6% incidence).
However, logistic regression using nutritional variables did not identify a statistically significant direct nutritional effect on mastitis incidence. Crude protein (p = 0.18256) and TDN (p = 0.16665) were not significant, and NDICP, ADICP, and NEL3X did not yield stable independent p-values. Cow status based on the final ration/farm grouping is shown in Figure 6. The observed mastitis incidence was 50.0% for Ration A (Delta Misr) and 35.6% for Ration B (Copenhagen).
Figure 7 compares the crude protein (CP%) and total digestible nutrients (TDN%) between Ration A and Ration B. Ration A contained a higher crude protein content (16.13%) than Ration B (15.02%), while both rations showed nearly identical TDN values (71.51% and 71.77%, respectively), indicating similar energy availability.
Figure 8 illustrates the relationship between dietary crude protein content and model-predicted mastitis probability for the two ration types. The plot shows different predicted probabilities for Ration A and Ration B, but this relationship was not statistically significant in the logistic regression analysis.
The final nutritional model did not provide evidence that crude protein, TDN, or the other evaluated ration variables independently predicted mastitis status at p < 0.05.

4. Discussion

4.1. Interpretation of Machine-Learning Performance

The present study evaluated infrared thermography combined with EfficientNetB3-based feature extraction and conventional machine-learning classifiers for mastitis screening in dairy cows. Among the ten evaluated classifiers, the MLP model achieved the highest image-level hold-out performance, with an accuracy of 86.22% and an AUC of 0.9184. This finding suggests that nonlinear classifiers can capture discriminative thermal-image patterns associated with mastitis-related udder changes.
The superior performance of MLP may be related to its ability to model nonlinear interactions among the high-dimensional EfficientNetB3 thermal features, whereas simpler classifiers may be less able to capture complex thermal-pattern relationships associated with mastitis. However, these values should be interpreted as hold-out image-level performance rather than independent external validation. The additional animal-level 5-fold stratified group cross-validation yielded a lower mean AUC, indicating that the image-level MLP result should be considered an optimistic within-dataset estimate, whereas the group-based cross-validation provides a more conservative estimate of animal-level generalization to unseen animals.
The confusion matrix of the selected MLP model also indicates that false-negative classifications remain clinically important. For the selected MLP model, 14 mastitic images were misclassified as healthy. These false-negative cases are clinically important because missed mastitis may delay intervention and allow deterioration of udder health. Therefore, despite the relatively high specificity, the sensitivity of 74.07% indicates that the model should be considered a supportive screening tool rather than a standalone diagnostic replacement. Its potential value lies in assisting herd health monitoring and prioritizing cows for further confirmatory testing using established methods such as SCC, CMT, bacteriological culture, or molecular diagnostics.

4.2. Biomarker and SCC-Association Findings

The exploratory biomarker analyses did not identify statistically significant associations between the evaluated blood/milk parameters and ML-predicted mastitis status. Although SCC, immunoglobulins, acute-phase proteins, metabolic indicators, and enzyme markers are biologically relevant to mastitis and inflammatory status, the present dataset did not provide sufficient statistical evidence to confirm biomarker-model concordance. The lack of significance may reflect limited sample size, class structure, biological variability, complete-separation issues in logistic regression, or farm-level confounding.
Figure 5 compares the logistic regression relationships between SCC and the traditional California Mastitis Test (CMT) and between SCC and ML predictions. In this analysis, neither association reached statistical significance (CMT: p = 0.88908; ML prediction: p = 0.67885). Thus, the present dataset does not support the claim that ML predictions have a stronger relationship with SCC than CMT. Rather, these findings suggest that the available sample size and predictor structure were insufficient to establish robust biomarker-model concordance. Therefore, the biomarker findings should be interpreted as exploratory biological contextualization rather than confirmatory validation of the ML model.
Santana et al. [16] evaluated the use of XGBoost for diagnosing bovine subclinical mastitis from udder thermograms collected over 14 months from dairy cows in an automatic milking system. Their model incorporated thermographic, environmental, production, and animal-level variables and achieved an AUC of 0.843, with high specificity. The coldest udder-region temperature was identified as an important predictor. These findings support the potential of combining IRT with machine-learning approaches for mastitis screening, although broader validation remains necessary before routine field application.
Previous studies have also investigated mastitis screening using SCC, EC, pH, behavioral data, and thermal imaging with different machine-learning models. Tian et al. [38] implemented a KNN model with EC and pH inputs. Bobbo et al. [39] achieved an accuracy of 79.7% and a sensitivity of 52.4% using linear discriminant analysis (LDA) for mastitis diagnosis. Recently, Khan et al. [40] developed an SVM classifier based on cow behavior data. Pan et al. [41] used a machine learning-based diagnostic framework integrating logistic regression (LR), support vector machines (SVMs), and feedforward neural networks (FNNs) to evaluate mastitis detection performance with EC, SCC, and their combined inputs. The SVM model achieved the highest accuracy (95.6%) and sensitivity (100%), with SCC as the primary input, while the FNN model delivered the best overall performance with an AUC of 0.981, highlighting its ability to capture complex patterns. These results underscore the value of SCC as a reliable and specific indicator of mastitis, being less affected by non-infectious factors than EC.
Although model performance varied, these studies demonstrate that combining multiple indicators through ML-based methods offers a promising route toward more robust and accurate mastitis screening compared to single-threshold systems. Several variables may affect model accuracy, including large, high-quality annotated datasets for training, data collection under diverse climates, breeds, management practices, and human factors [42]. In agreement with this broader literature, the present study supports the potential of combining IRT with machine-learning models, while also emphasizing that broader validation remains necessary before routine field application.

4.3. Feeding-System Findings and Farm-Level Confounding

The feeding-system analysis showed a higher observed mastitis incidence in Ration A than in Ration B; however, the statistical analysis did not identify a significant direct nutritional predictor of mastitis status. Common indicators used to assess mastitis include SCC, EC, CMT, and milk microbiology analysis [43]. Accurate, rapid, and timely screening tools such as IRT may help characterize udder-health risk and support earlier management decisions, although pathogen identification still requires microbiological or molecular testing.
Although high dietary crude protein has been discussed in relation to metabolic load and mastitis susceptibility, the present analysis does not demonstrate a significant direct effect of crude protein on mastitis risk. Any biological interpretation must therefore remain cautious until tested in a larger dataset with balanced farm, ration, production, and management variables [44,45]. The nutritional composition of the two feeding systems may still be relevant to udder health, but the statistical analysis suggests that the observed farm differences cannot be attributed confidently to ration composition alone.
Farm-level confounders may have contributed to the observed variation, including management practices, environmental conditions, housing, hygiene, milking routine, parity, lactation stage, and sampling differences. Consequently, the feeding-system findings should be interpreted as descriptive and hypothesis-generating rather than evidence of a causal dietary effect. Future studies should use larger, prospectively balanced farm/ration designs to distinguish nutritional effects from broader management and environmental influences.
Previous research has also shown that behavioral and feeding-related indicators may contribute to mastitis monitoring. Analysis using artificial neural networks and logistic regression models has demonstrated that the time spent feeding and resting are significant behavioral indicators for mastitis detection [46]. In addition, mastitis has been associated with reduced feed intake before clinical diagnosis [47]. In this context, IRT and machine-learning approaches may complement other precision-livestock technologies by providing non-invasive thermal information related to udder health.
Recent advances in machine learning and artificial intelligence show potential for automating the inspection, detection, and analysis of thermal images and videos, supporting herd health monitoring processes [48]. IRT cameras can work synergistically with modern machine-learning models to extract thermal details and assess animal-health status [48]. Such automated systems may enable the acquisition and processing of large amounts of data, enhancing efficiency, reducing labor, and improving animal-production practices [49]. However, practical deployment should be considered a potential future application rather than an established outcome of the present pilot study.

4.4. Limitations and Future Work

This study has several limitations. First, mastitis labels were assigned at the cow/whole-udder level rather than at the udder-quarter level, and clinical and subclinical mastitis cases were not analyzed separately. Second, although imaging conditions were standardized as much as possible, exact ambient temperature and humidity ranges were not continuously recorded. Third, the models were trained on BMP thermal image representations rather than raw radiometric temperature matrices. Fourth, the image-level hold-out results should be interpreted as within-dataset performance, while the animal-level group cross-validation provides a more conservative estimate of generalization. Fifth, the 170 serial thermal images were used as a qualitative external application and should not be interpreted as independent external validation.
In addition, no formal model-explainability analysis, such as Grad-CAM or saliency mapping, was performed; therefore, future work should verify that model attention is concentrated on udder thermal regions rather than background or camera-related artifacts. Finally, the biomarker, bacteriological, and nutritional analyses were exploratory and may be confounded by farm-level differences in management, housing, hygiene, milking routine, and environmental conditions. Future studies should prioritize larger labeled multicenter datasets, standardized environmental recording, quarter-level diagnostic labeling, independent external validation, and prospective on-farm testing before routine practical implementation.

5. Conclusions

This study suggests that IRT combined with EfficientNetB3-based feature extraction and ML classification can identify mastitis-related thermal-image patterns in dairy cows within the present dataset. Among the evaluated classifiers, MLP achieved the best image-level performance, whereas animal-level group cross-validation provided a more conservative estimate of generalization. The biomarker and nutritional analyses did not identify statistically significant associations at p < 0.05 and should therefore be interpreted as exploratory. Overall, this work presents a transparent, leakage-aware IRT-ML pipeline for supportive mastitis screening. Larger labeled multicenter datasets and independent external validation are required before routine on-farm implementation.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/vetsci13070640/s1, Table S1: Key technical specifications of the UNI-T UTx313 thermal imaging camera relevant to udder thermal-image acquisition; Table S2: Ingredients of the TMR and chemical composition on a DM basis for lactating dairy cows in the biological-assessment group from Copenhagen and Delta Misr farms. Figure S1: ROC curves for image-level hold-out evaluation of the ten machine-learning models on the unaugmented test set. The curves represent corrected within-dataset performance and should not be interpreted as independent external validation.

Author Contributions

S.M.A.S. and E.E.M.B.: conceptualization and funding acquisition; S.M.A.S., E.E.M.B., A.M.A. and E.A.E.: supervision of the whole study, validation, writing—review and editing; A.S.E. and M.F.A.A.: data collection and curation, investigation, methodology, software and writing—original draft; A.T.E.: data augmentation, modelling, and statistics. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the Deanship of Graduate Studies and Scientific Research at Qassim University (QU-APC-2026).

Institutional Review Board Statement

All procedures were approved and authorized by the Institutional Animal Care and Use Committee of Alexandria University (protocol ID: Alex. Agri. 082410121; approval date: 13 October 2024).

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Materials. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

IRT: infrared thermography; AI: artificial intelligence; ML: machine learning; SLF: smart livestock farming; PLF: precision livestock farming; CMT: California Mastitis Test; SCC: somatic cell count; SCM: subclinical mastitis; DM: dry matter; OM: organic matter; EE: ether extract; CP: crude protein; NDF: neutral detergent fiber; ADF: acid detergent fiber; TMR: total mixed ration; TDNs: total digestible nutrients; EC: electrical conductivity; SNF: solids-not-fat; BHBA: beta hydroxybutyrate; NEFAs: non-esterified fatty acids; LDH: lactate dehydrogenase; BMP: bitmap; CLAHE: Contrast Limited Adaptive Histogram Equalization; GAP: Global Average Pooling; GMP: Global Max Pooling; TP: true positive; TN: true negative; FP: false positive; FN: false negative; Acc: accuracy; Se: sensitivity; Sp: specificity; AUC-ROC: Area Under the Receiver Operating Characteristic Curve; AUC: Area Under the Curve; ROC: Receiver Operating Characteristic Curve; APPs: Acute Phase Proteins; LDA: linear discriminant analysis; KNN: K-nearest neighbors; SVM: support vector machine; MLP: multi-layer perceptron.

References

  1. Velasco-Bolaños, J.; Ceballes-Serrano, C.C.; Velásquez-Mejía, D.; Riaño-Rojas, J.C.; Giraldo, C.E.; Carmona, J.U.; Ceballos-Márquez, A. Application of udder surface temperature by infrared thermography for diagnosis of subclinical mastitis in Holstein cows located in tropical highlands. J. Dairy Sci. 2021, 104, 10310–10323. [Google Scholar] [CrossRef] [PubMed]
  2. Franco-Martínez, L.; Muñoz-Prieto, A.; Contreras-Aguilar, M.D.; Želvytė, R.; Monkevičienė, I.; Horvatić, A.; Kuleš, J.; Mrljak, V.; Cerón, J.J.; Escribano, D. Changes in saliva proteins in cows with mastitis: A proteomic approach. Res. Vet. Sci. 2021, 140, 91–99. [Google Scholar] [CrossRef] [PubMed]
  3. Adkins, P.R.F.; Middleton, J.R. Methods for diagnosing mastitis. Vet. Clin. Food Anim. Pract. 2018, 34, 479–491. [Google Scholar] [CrossRef] [PubMed]
  4. Zigo, F.; Vasil’, M.; Ondrašovičová, S.; Výrostková, J.; Bujok, J.; Pecka-Kielb, E. Maintaining Optimal Mammary Gland Health and Prevention of Mastitis. Front. Vet. Sci. 2021, 8, 607311. [Google Scholar] [CrossRef] [PubMed]
  5. El-Sayed, A.; Kamel, M. Bovine Mastitis Prevention and Control in the Post-Antibiotic Era. Trop. Anim. Health Prod. 2021, 53, 236. [Google Scholar] [CrossRef] [PubMed]
  6. Coşkun, G.; Aytekin, İ. Early Detection of mastitis by using infrared thermography in holstein-friesian dairy cows via classification and regression tree (CART) analysis. Selçuk. J. Agric. Food Sci. 2021, 35, 115–124. [Google Scholar] [CrossRef]
  7. Swami, S.V.; Patil, R.A.; Gadekar, S.D. Studies on the prevalence of subclinical mastitis in dairy animals. J. Entomol. Zool. Stud. 2017, 5, 1297–1300. [Google Scholar]
  8. Duarte, C.M.; Freitas, P.P.; Bexiga, R. Technological advances in bovine mastitis diagnosis: An overview. J. Vet. Diagn. Investig. 2015, 27, 665–672. [Google Scholar] [CrossRef] [PubMed]
  9. Lakshmi, R. Bovine mastitis and its diagnosis. Int. J. Appl. Res. 2016, 2, 213–216. [Google Scholar] [CrossRef]
  10. Araújo, V.M.; Rili, I.; Gisiger, T.; Gambs, S.; Vasseur, E.; Cellier, M.; Diallo, A.B. AI-Powered Cow Detection in Complex Farm Environments. Smart Agric. Technol. 2025, 10, 100770. [Google Scholar] [CrossRef]
  11. Khoei, T.T.; Slimane, H.O.; Kaabouch, N. Deep learning: Systematic review, models, challenges, and research directions. Neural Comput. Appl. 2023, 35, 23103–23124. [Google Scholar] [CrossRef]
  12. Alshehri, M. Blockchain-assisted internet of things framework in smart livestock farming. Internet Things 2023, 22, 100739. [Google Scholar] [CrossRef]
  13. Kok, Z.H.; Mohamed Shariff, A.R.; Alfatni, M.S.M.; Khairunniza-Bejo, S. Support vector machine in precision agriculture: A review. Comput. Electron. Agric. 2021, 191, 106546. [Google Scholar] [CrossRef]
  14. Polat, B.; Colak, A.; Cengiz, M.; Yanmaz, L.E.; Oral, H.; Bastan, A.; Kaya, S.; Hayirli, A. Sensitivity and Specificity of Infrared Thermography in Detection of Subclinical Mastitis in Dairy Cows. J. Dairy Sci. 2010, 93, 3525–3532. [Google Scholar] [CrossRef] [PubMed]
  15. Usamentiaga, R.; Venegas, P.; Guerediaga, J.; Vega, L.; Molleda, J.; Bulnes, F. Infrared Thermography for Temperature Measurement and Non-Destructive Testing. Sensors 2014, 14, 12305–12348. [Google Scholar] [CrossRef] [PubMed]
  16. Santana, R.C.M.; Guimarães, E.d.; Caracuschanski, F.D.; Brassolatti, L.C.; Silva, M.L.d.; Garcia, A.R.; Pezzopane, J.R.M.; Alves, T.C.; Tholon, P.; Santos, M.V.d.; et al. Machine learning techniques associated with infrared thermography to optimize the diagnosis of bovine subclinical mastitis. Vet. Med. Int. 2025, 2025, 5585458. [Google Scholar] [CrossRef] [PubMed]
  17. Colak, A.; Polat, B.; Okumus, Z.; Kaya, M.; Yanmaz, L.E.; Hayirli, A. Short Communication: Early Detection of Mastitis Using Infrared Thermography in Dairy Cows. J. Dairy Sci. 2008, 91, 4244–4248. [Google Scholar] [CrossRef] [PubMed]
  18. Lokhorst, C.; De Mol, R.M.; Kamphuis, C. Invited Review: Big Data in Precision Dairy Farming. Animal 2019, 13, 1519–1528. [Google Scholar] [CrossRef] [PubMed]
  19. Ebrahimie, E.; Ebrahimi, F.; Ebrahimi, M.; Tomlinson, S.; Petrovski, K.R. A large-scale study of indicators of subclinical mastitis in dairy cattle by attribute weighting analysis of milk composition features: Highlighting the predictive power of lactose and electrical conductivity. J. Dairy Res. 2018, 85, 193–200. [Google Scholar] [PubMed]
  20. Zhou, X.; Xu, C.; Wang, H. The Early Prediction of Common Disorders in Dairy Cows Monitored by Automatic Systems with Machine Learning Algorithms. Animals 2022, 12, 1251. [Google Scholar] [CrossRef] [PubMed]
  21. AOAC. Official Methods of Analysis of the Association of Official Analytical Chemists: Official Methods of Analysis of AOAC International, 21st ed.; AOAC: Washington, DC, USA, 2019. [Google Scholar]
  22. Van Soest, P.J.; Robertson, J.B.; Lewis, B.A. Methods for dietary fibre, neutral detergent fibre and nonstarch polysaccharides in relation to animal nutrition. J. Dairy Sci. 1991, 74, 3583–3597. [Google Scholar] [PubMed]
  23. NRC. Nutrient Requirements of Dairy Cattle; National Research Council, National Academy Press: Washington, DC, USA, 2001. [Google Scholar]
  24. ISO 4832; Microbiology of Food and Animal Feeding Stuffs—Horizontal Method for the Enumeration of Coliforms—Colony-Count Technique. The International Organization for Standardization: Geneva, Switzerland, 2006.
  25. ISO 6888-1; Microbiology of the Food Chain—Horizontal Method for Enumerating Coagulase Positive Staphylococci (Staphylococcus aureus and Other Species) Part 1: Method Using Baird-Parker Agar Medium. ISO: Geneva, Switzerland, 2021.
  26. Gornall, A.G.; Bardawill, C.J.; David, M.M. Determination of serum proteins by means of the biuret reaction. J. Biol. Chem. 1949, 177, 751–766. [Google Scholar] [CrossRef]
  27. Doumas, B.T.; Watson, W.A.; Biggs, H.G. Albumin standards and the measurement of serum albumin with bromcresol green. Clin. Chim. Acta 1971, 31, 87–96. [Google Scholar] [CrossRef] [PubMed]
  28. Trinder, P. Determination of blood glucose using an oxidase-peroxidase system with a non-carcinogenic chromogen. J. Clin. Pathol. 1969, 22, 158–161. [Google Scholar] [PubMed]
  29. Taggart, A.K.; Kero, J.; Gan, X.; Cai, T.Q.; Cheng, K.; Ippolito, M.; Ren, N.; Kaplan, R.; Wu, K.; Wu, T.J.; et al. (D)-beta-Hydroxybutyrate inhibits adipocyte lipolysis via the nicotinic acid receptor PUMA-G. J. Biol. Chem. 2005, 280, 26649–26652. [Google Scholar] [PubMed]
  30. Elshafey, B.G.; Elfadadny, A.; Metwally, S.; Saleh, A.G.; Ragab, R.F.; Hamada, R.; Mandour, A.S.; Hendawy, A.O.; Alkazmi, L.; Ogaly, H.A.; et al. Association between biochemical parameters and ultrasonographic measurement for the assessment of hepatic lipidosis in dairy cows. Ital. Ital. J. Anim. Sci. 2023, 22, 136–147. [Google Scholar] [CrossRef]
  31. Girdauskaitė, A.; Grigė, S.; Sabeckienė, I.; Džermeikaitė, K.; Krištolaitytė, J.; Miknienė, Z.; Antanaitis, R. Associations of Blood Lactate Dehydrogenase Activity with Blood Biochemical and Automated Milk Monitoring Parameters in Early-Lactation Dairy Cows. Agriculture 2026, 16, 502. [Google Scholar] [CrossRef]
  32. Ježek, J.; Malovrh, T.; Klinkon, M. Serum immunoglobulin (IgG, IgM, IgA) concentration in cows and their calves. Acta Agric. Slov. Supl. 2012, 100, 295–298. [Google Scholar]
  33. Zuiderveld, K. Contrast limited adaptive histogram equalization. In Graphics Gems IV; Academic Press: Cambridge, MA, USA, 1994; pp. 474–485. [Google Scholar]
  34. Tan, M.; Le, Q.V. EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), Beach, CA, USA, 9–15 June 2019. [Google Scholar]
  35. Raghu, M.; Zhang, C.; Kleinberg, J.; Bengio, S. Transfusion: Understanding transfer learning for medical imaging. arXiv 2019, arXiv:1902.07208. [Google Scholar]
  36. Hosmer, D.W.; Lemeshow, S. Applied Logistic Regression, 2nd ed.; John Wiley & Sons: Hoboken, NJ, USA, 2000. [Google Scholar]
  37. Firth, D. Bias reduction of maximum likelihood estimates. Biometrika 1993, 80, 27–38. [Google Scholar] [CrossRef]
  38. Tian, F.; Wang, Z.; Yu, S.; Xiong, B.; Wang, S. Clinical mastitis detection by on-line measurements of milk yield, electrical conductivity and deep learn. J. Phys. Conf. Ser. 2020, 1635, 012046. [Google Scholar] [CrossRef]
  39. Bobbo, T.; Biffani, S.; Taccioli, C.; Penasa, M.; Cassandro, M. Comparison of machine learning methods to predict udder health status based on somatic cell counts in dairy cows. Sci. Rep. 2021, 11, 13642. [Google Scholar] [CrossRef] [PubMed]
  40. Khan, M.F.; Thorup, V.M.; Luo, Z. Delineating mastitis cases in dairy cows: Development of an IoT-enabled intelligent decision support system for dairy farms. IEEE Trans. Ind. Inform. 2024, 20, 9508–9517. [Google Scholar] [CrossRef]
  41. Pan, L.; Chen, X.; Han, D.; Li, N.; Chen, D.; Wang, J.; Chen, J.; Huo, X. Machine learning-based clinical mastitis detection in dairy cows using milk electrical conductivity and somatic cell count. Front. Vet. Sci. 2025, 12, 1671186. [Google Scholar] [CrossRef] [PubMed]
  42. Asogan, A.; Sazali, N.; Veerendra, A.S.; Samylingam, L.; Aslfattahi, N.; Kok, C.K.; Kadirgama, K. A review on the impact of AI-enabled thermal imaging and IoT sensor fusion on early detection of mastitis in dairy cattle. Biosens. Bioelectron. 2026, 28, 100735. [Google Scholar] [CrossRef]
  43. Bobbo, T.; Matera, R.; Biffani, S.; Gómez, M.; Cimmino, R.; Pedota, G.; Neglia, G. Exploring the sources of variation of electrical conductivity and total and differential somatic cell count in Italian Mediterranean buffaloes. J. Dairy Sci. 2024, 107, 508–515. [Google Scholar] [CrossRef] [PubMed]
  44. Dreyer, C.B.; Losand, H.; Spiekers, H.; Hummel, J. Influence of fat-to-protein ratio and udder health parameters on the milk urea content of dairy cows. J. Dairy Sci. 2025, 108, 2527–2546. [Google Scholar] [PubMed]
  45. Zeleke, A.W.; Dimonaco, N.J.; Lawther, K.; Lavery, A.; Ferris, C.; Moorby, J.; Huws, S.A. Reducing crude protein content in the diet of lactating dairy cows improved nitrogen-use-efficiency and reduced N excretion in urine, whilst having no obvious effects on the rumen microbiome. J. Anim. Sci. Biotechnol. 2025, 16, 113. [Google Scholar] [PubMed]
  46. Grodkowski, G.; Szwaczkowski, T.; Krzysztof Koszela, K.; Wojciech Mueller, W.; Tomaszyk, K.; Ton Baars, T.; Sakowski, T. Early detection of mastitis in cows using the system based on 3D motions detectors. Sci. Rep. 2022, 12, 21215. [Google Scholar] [CrossRef] [PubMed]
  47. Sepúlveda-Varas, P.; Proudfoot, K.L.; Weary, D.M.; von Keyserlingk, M.A.G. Changes in behaviour of dairy cows with clinical mastitis. Appl. Anim. Behav. Sci. 2016, 175, 8–13. [Google Scholar] [CrossRef]
  48. Wang, M.; Tan, H.; Li, Y.; Chen, X.; Chen, D.; Wang, J.; Chen, J. Toward five-part differential of leukocytes based on electrical impedances of single cells and neural network. Cytom. Part A 2023, 103, 439–446. [Google Scholar]
  49. Pacheco, V.M.; de Sousa, R.V.; da Silva Rodrigues, A.V.; de Souza Sardinha, E.J.; Martello, L.S. Thermal Imaging Combined with Predictive Machine Learning Based Model for the Development of Thermal Stress Level Classifiers. Livest. Sci. 2020, 241, 104244. [Google Scholar] [CrossRef]
Figure 1. Dataset flow and model-evaluation workflow showing cow-level splitting before augmentation, training-only augmentation, held-out testing, animal-level stratified group cross-validation, and qualitative external application of the selected model.
Figure 1. Dataset flow and model-evaluation workflow showing cow-level splitting before augmentation, training-only augmentation, held-out testing, animal-level stratified group cross-validation, and qualitative external application of the selected model.
Vetsci 13 00640 g001
Figure 2. Confusion matrix of the MLP classifier in the 20% cow-level held-out test set.
Figure 2. Confusion matrix of the MLP classifier in the 20% cow-level held-out test set.
Vetsci 13 00640 g002
Figure 3. Prediction distribution for 170 serial thermal udder images used as a qualitative external application.
Figure 3. Prediction distribution for 170 serial thermal udder images used as a qualitative external application.
Vetsci 13 00640 g003
Figure 4. Biomarker p-values from L1-penalized logistic regression. No evaluated biomarker reached p < 0.05.
Figure 4. Biomarker p-values from L1-penalized logistic regression. No evaluated biomarker reached p < 0.05.
Vetsci 13 00640 g004
Figure 5. Logistic regression relationships between somatic cell count (SCC), CMT, and ML prediction. Neither CMT (p = 0.88908) nor ML prediction (p = 0.67885) showed a statistically significant SCC association.
Figure 5. Logistic regression relationships between somatic cell count (SCC), CMT, and ML prediction. Neither CMT (p = 0.88908) nor ML prediction (p = 0.67885) showed a statistically significant SCC association.
Vetsci 13 00640 g005
Figure 6. Cow status by ration/farm group: (a) absolute counts and (b) observed mastitis incidence rate.
Figure 6. Cow status by ration/farm group: (a) absolute counts and (b) observed mastitis incidence rate.
Vetsci 13 00640 g006
Figure 7. Comparison of (a) TDN and (b) crude protein with different ration types.
Figure 7. Comparison of (a) TDN and (b) crude protein with different ration types.
Vetsci 13 00640 g007
Figure 8. Relationship between crude protein and predicted mastitis probability.
Figure 8. Relationship between crude protein and predicted mastitis probability.
Vetsci 13 00640 g008
Table 1. Image dataset characteristics and preprocessing parameters.
Table 1. Image dataset characteristics and preprocessing parameters.
ParameterValue
Image formatBMP (Bitmap)
Input resolution224 × 224 pixels
Color spaceBGR → LAB (CLAHE on L channel) → BGR
ClassesHealthy, Mastitis
Original images976 (708 healthy, 268 mastitis)
Training split before augmentation780 images (566 healthy, 214 mastitis)
Held-out test split196 original images (142 healthy, 54 mastitis)
Training set after augmentation4708 images (2354 per class)
Table 2. Machine-learning classifiers and their hyperparameters.
Table 2. Machine-learning classifiers and their hyperparameters.
ClassifierTypeKey Parameters
Logistic RegressionLinearmax_iter = 2000, L2 penalty
Random ForestEnsemble (Bagging)n_estimators = 300
Extra TreesEnsemble (Bagging)n_estimators = 300
AdaBoostEnsemble (Boosting)Default (SAMME.R)
SVM (RBF)Kernel-basedkernel = rbf, probability = true
KNNInstance-basedk = 7, Euclidean distance
MLPNeural Networklayers = (256,128), max_iter = 500
Decision TreeTree-basedGini impurity
Gaussian Naive BayesProbabilisticGaussian likelihood
LDADiscriminant analysisSVD solver
Table 3. Milk and blood parameters used for evaluating dairy cow health.
Table 3. Milk and blood parameters used for evaluating dairy cow health.
VariableUnitCategory
SCCcells/mLMilk Quality
IgGmg/100 mLImmunoglobulin
IgAmg/100 mLImmunoglobulin
IgEIU/mLImmunoglobulin
APP%Acute Phase Protein
BHBAmmol/LMetabolic
NEFAµmol/LMetabolic
LDHIU/LEnzyme
Glucosemmol/LMetabolic
Albuming/dLProtein
Globuling/dLProtein
Total Proteing/dLProtein
Fat%Milk Composition
SNF%Milk Composition
Protein%Milk Composition
Lactose%Milk Composition
ECmS/cmElectrical Conductivity
Densityg/mLMilk Physical
Table 4. Confusion-matrix counts of the 10 machine-learning classifiers on the unaugmented held-out test set.
Table 4. Confusion-matrix counts of the 10 machine-learning classifiers on the unaugmented held-out test set.
ModelTPFPTNFN
Logistic Regression382411816
Random Forest261912328
Extra Trees272012227
AdaBoost312511723
SVM (RBF)411912313
KNN351812419
MLP401312914
Decision Tree323610622
Gaussian NB212012233
LDA333111121
Table 5. Performance of the 10 machine-learning models on the unaugmented held-out test set.
Table 5. Performance of the 10 machine-learning models on the unaugmented held-out test set.
ModelCasePrecisionRecallF1-ScoreSupportAccAUCAUC 95% CI
Logistic RegressionHealthy0.880.830.861420.79590.86330.8033–0.9185
Mastitis0.610.700.6654
Random ForestHealthy0.810.870.841420.76020.83030.7734–0.8823
Mastitis0.580.480.5354
Extra TreesHealthy0.820.860.841420.76020.84830.7928–0.8968
Mastitis0.570.500.5354
AdaBoostHealthy0.840.820.831420.75510.77840.7076–0.8459
Mastitis0.550.570.5654
SVM (RBF)Healthy0.900.870.881420.83670.89630.8425–0.9408
Mastitis0.680.760.7254
KNNHealthy0.870.870.871420.81120.84390.7798–0.9018
Mastitis0.660.650.6554
MLPHealthy0.900.910.911420.86220.91840.8740–0.9557
Mastitis0.750.740.7554
Decision TreeHealthy0.830.750.791420.70410.66950.5869–0.7461
Mastitis0.470.590.5254
Gaussian NBHealthy0.790.860.821420.72960.81460.7531–0.8689
Mastitis0.510.390.4454
LDAHealthy0.840.780.811420.73470.75180.6720–0.8253
Mastitis0.520.610.5654
Table 6. Statistical significance of biomarkers associated with mastitis prediction.
Table 6. Statistical significance of biomarkers associated with mastitis prediction.
Biomarkerp-ValueInterpretation
APP 0.16788Not significant
LDH (IU/L)0.19468Not significant
BHBA (mmol/L)0.21476Not significant
SCC (cells/mL)0.27087Not significant
IgE (IU/mL)0.53880Not significant
Globulin (g/dL)0.56715Not significant
Albumin (g/dL)0.67739Not significant
Total protein (g/dL)0.73666Not significant
Glucose (mmol/L)0.83311Not significant
IgG (mg/100 mL)0.86934Not significant
IgA (mg/100 mL)N/ANot significant; stable p-value not returned
NEFA (µmol/L)N/ANot significant; stable p-value not returned
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Elmasry, A.S.; Elwakeel, E.A.; Allam, A.M.; Attia, M.F.A.; Elmaria, A.T.; Badr, E.E.M.; Sallam, S.M.A. Infrared Thermography and Machine Learning for Mastitis Detection in Dairy Cows: A Pilot Case Study in Egyptian Farms. Vet. Sci. 2026, 13, 640. https://doi.org/10.3390/vetsci13070640

AMA Style

Elmasry AS, Elwakeel EA, Allam AM, Attia MFA, Elmaria AT, Badr EEM, Sallam SMA. Infrared Thermography and Machine Learning for Mastitis Detection in Dairy Cows: A Pilot Case Study in Egyptian Farms. Veterinary Sciences. 2026; 13(7):640. https://doi.org/10.3390/vetsci13070640

Chicago/Turabian Style

Elmasry, Aya S., Eman A. Elwakeel, Ali M. Allam, Marwa F. A. Attia, Alaa. T. Elmaria, Elsayed. E. M. Badr, and Sobhy M. A. Sallam. 2026. "Infrared Thermography and Machine Learning for Mastitis Detection in Dairy Cows: A Pilot Case Study in Egyptian Farms" Veterinary Sciences 13, no. 7: 640. https://doi.org/10.3390/vetsci13070640

APA Style

Elmasry, A. S., Elwakeel, E. A., Allam, A. M., Attia, M. F. A., Elmaria, A. T., Badr, E. E. M., & Sallam, S. M. A. (2026). Infrared Thermography and Machine Learning for Mastitis Detection in Dairy Cows: A Pilot Case Study in Egyptian Farms. Veterinary Sciences, 13(7), 640. https://doi.org/10.3390/vetsci13070640

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop