1. Introduction
Automatic milking systems (AMSs) are widely used on dairy farms, where they influence work organization, cow behavior, and production monitoring [
1]. However, the development of this technology leads to a significant change in the nature of the interaction between the cow and the production environment, which means that the traditional criteria for assessing the performance value of cows may not fully reflect their functioning under automated milking conditions [
2]. It should also be noted that the efficiency and functionality of AMSs may differ depending on the technical characteristics of robotic systems, including teat detection accuracy, sensor technologies, robotic arm design, and software algorithms supporting the milking process. These differences may affect both milking performance and the adaptation of cows to AMSs.
The performance of cows in AMSs is determined by complex interactions between physiological, morphological, and behavioral factors [
3,
4]. Particular importance is attributed to characteristics describing the milking process, such as the milk flow rate, milking duration, and robot utilization efficiency [
1,
5]. At the same time, a key factor determining the efficiency of the AMS remains the process of attaching the teat cups, the course of which depends on both the animal’s characteristics and its interaction with the technological system [
6].
Conformation traits, particularly udder structure, have long been recognized as key factors in assessing the production value of dairy cows, influencing milk yield, udder health, and longevity [
7]. In the context of AMSs, however, their significance is more complex, as it also encompasses interaction with the technological system, including the ability to correctly locate teats and the process of automatic attachment of teat cups [
8].
Previous research on cow performance in AMSs has focused primarily on the analysis of individual traits, such as milk yield, milking characteristics, or selected behavioral parameters [
6,
9]. Attempts to integrate this information for a comprehensive assessment of cows’ adaptation to AMSs have been much less frequent, even though this process is multidimensional [
10]. As a result, there is still a lack of synthetic indicators integrating different aspects of the milking process, as well as studies comparing the usefulness of statistical and machine learning approaches in the assessment of cows’ adaptation to AMSs [
11]. In particular, there is a lack of synthetic indices combining various aspects of the milking process and studies evaluating their utility in predictive modeling [
12].
The rapid development of AMSs enables the collection of large sets of production data, which opens up new opportunities for applying advanced data analysis methods, including machine learning techniques [
13,
14]. These methods allow for the modeling of complex, nonlinear relationships between animal traits and their performance in production systems, as well as the identification of key predictors of the analyzed phenomena. For example, Aerts et al. [
5] used decision tree algorithms to model milking efficiency in an AMS, while Dharejo et al. [
15] applied machine learning models to predict the occurrence of mastitis based on data from AMSs, achieving high prediction accuracy. Literature reviews also point to the growing importance of modeling and machine learning methods in the analysis of data from AMSs [
16].
Despite the growing number of studies utilizing machine learning methods, their advantage over classical statistical methods in the analysis of data from AMSs is not clear cut, particularly regarding model interpretability and the identification of biologically relevant predictors [
11,
17]. Therefore, there is a need for the systematic evaluation of the usefulness of both analytical approaches, with particular emphasis on their predictive power and the interpretability of results [
11,
18].
The objectives of this study were four-fold: (1) to identify factors influencing the performance of dairy cows in an AMS, (2) to construct a synthetic index of cows’ adaptation to the AMS (i.e., the robotic adaptability index, RAI) based on principal component analysis (PCA), (3) to evaluate the predictive capabilities of traits describing the milking process and the RAI using statistical methods and machine learning methods, and (4) to compare the predictive power and interpretability of different modeling approaches.
This study employed a comparative approach involving multiple linear regression, least absolute shrinkage and selection operator (LASSO) regression, and selected machine learning algorithms (random forest, extreme gradient boosting—XGBoost—and artificial neural networks), which made it possible to assess the extent to which increasing model complexity translates into improved prediction quality in the analysis of data from AMSs. The objectives related to identifying factors affecting cow performance in AMSs and evaluating the predictive usefulness of different modeling approaches had primarily a scientific and cognitive character. In contrast, the construction of the RAI and the evaluation of predictive models for assessing cow functioning in AMSs had a practical and application-oriented dimension.
2. Materials and Methods
2.1. Animals and Data
The study was conducted on a dataset comprising 796 primiparous Polish Holstein–Friesian cows maintained on seven commercial dairy farms located in central and eastern Poland (Mazowieckie, Wielkopolskie, Lubelskie, and Podlaskie regions) and equipped with the Lely Astronaut A4 automatic milking system (Lely Industries N.V., Maassluis, The Netherlands). The data were obtained from the milking robot management system and included 40,233 individual milkings. Data were collected over a 12-month period.
The analyzed farms differed in herd size and the number of cows assigned per milking robot, reflecting typical commercial AMS production conditions in Poland. The number of cows assigned per AMS unit ranged from 55 to 58 animals, indicating a comparable workload among the analyzed milking robots. Cows were housed in free-stall barn systems and managed according to standard commercial dairy farming practices. Cows did not have access to pasture or outdoor paddocks during the data collection period. Animals had permanent access to water and were fed a total mixed ration (TMR) balanced according to their nutritional requirements. Resting areas were equipped with bedding materials commonly used in commercial dairy production systems. All farms operated in accordance with national and European Union regulations concerning dairy cattle welfare.
Given the study’s objective, which was to assess the technical adaptation of cows to AMSs at the individual level, data from individual milkings were aggregated to the cow level by calculating average values for specific milking process parameters (
Table 1). As a result, each observation in the analysis corresponded to a single cow. This approach also resulted from the fact that conformation traits were recorded once for each cow, whereas AMS-derived variables represented repeated measurements. Without aggregation, the dataset would contain multiple AMS records linked to a single conformation score, leading to pseudo-replication and violation of the independence assumption. Although this approach reduced temporal and within-cow variability, it enabled stable long-term characterization of individual cows’ performance in AMSs.
The dataset included the cow ID, farm code, and a set of traits describing the milking process in the AMS, as well as functional and conformation traits.
The cows’ adaptation to the AMS was assessed based on three variables describing the milking process:
MilkingEfficiency (ME)—operationally defined as daily milk yield divided by the total box time (time spent by a cow per day in the AMS milking box, [kg/min]).
AverageNumber (AA)—the average number of attempts to attach the teat cups per milking session [n/milking].
AverageTime (AT)—the average time required to attach the teat cups during a single milking [s/milking].
Higher ME values indicated a more efficient milking process, while higher AA and AT values indicated a more difficult process of attaching the teat cups.
The analysis also included variables describing the cows’ physiological conditions and the way the milking robot was used (functional characteristics):
DIM—number of days in milk [day].
MilkingFreq—average number of milkings per day [n/day].
MilkFlow—milk flow rate [kg/min].
The analysis also included variables describing cow–robot interaction and behavioral responses in the AMS, including the following:
Refusal—number of refusals to enter the milking robot [n/day].
Failure—number of failed milking attempts [n/day].
Linear conformation scores describing the cows’ body structures and udders were also determined and included in the analysis:
HeightSacrum—height at the sacrum [cm].
BodyDepth—body depth [pts].
ChestWidth—chest width [pts].
RumpAngle—rump angle [pts].
RumpWidth—rump width [pts].
RearLegsSet—rear leg conformation (side view) [pts].
FootAngle—foot angle [pts].
RearLegsRearView—rear leg alignment (rear view) [pts].
ForeAttachment—fore udder attachment [pts].
RearUdderHeight—rear udder attachment height [pts].
CentralLigament—central udder ligament [pts].
UdderDepth—udder depth [pts].
RearUdderWidth—width of the rear part of the udder [pts].
RearTeatPlacement—rear teat placement [pts].
FrontTeatPlacement—front teat placement [pts].
TeatLength—teat length [pts].
Angularity—angularity [pts].
These traits were evaluated using a standard linear scale applied in the conformation evaluation of dairy cows. The conformation assessment was carried out by trained experts from the Polish Federation of Cattle Breeders and Dairy Farmers, in accordance with ICAR guidelines.
2.2. Development of an Adaptation Index for the AMS
The adaptation of cows to an AMS is a multifactorial process that encompasses both milking efficiency and the course of the teat cup attachment procedure. Therefore, the three variables describing the milking process (ME, AA, and AT) were treated as different aspects of the same biological phenomenon.
PCA was used to integrate the information contained in these variables. The analysis was performed on a correlation matrix, which means that all variables were standardized beforehand. Standardization was performed according to the following formula:
where
—the trait value for the i-th cow;
—the population mean (n = 796);
—the standard deviation of the trait.
The first principal component (PC) was used to construct the RAI, which was defined as a linear combination of standardized variables:
where
- ■
ZME—standardized ME value;
- ■
ZATT—standardized AA value;
- ■
ZTIME—standardized AT value.
Higher index values indicated greater milking efficiency, fewer attempts to attach the teat cups, and shorter attachment time, and, thus, the better adaptation of the cow to the AMS.
To improve breeding interpretability, the index was subsequently rescaled to a range of 0–100 according to the following formula:
where
RAImin and
RAImax denote the minimum and maximum values of the index in the analyzed population, respectively.
2.3. Predictive Modeling
To assess the predictive capabilities of traits describing cow performance in an AMS, an approach based on predictive modeling was employed. The analysis was conducted for three traits describing the milking process, ME, AA, and AT, as well as for the RAI constructed based on them.
The functional traits of cows and conformation traits describing the body and udder structure were used as explanatory variables.
To compare the predictive power of various analytical methods, models of increasing complexity were used.
2.3.1. Multiple Linear Regression
Multiple linear regression was applied as a reference model to describe the linear relationships between the analyzed variables. The model took the following form:
where
Yi denotes the dependent variable under analysis (ME, AA, AT, or
RAI),
X1i…
Xki denote the explanatory variables,
β0 is the intercept,
βi is the regression coefficient, and
ε is the random error.
2.3.2. LASSO Regression
Least absolute shrinkage and selection operator (LASSO) regression was used to select the most important predictors and to mitigate the problem of multicollinearity. The method involves estimating regression coefficients with an additional penalty term (λ∑∣βj∣), which leads to the reduction in some coefficients to zero (βj denotes the regression coefficient for the jth explanatory variable, while λ is the regularization parameter controlling the strength of the penalty imposed on the regression coefficient values; the value of the regularization parameter λ was determined automatically based on cross-validation).
2.3.3. Random Forest
Random forest (RF) is an ensemble algorithm based on the construction of multiple decision trees created on random subsamples of the data. The final prediction is the average of the predictions from all the trees included in the model. The number of randomly selected predictors for each split in a single tree was five, and the proportion of randomly selected cases for each tree was 50%. In addition, 30% of the cases from the training set were randomly assigned to the validation set, which was used to determine the optimal number of component trees and prevent overfitting, while the maximum number of component trees was 1000. The minimum size of the parent and child nodes was five, with this parameter controlling the minimum number of cases in a tree node required for its splitting and thus limiting the tree’s complexity (the number of nodes and levels). New component trees were added to the model until the mean squared error on the validation set did not fall below 5.0% over ten consecutive iterations (of adding new component trees to the model) [
19].
2.3.4. Extreme Gradient Boosting
Extreme gradient boosting (XGBoost) is a boosting algorithm based on the sequential construction of decision trees that minimize the loss function through gradient optimization. The learning rate, which determines the weight with which component trees were added to the model, was 0.1, while the proportion of randomly selected cases for each component tree was 50.0%. Furthermore, 30.0% of the cases from the training set were randomly assigned to the validation set, and the maximum number of component trees was 1000. The minimum size of the parent and child nodes was five and one, respectively, while the maximum number of nodes in a component tree was three. New trees were added to the model until the lowest mean polynomial error on the validation set was achieved [
19]. In both tree-based methods, the relative importance of predictors was determined based on the total reduction in prediction error (e.g., MSE) obtained as a result of splits using a given variable, summed across all nodes and trees, and then expressed relative to the maximum value.
2.3.5. Artificial Neural Networks
The Broyden–Fletcher–Goldfarb–Shanno algorithm was used to train the multilayer perceptron (MLP) [
20]. Statistica’s automatic designer enabled the automatic selection of the optimal network structure and parameters (the number of neurons in the hidden layer, type of postsynaptic potential function and activation function, number of training epochs, etc.). All networks were trained until the lowest possible error (sum of squares) was achieved on the validation set (a portion of the entire dataset used to prevent overfitting). The relative importance of predictors was determined based on the error magnitude for the network with the predictor removed. The predictors were then sorted by the value of the error ratios (error for the model with the predictor removed relative to the full model).
For each of the analyzed dependent variables (ME, AA, AT, and RAI), all of the aforementioned models were fitted using the same set of explanatory variables.
To evaluate the predictive ability of the models, 10-fold cross-validation was used. The data were divided into ten subsets, nine of which were used to train the model, and one for testing. The procedure was repeated ten times so that each subset served as the test set exactly once.
The predictive ability of the models was assessed based on the coefficient of determination (R2), the root mean square error (RMSE), and the mean absolute error (MAE).
The models were compared in terms of prediction accuracy to determine which approach best described the relationships between the analyzed features.
For the multiple regression models, standardized β coefficients are presented, whereas for the LASSO models, unstandardized coefficients are shown, which resulted from the different assumptions and methods of parameter estimation in both approaches.
4. Discussion
The aggregation of individual milking data to the cow level enabled the characterization of stable, long-term differences between animals and reduced the influence of short-term variability related to lactation stage or environmental conditions. Such an approach is commonly used in studies focused on the general functioning of cows in AMSs and facilitates the comparison of animals at the individual level [
1,
3].
At the same time, it should be emphasized that data aggregation leads to the loss of information regarding within-cow variability and does not account for the hierarchical structure of the data (individual milkings nested within cows). Consequently, the obtained results describe the average level of cow performance rather than the temporal dynamics of the milking process [
21]. Similar limitations have been highlighted in studies based on AMS data, which indicate the importance of longitudinal and multilevel approaches for a more comprehensive analysis of production processes and cow behavior over time [
14,
16]. Nevertheless, the choice of this analytical approach was consistent with the primary objective of the present study, which was to evaluate long-term adaptation of cows to AMSs at the individual level rather than to model variability between single milking events. Future studies should therefore incorporate longitudinal or mixed-effects modeling approaches that allow simultaneous consideration of within- and between-cow variability.
One of the main goals of this study was to construct a synthetic index describing cow adaptation to AMSs (RAI). Principal component analysis enabled dimensionality reduction and the integration of information contained in three variables describing the milking process. The results indicate that the RAI was determined primarily by variables related to the teat cup attachment process (AA and AT), whereas the contribution of milking efficiency (ME) was relatively smaller. Therefore, the proposed index should be interpreted mainly as an indicator of technical adaptation to AMSs associated with the efficiency and stability of teat cup attachment rather than as a universal measure of technical adaptation to robotic milking systems.
Attempts have been made in the literature to synthetically assess cow performance in AMSs, including milking efficiency indicators [
5], characteristics of the milking process, and the classification of animals according to their performance in robotic milking systems [
6]. However, these approaches have mainly focused on the analysis of individual variables or their modeling, whereas integrated indicators describing the multidimensional nature of cow adaptation to AMSs remain relatively rare in the literature. From a practical perspective, this aspect is particularly important because problems related to teat cup attachment are among the major factors limiting the efficiency and functionality of AMSs [
5,
6].
The obtained results of the predictive models indicate a clear difference in the modeling ability of the analyzed traits. In the case of milking efficiency (ME), all the approaches used, both statistical and machine learning-based, were characterized by high prediction performance. This suggests that the relationships determining this trait are largely linear and can be effectively modeled using classical statistical methods [
13,
22]. A different situation was observed for the variables describing the teat cup attachment process (AA and AT) and RAI. The very high correlation observed between AA and AT additionally confirms that both variables describe closely related aspects of the teat cup attachment procedure, which justified their integration using PCA.
Prolonged teat cup attachment time may negatively affect milking dynamics and the neurohormonal mechanisms associated with milk ejection, which may consequently reduce the efficiency of the milking process. The observed variability of attachment-related traits may therefore reflect differences in cow morphology, teat placement, udder conformation, and cow–robot interaction during milking.
The highest predictive ability for these traits was obtained for the LASSO model, while more complex machine learning models (RF, XGBoost, MLP) did not demonstrate a clear advantage. These results are consistent with observations presented in the literature, which emphasize that in the case of data with moderate size and high multicollinearity of variables, regularization methods can outperform more complex machine learning algorithms [
14].
The lack of a clear advantage of machine learning models over statistical methods may be due to the characteristics of the analyzed dataset. First, the number of observations (
n = 796) was relatively small in terms of the use of nonlinear models, which require larger datasets for effective training and capturing complex relationships [
23]. Second, the explanatory variables exhibited significant multicollinearity, favoring the use of regularization methods such as LASSO regression, which enables the simultaneous selection of predictors and stabilization of parameter estimates [
24]. Third, the analysis showed that the relationships between the analyzed variables were largely linear or quasi-linear, limiting the potential advantage of more complex machine learning models. Under such conditions, simpler statistical models can achieve comparable or even higher prediction performance while maintaining greater interpretability [
22]. Similar conclusions are presented in reviews on the application of machine learning in agriculture, indicating that the advantage of machine learning models is not universal and strongly depends on the characteristics of the data [
13,
16].
In addition, the relatively small number of cows and the use of data originating from a limited number of farms may have limited the ability of more complex nonlinear models, such as MLP and XGBoost, to fully exploit their predictive potential. Although the dataset included more than 40,000 individual milking events, aggregation to the cow level reduced the effective number of observations used for model training. Moreover, the absence of an independent external test dataset means that the obtained results should be interpreted primarily as an evaluation of the relative performance of the analyzed models rather than their absolute generalization ability. Future studies should therefore include inter-farm validation using larger and more heterogeneous datasets differing in herd management and AMS configuration.
It should also be noted that only primiparous cows were included in the analysis, which reduced the biological variability associated with parity but may limit the direct applicability of the results to multiparous animals.
An additional limitation of the present study is the lack of information regarding udder health parameters, particularly somatic cell count (SCC), which is an important indicator of milk quality, mastitis risk, and cow functioning in AMSs. The primary objective of the present study was focused on modeling technical and functional aspects of cow adaptation to AMSs; therefore, health-related variables were not included in the predictive analyses. Future studies should integrate production, behavioral, and udder health data to provide a more comprehensive assessment of cow performance in AMSs.
Predictor importance analysis clearly demonstrated that functional variables, particularly milking speed (MilkFlow) and the number of unsuccessful milking attempts (Failures), played a key role in determining milking performance in an AMS. These results are consistent with previous studies indicating that cow behavior and their interaction with the AMS are fundamental to milking efficiency [
1,
9].
The significance of conformation traits also proved limited, suggesting that functional variables and cow–robot interaction indicators (Refusal, Failure) played a greater role in AMS performance than classical morphological traits.
The relatively limited importance of conformation traits observed in the present study may result from both biological and technological factors. Modern AMSs are increasingly capable of compensating for moderate morphological differences between cows due to improvements in sensor precision, teat localization systems, and robotic arm functionality [
3,
11]. Consequently, moderate variation in udder and body conformation may currently have a smaller impact on milking performance than in earlier generations of robotic milking systems.
In contrast, functional and behavioral traits associated with cow–robot interaction, such as milk flow characteristics, milking frequency, and unsuccessful milking attempts, may better reflect the actual efficiency of cow adaptation to AMSs under commercial production conditions [
1,
5]. This may explain why variables directly describing cow performance within the AMS environment showed greater predictive power than classical conformation traits traditionally considered important in dairy cattle evaluation.
These findings are consistent with recent studies demonstrating the growing importance of behavioral and sensor-derived data in the analysis of cow performance within precision livestock farming systems [
8,
14]. The methods used to assess predictor importance allowed the identification of variables with the greatest predictive power; however, they did not provide a complete interpretation of the direction and nature of relationships between variables. This limitation is typical of more complex machine learning models, which often achieve high predictive performance at the expense of interpretability [
25,
26]. In breeding applications, where understanding biological mechanisms is crucial, this may limit their practical usefulness.
It should also be noted that the current analysis did not use an independent test set, and the assessment of the models’ predictive ability was based solely on 10-fold cross-validation. Although this method is widely used and allows for effective utilization of available data [
23], it may lead to optimistic estimates of predictive performance, particularly in more complex models [
22]. Therefore, future studies should include independent external validation datasets to enable a more reliable assessment of model generalization ability [
14,
16]. Finally, it should be mentioned that only primiparous cows were included in the analysis, which reduced the biological variability associated with parity but may limit the direct applicability of the results to multiparous animals.