1. Introduction
Many studies have documented associations between cows’ physiological status, environment, genetics, and management factors and their reproductive outcomes (briefly reviewed in Shahinfar et al. [
1] and Hempstalk et al. [
2], for example). Quantified associations between health-related events (HRE), relative estrus intensity (REI), and probability of pregnancy per artificial insemination (P/AI) can aid in decision-making of insemination decisions for individual or groups of cows. Examples where predicted P/AI may be useful are the choice of the voluntary waiting period, hormonal intervention, type of semen to use, do-not-breed decisions, and removal of cows [
1,
2,
3,
4]. Recently, individual predictions of P/AI have been a component of targeted reproductive management which aims to improve the reproductive efficiency of (non-organic) dairy cows by applying customized group-level management strategies based on hormonal therapy and expected reproductive performance [
5,
6]. However, tools to easily predict the individual P/AI including HRE and REI information in organic dairy cows are not readily available.
Greater activity may be an indicator of stronger estrus display and greater fertility. Cows with greater REI, measured for example in walking steps/h compared to a baseline, had greater P/AI [
7,
8,
9]. A high correlation has been found between estradiol level and changes in cows’ behavior during the estrus period [
10]. Automated activity monitors (AAM) have been used for decades to measure the activity of dairy cows, and they have become reliable tools for detecting estrus expression using immediate data [
7,
11]. In addition to low REI, many diseases (HRE) are associated with reduced fertility [
12,
13,
14]. Multiple HRE may further reduce P/AI [
15]. We have also seen greater emphasis on predictive models of P/AI for individual inseminations. This development has been fueled by an interest in precision dairy management and an increase in the volume and variety of data collected that may be associated with P/AI, such as data from AAM, genomic testing and cameras [
16,
17].
Organic dairy production is bound by strict regulations [
18]. For example, organic dairy production requires that dairy cows graze not less than 120 days per calendar year. Grazing may reduce estrus detection compared to confined housing [
19]. The use of reproductive hormones to manage estrus and ovulation, and antimicrobials to treat reproductive disease, is not allowed [
20,
21]. In addition, organic dairy cows with reproductive failure may be inseminated more often than conventional cows because of their greater cost to replace for failure to conceive [
22]. Grazing may also have positive effects on health, welfare, and stress reduction and therefore be beneficial to reproductive efficiency. Collectively, however, lower P/AI may be expected in some organic dairy herds than in conventional dairy herds because the technological limitations outweigh the possible fertility advantages of the organic dairy production system.
Machine learning methods other than the more traditional logistic regression have been employed because they can easily handle heterogenous data (see for example Shahinfar et al. [
1], Fenlon et al. [
23], and Barden et al. [
4]). While predictions of P/AI may or may not be more accurate with newer machine learning methods, these methods are often not easy to deploy because they cannot be recreated from only the published study. However, parameter estimates of logistic regression can be published and are easy to incorporate in decision support software. Attention must be paid to the proper vetting of prediction models, including application of discrimination and calibration methods [
24,
25] as is common in the machine learning literature. This has not been common practice for prediction models of P/AI for dairy cattle, which has made comparisons of various prediction models wanting.
To our knowledge, there are no other studies evaluating the relationship between REI, HRE, and their combination and P/AI in organic dairy cows. Therefore, the first objective of this study was to explore associations between REI, HRE, and their combinations, and P/AI in organic certified dairy cows. Two other objectives were to create and report easily deployable equations for prediction of P/AI and to report discrimination and calibration statistics for the prediction models that are also used in the machine learning literature such that results can be easier compared across prediction methods.
We hypothesized that healthy cows would have greater P/AI than cows with at least one HRE. Moreover, cows with a high REI would have greater P/AI than those with a low REI. We also expected an additive effect of HRE and REI, where healthy cows with a high REI would have markedly greater P/AI than cows with at least one HRE and a low REI.
2. Materials and Methods
2.1. Animals and Housing
This observational retrospective cross-sectional study used insemination, health and activity data collected from 2019 to 2021 from 13,365 certified organic dairy cows located in Colorado, USA. All cows belonged to the same owner. Cows were milked twice daily and housed in free-stall barns with sand bedding when indoors and access to a patio for exercise. Alleys were cleaned once daily. The bedding was cleaned daily, and sand was added 3 times per week. During the winter, cows stayed 24 h per day in the free-stall barns.
A total mixed ration (TMR) was fed twice daily to meet or exceed the nutritional requirements for lactating Holstein cows producing 30 kg/d of milk. Nongrazing diets for lactating cows were based on corn silage, wheat silage, grain mix containing soybean, soy hulls, corn wheat, minerals and vitamins, sorghum silage, alfalfa hay, and grass hay. During grazing season, the diets included corn silage, wheat silage, grain mix containing soybean, soy hulls, corn, wheat, and minerals and vitamins, sorghum silage, alfalfa hay, or grass hay. During the grazing season (approximately April to October), cows were offered fresh TMR for approximately 4 h before being moved to pasture where they stayed for at least 8 h before returning for the second milking. Grazing provided approximately 35% of the daily dry matter intake. The grazing management included rotational grazing on various types of grass pastures. Water was always available ad libitum.
Rolling herd milk production was approximately 9000 kg per year. Individual cow milk yields were not available for many cows and were therefore not included in this study. We also had no access to data on body condition scores, body weights, calf birth weights, and other data that potentially might improve the estimation of P/AI.
2.2. Reproductive Management
Cows were inseminated throughout the year. Pregnancy per AI was lower in the summer due to heat stress, resulting in a slightly seasonal calving pattern with heavier calving in the winter. Artificial insemination was performed upon the AAM alerts or visual estrous detection. A list of cows eligible for insemination was generated every morning based on the alerts provided by the cloud-based CowAlert AAM system (IceRobotics, Stirling, UK,
www.peacocktechnology.com/cowalert, accessed on 15 May 2025) software. Visual estrous detection was also performed by the farm staff every morning before the start of inseminations to ensure that cows that were in estrus but not found by the AAM system were also inseminated. Cows on the AAM alert list were inseminated regardless of visual estrus verification by the farm staff. Artificial insemination was performed after cows left the milking parlor using manual catch headgate chutes. Re-insemination was performed upon continued observation of estrous signs after the first insemination. Pregnancy confirmation was done by ultrasonography approximately 32 ± 3 days after insemination. Pregnancy was re-evaluated 67 ± 3 days after insemination.
2.3. Activity and Health-Event Related Data
Activity data were obtained using i-QUBE sensors (IceRobotics, Stirling, UK) worn on the cows’ rear legs to monitor each movement. Some cows received activity sensors starting in 2018, but not all cows wore sensors when the data for this study were collected. Activity data were sent wirelessly to the CowAlert software at every milking. Activity data produced by the sensor included motion index, number of steps, standing time, lying time, standing change, transitions down, and transitions up. The motion index was calculated by a proprietary, undisclosed algorithm. These data were reported per 15 min for each cow, including start and ending times. For our study, we used only the number of steps. We counted the number of steps per hour. Thus, every cow had 24 observations per day, unless observations were missing.
For analysis of each insemination, we included the steps/h for the first 6 h of the AI-day, and 192 h (8 days) before the AI-day. Thus, a total of 198 steps/h for the 198 h prior to 6:00 o’clock on AI-day were used. We assumed that 6:00 o’clock on AI-day was the most recent time (deadline) that activity data were available for calculation of the REI. Inseminations occurred in the morning after the 6:00 o’clock deadline. This way no activity data post insemination was used in our analysis.
Next, we calculated REI as [steps/h at AI-day/steps/h at baseline], expressed as a percentage, following Madureira et al. [
26]. A REI > 100% indicated increased steps/h prior to AI. Steps/h at baseline represented the mean steps/h calculated from hours 31 to 198 (168 h = 7 d) prior to the 6:00 o’clock deadline. The most recent 30 h prior to the 6:00 o’clock deadline (6 h on AI-day and 24 h on the day before AI-day) were not used in the baseline calculation to avoid including possibly increased steps/h associated with estrus. Steps/h at AI-day was calculated as the highest 8 h moving mean in the most recent 30 h prior to the deadline. This calculation aimed to capture the greatest change in steps/h over a long enough period prior to insemination but was not necessarily the most recent 8 h prior to the 6:00 o’clock deadline. We categorized REI for every insemination in one of 4 categories as follows: ≤200%, >200–400%, >400–600%, and >600%.
Health-related events and inseminations were entered into DART software (DRMS, Raleigh, NC, USA,
www.drms.org/Products-Services/Software/Dart, accessed on 15 May 2025) by farm staff and later obtained for our study from DRMS. The same health events that were less than 14 days apart from the previous same health event were omitted. After this filtering, the remaining health events were assumed to be new cases. For each insemination, we obtained the 10 most recent health events between calving and the day prior to AI-day. Often fewer than 10 events were present. We then classified health events as clinical mastitis, metabolic disease (i.e., hypocalcemia, ketosis, displaced abomasum, digestive problems), reproductive disease (i.e., metritis, endometritis, pyometra, retained fetal membranes), and lameness combined with injury. Diseases were diagnosed following a standardized physical exam protocol upon visual detection (e.g., mastitis after the milking procedure, lameness as the cows walked back to the pens from the milking parlor) or fresh cow health check. Cows were treated following the farm protocol using only medications allowed by the USDA-NOP [
18].
Inseminations not having any health event in one of these 4 groups prior to AI-day were classified as healthy. If more than one class had at least 1 health event, such inseminations were classified as 2 or 3 and more HRE.
2.4. Records and Statistical Analysis
Data processing, filtering and statistical analyses were done with SAS 9.4 statistical software (SAS Institute, Cary, NC, USA). After initial filtering for errors and incomplete records, the dataset included 68,744 inseminations and their respective insemination outcomes (open or pregnant). The mean raw P/AI in this dataset was 0.31, calculated as number of inseminations diagnosed as pregnant. The median days in milk (DIM) at insemination was 105 d and the mean 125 d.
An important objective of this study was to create and report easily deployable equations. We therefore formulated three mixed logistic regression models using the GLIMMIX procedure of SAS where insemination was the event and pregnancy (conception) was the outcome (0/1). Prediction equations can be easily built from the published parameters alone. In all models, cow was used as a random effect to account for repeated inseminations within parity. Every model also included the continuous covariates of the natural logarithm of the DIM at insemination, the square of the natural logarithm of the DIM at insemination, and the uncorrected herd-average P/AI in the 3 mo before the month of the insemination. These covariates were known to be good predictors of P/AI.
In addition, every model included the categorical variables of parity (1, 2, 3, ≥4), insemination season (winter, spring, summer, fall) and interservice interval (days from calving to 1st insemination, or ≤17, 18–24, 25–35, 36–48, ≥49 d after previous insemination). These effects are referred to as the control model.
The first model added HRE as a categorical variable to the control model (categories: healthy (no diseases), or 1 case of lameness, mastitis, metabolic, or reproductive disease, or 2 HRE or ≥3 HRE). The second model added REI to the control model (categories: ≤200%, >200–400%, >400–600%, >600%). The third model added combinations of HRE and REI to the control model (the 4 REI categories combined with healthy, 1 HRE, or ≥2 HRE). Therefore, the equations used in the logistic regression models were:
where P/AI, DIM, HRE, REI, and COMBO were defined previously and HPAI is the previous 3-month herd P/AI, PARITY is parity, SEASON is insemination season, INTERSERVICEINTERVAL is interservice interval, and b
cow is the random effect.
The number of observations used in each model is shown in
Table 1. More HRE records were available than REI and COMBO records because initially not all cows were equipped with AAM sensors. Sensors were provided per pen when they became available. The REI and COMBO models included only cows with AAM sensors. The HRE model includes cows with and without sensors. We decided to use all the HRE records available to us to increase the power of the study. Therefore, not every cow was present in each of the three models.
The average cow in the study was approximately 3.7 years old. The number of inseminations per HRE category, REI category, or combined HRE and REI category by parity are in
Supplementary Tables S1–S3. For each of the 3 models, the records were split randomly by cow to create a training set (70% of cows) and a test set (30% of cows). All records belonging to the same cow were therefore assigned to one or the other data set.
The logistic regression models were applied to the training data sets. The results obtained from the logistic models applied to the training sets included parameter estimates expressed on the log-odds scale, odds ratios and the estimated marginal means (EMM) of the estimated P/AI for the categorical variables. This approach allowed us to interpret the effects of REI, HRE and COMBO directly in terms of differences in estimated P/AI which aids interpretation. The parameter estimates were also applied to the test sets to obtain the predicted P/AI for cows that were not included in the training of the logistic regression models.
2.5. Model Evaluation
Discrimination statistics, such as sensitivity and specificity, measure the ability of the model to predict actual insemination outcomes. If the predicted P/AI > cutoff probability, then a pregnancy is predicted, otherwise failure of pregnancy is predicted. We used the standard cutoff probability of 0.5 and the cutoff probability equal to the mean P/AI in each relevant test set (0.27 to 0.31). The standard cutoff probability of 0.5 was added because is often used instead of the true probability of success. However, our study is concerned with the quality of the predicted P/AI, not classification of binary predictions and outcomes. Discrimination metrics cannot confirm that the predicted P/AI are without bias. Therefore, calibration tests were added to estimate accuracy and lack of bias of the predicted P/AI themselves [
24,
25]. All model evaluations were performed on the test sets.
2.5.1. Discrimination
Discrimination calculations were based on tabulating the number of true positive (TP), false positive (FP), true negative (TN), and false negative (FN) cases, where “true” and “false” indicate agreement between predicted and actual insemination outcomes. A “positive” prediction corresponds to pregnancy, whereas a “negative” prediction corresponds to failure of pregnancy. Predicted insemination outcomes follow from whether the model’s predicted P/AI is above or below the cutoff probability.
Calculated discrimination statistics were:
The F1-score is the harmonic mean of precision and recall. These are all common statistics in the machine learning literature, or epidemiology literature, or both.
Precision, recall, specificity, accuracy, negative predictive value, positive predictive value and the F1-score vary between 0 and 1, with a value of 1 indicating excellent classification depending on no FP, FN or both. Some measures can be misleading in unbalanced data sets, such as the measures found in our study, for example, leading to high accuracy and specificity when the actual probability of success is much lower than 0.5. Matthews correlation coefficient is a balanced measure of discrimination success even when classes are imbalanced, as in our study, with a result > 0 indicating that the prediction is better than random and a value of 1 implies perfect prediction.
In addition, we calculated the area under the receiver operating characteristic curve (AUROC), which is equivalent to the concordance statistic. This curve plots the sensitivity vs. 1-specificity for cut-off probabilities varying from 0 to 1. In the context of this study, the AUROC is the probability that the model will predict a higher P/AI for a randomly chosen insemination that resulted in pregnancy than for a randomly chosen insemination that resulted in failure to get pregnant [
23]. A result >0.5 indicates that the model is a better classifier than the population mean and a value of 1 implies perfect classification. The Youden’s J Index is the cut-off probability where the vertical distance between the curve and the diagonal non-discrimination line is the greatest.
Finally, we calculated:
where
N is the number of inseminations, f
i is the predicted P/AI and o
i is the actual outcome of insemination (1 if pregnant, 0 if failure to get pregnant). A Brier score is compared to the baseline Brier score, calculated as:
where
p is the population mean P/AI. Lower Brier scores than the baseline indicate better predictive performance of the model than the population mean. A Brier score of 0 means perfect prediction. The AUROC and Brier scores are independent of the choice of the cutoff-probability.
2.5.2. Calibration
Calibration was evaluated using the Hosmer–Lemeshow goodness-of-fit test, Brier score, the mean absolute calibration error percentage, and plots. A large, non-significant
p-value of the Hosmer–Lemeshow chi-square statistic indicates good calibration. The lower the Brier score, the better the predictions are calibrated. The mean absolute calibration error percentage, calculated as:
measures the calibration agreement between predicted and observed probabilities.
Calibration was also assessed visually with calibration plots for the 3 logistic regression models following Fenlon et al. [
27]. A calibration plot was drawn by grouping the observations into 25 equi-interval bins based on the predicted P/AI and plotting the mean predicted P/AI against the actual proportion of insemination that resulted in a pregnancy within each bin. Points along the diagonal with 95% confidence intervals around the diagonal indicate good calibration.
Finally, to further assess calibration, we created binned residual deviance plots for the 3 logistic regression models following Collet [
28]. We grouped the observations into 25 equi-interval bins and plotted the mean predicted P/AI against the mean residual deviance within each bin. Residuals randomly scattered around zero with 95% confidence interval consistently covering 0 indicate good model fit.
4. Discussion
Our study aimed to quantify the associations between REI, HRE, and their combinations, and P/AI in organic certified dairy cows. We were primarily interested in estimating the magnitude of these associations rather than focusing on hypothesis testing of explanatory variables. We also wanted to create and report easily deployable equations to predict P/AI for organic cows not in the datasets. In addition, we aimed to report discrimination and calibration statistics that are also used in the machine learning literature such that results can be easily compared across studies when other methods than logistic regression are used.
The discrimination and calibration results for the 3 models cannot be directly compared to judge the quality of the linear regression models because the test datasets did not contain the same records. We chose to use all records that met the criteria for the individual models to obtain the maximum power. Far fewer records contained activity data than health data.
The predictive power of our three logistic regression models was weak, as judged by the low AUROC. Hempstalk et al. [
2] reported AUROC ranging from 0.487 to 0.675 for eight different machine learning algorithms and scenarios predicting insemination outcomes using pasture-based, seasonal breeding data from Ireland. In their study, logistic regression was generally the best-performing algorithm, however, Shahinfar et al. [
1] reported AUROC varying from 0.61 to 0.76 for prediction of insemination outcomes in conventionally housed US dairy cows. Barden et al. [
4] summarized the state-of-the art in 2024 and concluded that no predictive models for insemination outcomes have demonstrated high levels of discrimination, despite the use of advanced methods and large heterogenous datasets. It is likely that characterization of the underlying physiological processes is necessary to further discriminate between high and low P/AI [
4].
These comparable studies did not all have access to the same explanatory variables. For example, Hempstalk et al. [
2] included milk yield, mid-infrared spectral data, body condition score, and days open in the previous parity in their models, which were not available in our records. Barden et al. [
4] reported that daily milk yield was the second most important predictor of individual cow P/AI, after herd-average P/AI prior to the AI event. Greater daily milk yield is often associated with lower P/AI early in lactation. That study did not include HRE (except for somatic cell counts). Other differences explain some of the variation in performance reported by these studies. For example, Hempstalk et al. [
2] applied their trained models to data from other herds to measure prediction accuracy, likely leading to poorer performance, whereas Shahinfar et al. [
1] and our study used only internal cross validation data. In addition, Shahinfar et al. [
1] included herd–year–month as an independent variable, but this variable is not available when a model is applied to records from a new farm or year. Therefore, Hempstalk et al. [
2] did not include this variable. We followed Barden et al. [
4] and included herd average P/AI in the three months prior to insemination to make our model as robust and representative to new organic herds as possible. Nevertheless, the performance of our models when applied in other herds remains a question for now.
We used logistic regression to make our models as deployable as possible. Prediction functions to estimate P/AI can easily be constructed from the parameters estimates in
Table 2,
Table 3 and
Table 4 and build into software. Deployment of trained machine learning models, such as random forests and support vector machines, is more challenging even when these models are made available. Often trained machine learning models are not available, and the reader of a published study is tasked to repeat the study with their own data to obtain a prediction function. However, given the goal to predict P/AI, rather than inference on the importance of explanatory variables, other machine learning methods may be useful on the large amounts of heterogenous data that are increasingly collected on dairy farms. Therefore, we attempted to provide discrimination and calibration statistics that are also commonly used in the machine learning community than has been the practice in the veterinary and dairy science literature for similar studies.
Despite the low AUROC, our logistic regression models may still be useful. The predicted P/AI in the calibration bins ranged from approximately 0.25 to 0.45, which is sufficient variation to affect some insemination decisions such as the type of semen, sire choice, delay of insemination, or culling decisions. Barden et al. [
4] further discuss that prediction models with low AUROC, like in our study, may still be useful in practice.
As illustration, consider a cow with as inputs a previous three-month herd P/AI of 0.30, 90 DIM, second parity, inseminated during the spring, an interservice interval of 20 d (18–24 category), and a REI of >600%. The log-odds (
η) is obtained by multiplying each variable input by its coefficient from the REI model and adding the regression coefficients corresponding to the cow’s categorical levels:
The predicted P/AI is then obtained by applying the logit transformation:
Discrimination statistics that are calculated from a classification matrix which depends on a cutoff probability (
Table 5) were only of secondary interest. Using the standard cutoff probability of 0.5, the ratio of FN to FP is high due to the mean P/AI being around 0.3. The conclusion should not be to not inseminate a cow when her predicted P/AI is <0.5. Following Shahinfar et al. [
1], the cost of a FN (failing to inseminate a cow that would have become pregnant, i.e., the cost of a pregnancy [
29]) is much greater than the cost of a FP (inseminating a cow that would not have become pregnant, i.e., semen and labor cost). Therefore, FN should be avoided more than FP, which implies a lower cutoff than 0.5 if the purpose is to decide to inseminate the animal or not. In all these cases, accurate cashflows following each decision alternative are needed to inform the probability cutoff(s) if the goal was to inseminate or not. These insemination decision alternatives are beyond the scope of the current study.
Trends in the parameters of the HRE model were as expected. Healthy, first lactation cows had a greater estimated P/AI. The detrimental effect of suboptimal health on insemination outcomes has been observed widely, in grazing and conventional dairy farms (for example in Ribeiro et al. [
12]; Vergara et al. [
30]). The magnitude of the associations of HRE with P/AI are similar in size to those reported by Pinedo et al. [
14]. Therefore, our study does not provide evidence that the association between HRE and P/AI is different in organic herds compared to conventional herds.
The REI model explored how an increase in activity was associated with greater P/AI. Associations between REI, as measured in a relative increase in steps/h prior to insemination compared to the cow’s baseline activity level, and P/AI were weak and generally not different from inseminations at the lowest level of ≤200% REI. This was surprising because other studies have shown a clear increase in P/AI associated with greater activity [
7,
8,
26,
31,
32]. For example, Madureira et al. [
7], using a similar relative intensity measure as in our study, reported that cows with greater peak activity had greater P/AI than those with lower peak activity (37 vs. 25% for a collar-mounted accelerometer, 34 vs. 21% for a leg-mounted pedometer). Madureira et al. [
26] reported an increase in P/AI from 23% (for relative intensity 100–199%) to 38% (for a relative intensity ≥400%) of all insemination events. They also reported that greater estrous expression at timed AI improved ovulation rates and reduced pregnancy loss.
Our definition of REI was very similar to the calculation using steps/h described by Madureira et al. [
26]. On the study farm, cows were flagged for estrus based on an increase in the motion index, a proprietary algorithm with an unknown calculation for us. Although not identical, changes in the motion index prior to AI-day were like changes in the REI. The motion index is highly influenced by changes in steps/h. It is unlikely that increases in the motion index show a stronger agreement with P/AI than changes in the REI. Furthermore, our interest was not in quantifying the association between the motion index and P/AI because of the proprietary nature of the index.
Cows with low REI may not have been inseminated, and therefore not present in our observational study, if they had not been flagged by the motion index or diagnosed as in estrus by the farm staff. This contrasts with some studies (for example Madureira et al. [
7,
26]) where cows with low increases in activity were still inseminated because of research study design, likely leading to low P/AI. Such studies will have found a stronger agreement between activity and P/AI. It is possible that the weak association between REI and P/AI in our observational study was due to management decisions. The results of our study show that under practical conditions, the REI was not necessarily a good indicator of P/AI, and therefore of limited use for insemination decisions.
We decided to retain the REI fixed effect in the REI model and show the results in
Table 3, even though the REI effect had
p > 0.05. Because our emphasis was not hypothesis testing, we follow Wasserstein and Lazar [
33] who advise that business decisions should not be based only on whether a
p-value passes a specific threshold. Variations in REI may still affect P/AI, even though the evidence for it in our study is not very strong. A relatively large
p-value does not imply evidence in favor of the null hypothesis (no REI effect). However, the more parsimonious base model (
Table S4) offers similar estimated P/AI without the REI effect and might be used instead.
We hypothesized that the strength of the relationship between REI and P/AI might depend on the HRE category. However, we found no evidence for this using the logistic regression model that included both REI and HRE. This finding is likely due to the weak association between REI and P/AI, as shown in the REI model, and therefore the limited improvement in the COMBO model over the HRE model.
Assessment methods for the evaluation of binary prediction models, such as the logistic regression models developed in this study, can be divided into those for discrimination and calibration [
27,
34].
It is our impression that comprehensive model evaluation of binary prediction models for application in dairy science was generally not conducted until the recent past. For example, in a logistic regression model to predict chronic mastitis in dairy cows, Kristula et al. [
35] showed parameter estimates, odds ratios, and probability estimates. The Hosmer–Lemeshow equal-decile goodness-of-fit chi-square statistic was calculated. Calibration was not mentioned. Shahinfar et al. [
1] evaluated their models by the AUROC and the proportion of correctly classified instances (accuracy), but there is no evidence of any calibration evaluation. More recently, Fenlon et al. [
23] reported precision, recall, F1-score, and AUC for their predictive model of insemination outcomes. They used tests of overall goodness-of-fit, bias, and calibration error to measure the accuracy of the predicted probabilities. Fenlon et al. [
25] provide a discussion of calibration techniques for binary and categorical predictive models. These techniques have also been illustrated in Barden et al. [
4].
We acknowledge that our study presents several limitations that hinder its general applicability. First, the study is based on data from a single herd. Specific management practices might have influenced our results and may not be representative of other organic dairy farms. For instance, artificial insemination was performed not only upon estrus alerts but cows were also inseminated upon visual estrous detection by farm staff. Combining visual detection with automated alerts may produce better reproductive outcomes, limiting the generalizability of our results to other farms relying exclusively on automated detection. Second, our models lack external validation, which should caution the reader against the general applicability of the results without further validation. Third, our study lacks specific information of severity of diseases or etiological agents causing the diseases. A fourth limitation is the absence of individual milk production data and other variables that might increase the discriminatory power of our models.