Optimizing models to be used in the genetic evaluation of important economic traits is necessary since they can improve the estimation of genetic parameters, breeding values, and accuracies of predictions. An optimized model can impact the selection of animals to be the sires and dams of the next generations and, consequently, can lead to greater genetic gain and increased profitability. This study compared different models for the genetic evaluation of conformation traits, milking ease, and milking temperament in Dairy Gir cattle to identify the most suitable model. The models differed in their fixed effect structure and the inclusion of contemporary groups as fixed or random effects. Additionally, linear and multinomial threshold models were compared for the genetic evaluation of the categorical traits.
The results and discussion that follow present genetic parameter estimates across all traits and models, with a comparison of goodness of fit, average EBV accuracies, and bulls’ EBV rankings of linear models for groups of traits, followed by a comparison between linear and threshold models for categorical traits, and a summary of results and limitations of this study.
3.1. Genetic Parameter Estimates
The estimates of additive genetic, permanent environmental, and residual variances for all the evaluated models are shown in
Table 4. Heritability and repeatability estimates are presented in
Table 5 for both linear and threshold models. The estimated standard errors for the variance, heritability and repeatability estimates are presented in
Table S1 and Table S2 (Supplementary Material), respectively.
The estimates of genetic, permanent environmental, and residual variances were similar between the linear models for almost all the traits (
Table 4). Body length, rear legs–side view, rear legs–rear view, and teat diameter had a smaller estimated residual variance for linear models M2, M3, and M4 than for linear model M1. This result may be because for these traits, CG was not included in the linear model M1 but it was included in the linear models M2, M3, and M4 as a fixed or random effect, which could have reduce the residual variance.
The estimates of additive genetic variance for teat diameter and body length were higher for linear model M1 than for M2, M3, and M4. However, estimates of permanent environmental variance were higher for linear models M2, M3, and M4 than for linear model M1. The linear M2, M3, and M4 models included CG (herd–year evaluation) for both traits, either as a fixed or random effect, whereas linear model M1 included only the year of evaluation. The difference in the genetic additive variance is reflected in a higher heritability estimate for linear model M1 for these traits. However, the repeatability estimates were similar between linear models M1 and M2 (
Table 5).
Small differences in the repeatability estimates for the hook width between linear models M1 and M2 and linear models M3 and M4 were observed, which could be attributed to the inclusion of the evaluator as a fixed effect in M2 and M4.
The estimates of genetic additive and permanent environment variance were very similar for the threshold models M1 and M2 for all the traits, except for the rear legs–side view, in which M1 showed smaller estimates of genetic additive and permanent environment variance (0.47 and 1.68, respectively) than did M2 (0.71 and 2.50, respectively) (
Table 4). These differences may be attributed to the fixed effects (CG and quadratic effect of the dam’s age) added to the second model. When comparing threshold models M1 versus M3 and M2 versus M4, the estimates of genetic additive and permanent environment variance were smaller for M3 and M4 for almost all traits, except for rear legs–side view and rear udder width, which is expected since one more random effect (CG) was included in threshold models M3 and M4.
The estimates of genetic additive variance (
Table 4) and heritability (
Table 5) for rear legs–rear view were low and not reliable for all four models when both linear and threshold models were used, which may be due to the inability to record this trait correctly in the field or the poor definition of the trait. Traits related to feet and legs are highly influenced by the environment and management [
25,
26], which explains the low heritability estimates. Heritability estimates in the literature for rear legs–rear view are scarce; however, Duru et al. [
27] reported a heritability of 0.11 (±0.18), and Ptak et al. [
28] reported a heritability equal to 0.09 (±0.02) for this trait.
When threshold models M1 and M2 were compared, heritability and repeatability estimates were similar between models (
Table 5). As expected, thresholds M3 and M4, which included CG as a random effect, yielded estimates of heritability and repeatability smaller than those of M1 and M2, which included CG as a fixed effect (
Table 5). This result is expected because the inclusion of CG variance in the estimated phenotypic variance leads to a higher denominator in the fraction, resulting in a smaller heritability and repeatability estimate, which should be interpreted as estimates across CGs in contrast to the estimates within CG, which are obtained when CG is treated as a fixed effect and, therefore, does not contribute to the estimated phenotypic variance.
Heritability estimates for conformation traits, milking ease, and milking temperament vary in the literature, which could be due to several factors, including the statistical model, breed studied, fixed and random effects considered in the models, and data editing, among others [
29]. However, the heritability estimates in this study were generally similar to the results in the literature [
27,
30,
31,
32,
33,
34,
35,
36].
3.2. Heart Girth, Rump Length, Teat Length, and Navel Length
For heart girth, rump length, teat length, and navel length, all tested fixed effects were significant (
p < 0.05), resulting in the same models for M1 and M2 (
Table 2) and the same values of
. Consequently, M3 and M4 were also the same models with the same
values. When the linear models with CGs fitted as a fixed effect (M1 and M2) were compared with the linear models with CG fitted as a random effect (M3 and M4), the linear models M3 and M4 showed a better fit based on their higher
values (
Table 6). For these traits, the average EBV accuracies were the same between M3 and M4 (
Table 7). The differences in average EBV accuracy of bulls with at least 20 daughters with records between M1 and M3 or M2 and M4 were statistically significant (
p < 0.05) for all compared traits (
Table S3 in the Supplementary Material). The EBV Spearman rank correlation coefficients between M1 and M3 (or M2 and M4) for these traits ranged between 0.98 and 0.99 (
Table 8), which indicates that changing the CG from fixed to random effect had little impact on the ranking of bulls for these traits.
3.3. Stature, Pin Width, Rump Angle, Udder Depth, Milking Ease, and Milking Temperament
Stature, pin width, rump angle, udder depth, milking ease, and milking temperament showed the same
values between linear models M1 and M2 (
Table 6), which could indicate that differences in the structure of fixed effects for these traits were not sufficient to affect their
values. The
values between linear models M3 and M4 for these traits were the same. However, when the linear models M1 and M2 were compared with M3 and M4, respectively, the models that included CG as a random effect showed higher values of
, indicating a better fit of the linear models M3 and M4 with CG as a random effect. Although the
values were the same between M3 and M4, fitting nonsignificant fixed effects in a model can lead to increased standard error of the estimates. When the average EBV accuracies for these traits were compared between linear models M1 versus M3 and M2 versus M4, linear M3 and M4 presented higher averages for all traits, except udder depth and milking temperament, which presented similar values (
Table 7). The differences in average EBV accuracy for bulls with at least 20 daughters with records between M1 and M3 and between M2 and M4 were statistically significant (
p < 0.05) for all compared traits, except for udder depth (
Table S3 in the Supplementary Material). Similarly, the EBV Spearman rank correlation coefficient between linear M1 and M2 was greater than 0.99 (
Table 8) for these traits, indicating that differences in the fixed effects fitted in the model led to little or no reranking between both models, also supporting the use of more parsimonious models. The EBV Spearman rank correlation coefficients between linear models M1 and M3 and between M2 and M4 were similar and ranged from 0.97 to 0.99 (
Table 8), indicating that just a few bulls were reranked, especially for pin width and udder depth, when CG was considered a random effect.
3.4. Rear Legs–Side and –Rear Views
For the rear legs–side and –rear views, linear model M1 showed higher values of
than the linear model M2. When linear model M1 was compared to linear model M3, linear model M3 showed higher values of
(
Table 6). This result could indicate that considering CG as a random effect could have led to better model performance for the rear legs–side and –rear views. The average EBV accuracies for these traits were greater for linear model M1 (rear legs–side view) and M4 (rear legs–rear view) (
Table 7). A higher average EBV accuracy for linear model M1 for the rear legs–side view could be explained by a slightly higher additive genetic variance estimated for this model. Differences in average EBV accuracy for bulls with at least 20 daughters with records were statistically significant (
p < 0.05) for rear legs–side and –rear views between M1 and M2, M1 and M3, and M2 and M4 (
Table S3 in the Supplementary Material). The EBV Spearman rank correlation coefficient between linear models M1 and M2 indicated greater reranking when the fixed effects changed between models for rear legs–side view (0.91) than for rear legs–rear view (0.68). The removal of CG from the models, due to its lack of significance, impacted the bulls’ rankings, as well as the genetic parameter estimates, as discussed previously. A large reranking was also observed when linear models M1 and M3 were compared for rear legs–rear view (0.80,
Table 8) and again a much smaller re-ranking was found for rear legs–side view (0.94). Changing the CG from fixed to random effect impacted much less the ranking (0.95 and 0.99) when comparing linear models M2 and M4 for rear legs–rear and –side views, respectively.
3.5. Body Length, Hook Width, Foot Angle, Fore Udder Attachment, Rear Udder Width, and Teat Diameter
The body length, hook width, foot angle, fore udder attachment, rear udder width, and teat diameter showed higher values of
(
Table 6) for linear model M2 than for linear model M1. When the linear model M2 was compared to the linear model M4, the linear model M4 presented higher values of
, indicating a better fit of this model for these traits. The average EBV accuracies were higher for linear model M4 for foot angle and hook width. For body length and teat diameter, linear model M1 had the highest average EBV accuracies, and for fore udder attachment and rear udder width, linear model M2 showed the highest average EBV accuracies (
Table 7). The highest average EBV accuracies found for linear models M1 and M2 for body length and teat diameter and for fore udder attachment and rear udder width could be explained by higher additive genetic variance estimated for these models and traits, respectively. Differences in average EBV accuracy for bulls with at least 20 daughters with records were statistically significant (
p < 0.05) for all compared traits between M1 and M2, M1 and M3, and M2 and M4 (
Table S3 in the Supplementary Material). Removing CG as an effect from linear model M1 for body length and teat diameter may explain the moderately high reranking observed between linear models M1 and M2 (0.88 and 0.89) and linear models M1 and M3 (0.92) (
Table 8). Little reranking was observed between linear models M2 and M4 (0.98 to 0.95) for body length, hook width, foot angle, fore udder attachment, rear udder width, and teat diameter.
3.6. Linear and Threshold Models
A comparison of the
values between the linear and threshold models is not possible since they are in different scales. However, an evaluation of the impacts of both models on average EBV accuracies and the bull rankings is possible. When comparing the linear model with higher
values (M3: rear legs–side and –rear views; or M4: foot angle, fore udder attachment, rear udder width, udder depth, milking ease, and milking temperament) to their respective thresholds for the categorical traits, the average EBV accuracies were higher for the linear models (M3 and M4) than for the threshold models (M3 and M4), except for milking ease (
Table 7). The differences in average EBV accuracy for bulls with at least 20 daughters with records between the linear and threshold models were statistically significant (
p < 0.05) for all compared traits, except for milking temperament for linear model M3 versus threshold model M3 and milking ease for linear model M4 versus threshold model M4 (
Table S3 in the Supplementary Material). The EBV Spearman rank correlation coefficients were greater than 0.94 for all traits (
Table 8) indicating a low degree of reranking between linear and threshold models. The similar results found between linear and threshold models can also be attributed to the transformation of the categorical data to fit normality when linear models were applied. The use of non-normal distributed data can lead to biased genetic parameter estimates and improper animal ranking.
Meijering [
37] reported no advantage of using threshold models over linear models, although threshold models were theoretically a better option. Vanderick et al. [
38] observed that the threshold model had better goodness of fit than linear models. However, according to the authors, no clear advantage was found in terms of predictive ability, and linear models would be more suitable and practical for application in genetic evaluation. Weller et al. [
39] also identified several advantages of using threshold models over linear models, and even though threshold models estimate larger variance components, the rank correlation estimated between the two models was greater than 0.90. Threshold models can theoretically be the best fit for categorical traits. However, this study revealed that the lower average EBV accuracies and minimal changes in the animals’ rankings for most of the studied traits may not justify the implementation of these models in a genetic evaluation.
3.7. Summary and Study Limitations
In this study, models including all fixed effects (M2) performed better or the same as models that included only statistically significant fixed effects (M1) based on their values for most traits, except rear legs–side and –rear views. The inclusion or exclusion of fixed effects in a model needs to be performed carefully. The inclusion of nonsignificant fixed effects can lead to overparameterization of the statistical models and increase the standard error of the estimates. However, when important effects are not included in the model can lead to biases in the evaluation. For traits with the same values between linear models M1 and M2, using a model that includes only significant fixed effects may be recommended to reduce the standard error of the estimates.
Models incorporating CG as a random effect showed better fitting across all assessed traits in this study, as indicated by their higher
values. As suggested by Schaeffer [
13], when an animal model is used, CG should be treated as a random effect, whereas CG was commonly treated as a fixed effect in the past based on Henderson’s selection bias theory under a sire model. Fitting CG as a random effect can result in other benefits. For instance, when CG is treated as a random effect, the loss of information from the removal of CGs with a small number of individuals (e.g., less than 5) would be avoided [
40].
A limitation of this study is the use of to compare linear mixed models. This limitation could be resolved if the models’ maximum likelihood values, instead of restricted maximum likelihood values, were available, which could be used to compare models that differ in their fixed effects.