4.1. From ISA to ISA2
ISA2 preserves the original motivation of ISA while making the computational logic more explicit. The antecedent index established a useful premise: multiple physiological measurements can be weighted and integrated to rank cows exposed to wetland conditions [
3,
4]. ISA2 retains that premise but separates four decisions that should not be conflated. Predictive modelling determines which biomarkers help predict live weight and BCS; clinical information determines favourable biomarker regions; profile equations translate weighted desirability to a common scale; and the final integration rule determines how the two profiles trade off.
The large difference between antecedent [
4] and revised ranks shows why this separation matters. ISA2 is not a renamed or linearly transformed ISA. However, the low concordance cannot be interpreted as evidence that ISA2 selects biologically better cows. Several changes occurred together, and no independent fertility, calf or survival outcome was available. The appropriate conclusion is methodological: the revised architecture produces a materially different and more auditable ranking whose external value must now be tested.
The continuous desirability step addresses the clearest structural limitation. More than three quarters of seasonal records had at least one biomarker close to a clinical limit. With a binary , small analytical variation can reverse a contribution completely. Graded desirability preserves information about proximity to the favourable region. It does not, however, make the reference ranges themselves universally valid. The high proportions outside some laboratory intervals and the observed sensitivity to desirability parameters show that future use requires external calibration rather than uncritical transport of thresholds.
The antecedent ISA should therefore not be interpreted as obsolete or invalidated by ISA2. Its simpler threshold-based architecture remains defensible when the aim is transparent within-herd screening under conditions closely matched to the population, laboratory and management setting in which the ISA was calibrated, when computational or data resources are limited, or when continuity with earlier ISA-based monitoring is important [
3,
4]. In such settings, the lower analytical resolution may be an acceptable trade-off, provided that users recognize the dependence on reference limits and avoid transporting the score uncritically across laboratories, populations or production systems.
ISA2 is preferable when the objective is finer phenotypic ranking, particularly for animals close to physiological reference limits, and when the data and computational infrastructure required for continuous desirability functions and model-derived weights are available. The two tools should therefore be viewed as complementary generations of the same framework rather than as mutually exclusive alternatives. The marked rank reordering observed between ISA and ISA2 does not establish that ISA2 selects biologically superior cows, because neither index was compared here against an independent reproductive, productive, health-survival or economic endpoint. A prospective head-to-head comparison is needed to determine when the additional resolution of ISA2 produces a practically meaningful advantage over the simpler ISA. In this way, ISA2 is considered an evolved model derived from the ISA equation [
4], within the context of the continuous improvement of the process for identifying breeding cows that rank favourably on the measured adapto-productive profile under hot and humid conditions.
4.2. Why ISA2-B Was Preferred Without Claiming That A or C Are Invalid
The three ISA2 alternatives ranked this cohort similarly, but their meanings differ. ISA2-A is easy to communicate and intentionally permits compensation. ISA2-C is useful as an unsupervised comparison, but with only two standardized positively correlated profiles, PCA necessarily gives PC1 equal-magnitude favourable-direction loadings. In the present design, ISA2-C therefore represents a symmetric common-direction benchmark rather than a data-driven estimate of unequal importance for BP-LW and BP-BCS. Its 77.7% explained variance reflects the strength of correlation between the two standardized profiles. A future PCA formulation with more than two component profiles, or a deliberately unstandardized analysis, would answer a different methodological question and should be validated as a new alternative rather than interpreted as the current ISA2-C.
ISA2-B was retained because it matches the stated decision objective: high ranking should require proximity to favourable values on both biomarker-profile axes. The corrected formula measures distance from the joint ideal rather than magnitude from the origin. This resolves the geometry problem of the original quadratic proposal. The choice is therefore based on a prespecified interpretation, not on a marginally higher empirical correlation. Reporting all three alternatives remains useful because it shows whether the selected ranking depends materially on aggregation philosophy. We state plainly that on the present cohort ISA2-B offers no demonstrated practical advantage over the simpler and more directly interpretable ISA2-A: the two rankings agree at Spearman 0.9969 and share fifteen of sixteen shortlisted cows. The case for ISA2-B rests on behaviour this cohort does not exercise, namely the treatment of animals that are strong on one profile and weak on the other, and a user whose objective is explicitly compensatory should prefer ISA2-A. Reporting all three formulations remains informative because it reveals when the choice begins to matter.
Supplementary Figure S8 traces the three rules across the equal-sum path and shows where they separate.
4.3. Model Validation and Feature Importance Define the Evidential Ceiling
The weights inherited uncertainty from the predictive models. GBM had the lowest primary cross-validation error, but repeated grouped folds did not establish a clear advantage over RF. Held-out
values were modest, particularly for BCS. This is a central limit on interpretation, not a secondary technical issue. Two consequences follow. First, the component weight is a ratio of two cross-validated coefficients of determination and inherits their uncertainty, so its value should not be taken as evidence that live weight is meaningfully more predictable than body condition; what protects the ranking is not the precision of that weight but its limited influence, since the Spearman correlation with the retained ranking stays above 0.96 across the whole range examined. Second, permutation importance measures the loss in predictive accuracy when a predictor is scrambled, and when the model predicts only modestly that loss is small and its ordering correspondingly less firm. The weights are an attribution of predictive contribution within one fitted model on one cohort; they are not measurements of physiological importance and are not transportable to another herd without re-estimation. The comparison with random-forest weights reported in
Section 3.6 bounds how much of the ranking depends on the algorithm, but it does not remove this constraint.
Agricultural machine-learning studies increasingly emphasize validation that respects biological dependency. Becker et al. demonstrated the potential of machine learning for predicting heat stress in dairy cattle [
20], while Ribeiro et al. showed that cross-validation strategy can substantially alter estimated cattle-behaviour prediction quality when dependent records are involved [
21]. Grouping all seasons from the same cow in ISA2 therefore makes the reported performance more conservative and more relevant to ranking unseen animals.
Permutation importance should also remain a predictive attribution. The corrected workflow estimates it by cross-fitting exclusively within the training animals, so the independent testing cows no longer participate in biomarker selection or SI estimation. Correlated predictors can nevertheless share or obscure importance [
12], and Kaneko showed that small samples and correlation can destabilize conventional permutation importance [
13]. The repeated training-only importance-rank summaries quantify part of this uncertainty but do not make the weights universal. Larger datasets should compare conditional or grouped permutation approaches and test whether the same importance structure transports to new herds.
Section 3.6 quantifies what that uncertainty costs the ranking. Two learners applied to the same animals are two estimators of the same underlying quantity, so their disagreement measures the sampling variability of the weights rather than adjudicating between the algorithms, and a discrepancy of this size is what held-out coefficients of 0.265 and 0.123 from 45 training cows would lead one to expect. Two features of the comparison are worth separating. The set of retained biomarkers is comparatively stable: seven of ten agree in each profile and the leading biomarkers recur, so the panel of indicators the method identifies is more reproducible than the ordering it produces. The individual ranking is not stable at that level, and the upper quintile is where the instability is concentrated. We therefore retain gradient boosting as a declared convention rather than a justified optimum, since revising the selection on the basis of a comparison run after the fact would reintroduce exactly the contingency the prespecified rule was designed to exclude. The constructive response is not to choose between learners on a sample that cannot separate them, but to reduce the variance of the estimate: averaging scaled importance across several algorithms, or across bootstrap refits of one, would yield weights with a smaller sampling error than any single fit provides. We identify this as the most immediate methodological priority for a subsequent version.
4.4. Biological Plausibility and Recent Evidence on Adaptation
The multivariable biomarker approach is consistent with current work showing that thermal adaptation involves coordinated physiological, hematological and morphological responses. Cartwright et al. reviewed the impact of heat stress on dairy cattle and selection strategies for thermotolerance [
22]. Vieira et al. reported breed-dependent changes in physiological variables under increasing heat load in adapted Brazilian cattle [
23], and later thermography-based work showed heterogeneous responses across cattle breeds [
24]. Silveira et al. integrated machine-learning and path methods and identified hematological and coat-related traits as relevant to adaptive responses [
25]. Proteomic work in tropically adapted Caracu cattle also demonstrated broad plasma changes during heat stress and recovery [
26].
These studies support the general premise of a multivariable profile but do not validate the specific ISA2 weights. Alkaline phosphatase, hair length, hematocrit and creatinine were important because their disruption reduced model performance in this cohort. That is not evidence that manipulating them would change adaptation or productivity. The more defensible interpretation is that they are candidate predictive indicators whose value should be tested prospectively.
Recent sensor and machine-learning studies also illustrate how the framework could evolve. Da Costa Silva et al. used machine learning to estimate thermal-stress indicators in rams from environmental and thermographic measurements [
27], and Schmeling et al. quantified physiological and behavioural reactions of cattle to increasing pasture heat load [
28]. These approaches suggest that future ISA2 versions could integrate environmental and high-frequency sensor data rather than rely only on four seasonal snapshots.
4.6. Measurement, Ranking Stability and Practical Use
The exact invariance to affine BCS relabelling should not be confused with measurement reliability. Manual BCS is observer-dependent. Siachos et al. recently developed and externally tested a machine-learning imaging system for automated dairy-cow BCS [
32], showing how repeated standardized measurements may reduce observer dependence. Future ISA2 studies should compare multiple observers or automated BCS rather than simulate country-labelled scoring conventions.
Ranking stability is equally important. ISA2-B retained most upper-quintile cows under small simulated biomarker error, but stability declined as noise increased, and the desirability-parameter analysis showed that shortlist membership can change under less favourable specifications. The former ordinary bootstrap inclusion probabilities are not reported because resampling animals with replacement can confound uncertainty in rank with the probability that a cow is sampled at all. For the fixed set of 80 cows, measurement perturbation and agreement across ISA2-A, ISA2-B and ISA2-C provide more directly interpretable stability evidence. The upper-20% threshold itself remains a management convention; applications should examine several retention fractions, and a utility-based decision threshold would be preferable if economic information becomes available.
Applying ISA2 to a different herd or laboratory requires deliberate recalibration rather than direct transfer, and the two dependencies involved behave differently. The min–max anchoring is a pure cohort dependency: the four anchoring constants are the observed extremes of the raw profiles, so adding a cow outside that range rescales every other animal. Within one cohort this is harmless because the ranking is invariant to it, but two scores of 70 computed on different cohorts are not comparable quantities. The reference intervals are external to the data and do not shift with the cohort, yet they were obtained from one analyzing laboratory, and the desirability parameters are analyst choices to which the ranking is more responsive than it is to the component weight. We therefore recommend four steps. Freeze the biomarker set, the desirability directions and the two desirability parameters before any new animals are scored. Re-derive the reference intervals from the laboratory that will perform the assays, or establish transferability with shared samples or standardized controls. Decide explicitly whether the anchoring constants are re-estimated on the new cohort, which preserves comparability within it, or held fixed from a reference cohort, which preserves comparability across cohorts at the cost of values falling outside 0–100. And re-estimate the weights, since they are model-specific and cohort-specific, unless the deliberate objective is to transport the earlier weights and test whether they still perform.
4.7. Limitations and Future Research
The present study has ten principal limitations. First, the cohort is small and regional, and predictive performance on unseen cows was modest. Second, GBM was not clearly superior to RF when fold uncertainty was considered. Third, cohort min–max anchors make the 0–100 profiles population-dependent. Fourth, clinical reference intervals came from the source laboratory framework and may not transfer across laboratories or management systems; their midpoints were adopted as operational symmetric targets rather than demonstrated biological optima. Fifth, desirability direction and shape are partly specified and changed the ranking under sensitivity analysis. Sixth, even training-only cross-fitted permutation importance remains affected by correlated predictors and algorithm choice. Seventh, α and β are predictability weights rather than economic, welfare or genetic weights. Eighth, BCS was scored by consensus among four assessors rather than by independent raters, so no measure of inter-observer agreement can be computed and the observer component of the error is unquantified. Ninth, with only two standardized positively correlated profiles, ISA2-C necessarily uses equal-magnitude PC1 loadings and therefore does not learn differential profile weights. Tenth, and most importantly, no independent reproductive, productive, health-survival or economic endpoint was available. Internal stability is not external biological validity.
Two further limitations follow from the design of the field campaign. Genotype is partially confounded with ranch: Brangus animals were drawn from all four ranches, whereas each of the remaining genotypes came from a single ranch, so a difference attributed to genotype could in principle reflect the ranch. This is a consequence of where these populations exist, since the Criollo remnants are not distributed across the region, and it is a further reason for reporting the genotype composition of the upper quintile as descriptive. Selection was also purposive rather than random: the cows were drawn from breeding herds already judged functional by their producers, so the cohort describes the ordering of animals that are already performing and is not representative of the regional population.
An eleventh limitation belongs with the ten above and is among the more consequential. The weights are estimated quantities with substantial sampling error, and the ranking inherits it. Re-estimating them under an alternative but equally defensible modelling protocol, with every other element of the construction held fixed, changed seven of the sixteen shortlisted cows and displaced 39 of 80 animals by more than ten rank positions. This exceeds the effect of varying the component weight or the aggregation rule and is a property of estimating weights from models of modest accuracy on a small number of animals, not a defect of this cohort. It reinforces rather than alters the conclusion already drawn: ISA2 should not be used for individual replacement decisions before the weights are stabilized and the ranking is validated against independent outcomes.
These limitations define a clear validation program. The current biomarker definitions, desirability parameters, profile anchors and primary ISA2-B equation should first be frozen, then applied prospectively to cows from independent farms and years. Outcomes should include pregnancy, conception interval, calving interval, calf survival and weaning weight, cow health, longevity and treatment costs. Leave-farm-out and leave-year-out validation should complement animal-grouped cross-validation. Inter-laboratory reproducibility should be quantified with shared samples or standardized controls, and BCS should be assessed by multiple observers or automated imaging [
32].
That validation program should include a frozen, prospective comparison of the antecedent ISA with ISA2-A, ISA2-B and ISA2-C in the same animals, without retrospective retuning to the validation outcomes. For each external endpoint, future studies should compare discrimination, calibration, rank stability, overlap across practically relevant replacement fractions and the consequences of reclassification for management decisions. This design would show whether the simpler ISA remains adequate for screening in low-resource or continuity-focused applications, or whether the additional resolution of ISA2 translates into measurable gains in fertility, calf performance, health, survival, longevity or economic return.
A second research line should test alternative weighting philosophies. Conditional or grouped feature importance can be compared with the current permutation weights, while production-system economics and welfare priorities can be introduced separately from predictability. If pedigree or genomic information becomes available, formal heritability and genetic-correlation analyses should precede any claim that ISA2 is a breeding-value index. Resilience studies based on dense activity or milk-yield trajectories provide a precedent for this transition from an observed indicator to a genetically evaluated phenotype [
29,
30,
31].
Finally, ISA2 can be extended from four seasonal measurements to dynamic profiles combining activity sensors, automated BCS, thermography and environmental data. Such an extension should focus on deviation and recovery, not simply on adding more predictors. The objective is not to maximize algorithmic complexity but to determine whether a transparent multivariable phenotype predicts outcomes that matter to producers and animal welfare better than simpler measurements alone.