1. Introduction
Colorectal cancer is one of the major global disease burdens, accounting for 9.6% of newly diagnosed cancer cases and 9.3% of cancer-related deaths worldwide according to GLOBOCAN 2022 [
1]. Molecular stratification has become increasingly important for prognosis and treatment selection in this disease [
2,
3]. In particular, stratification of MSI-H status is especially important because immune checkpoint inhibition has been shown to provide clinical benefit in appropriately selected patients [
4,
5]. Assessment of MSI (microsatellite instability) status using polymerase chain reaction (PCR) has been widely applied in colorectal cancer screening because it contributes to the evaluation of Lynch syndrome and may provide information for treatment decision making [
6,
7,
8,
9]. Although this method is a standard for clinical reference, it requires additional genetic testing or immunohistochemical procedures, time, and technical resources, and MSI testing may not be consistently completed in every eligible patient [
3,
10,
11,
12]. Current professional guidance continues to emphasize appropriate MMR/MSI testing for clinical decision making [
13]. Therefore, developing computational prescreening strategies based on routinely available hematoxylin- and eosin-stained histopathology slides may reduce these limitations during the initial stage of molecular screening [
14,
15,
16].
Studies have demonstrated that deep learning models can predict MSI status directly from H&E-stained histological slides [
15,
16,
17,
18]. Early approaches generally relied on convolutional neural networks trained or fine-tuned for a specific prediction task [
15,
16,
17,
18]. In a large international study, deep learning achieved the ability to detect colorectal cancer samples with dMMR or MSI with a mean AUROC of 0.92 and could stratify colorectal cancer at large scale and low cost [
15]. Subsequent studies have further investigated automated MSI prediction and its potential use for prescreening or prioritization of molecular testing [
14,
19,
20]. More recently, deep learning models have increasingly been pretrained on larger and more diverse input datasets, such as large collections of unlabeled images or heterogeneous whole-slide images, to learn diverse morphological tissue features [
21,
22,
23,
24,
25,
26,
27,
28]. Deep learning models such as UNI, CONCH, Virchow, Virchow2, and Phikon differ substantially in pretraining data, objective functions, model scale, and image resolution [
21,
22,
23,
29,
30], but each aims to provide reusable features for subsequent computational pathology tasks. Representations extracted by frozen image encoders can be used as inputs for downstream linear classifiers, nonlinear machine learning algorithms, or multiple-instance learning models, as increasingly examined in standardized pathology foundation model benchmarks [
31,
32,
33,
34,
35]. For whole-slide analysis, clustering-constrained attention multiple-instance learning aggregates variable numbers of patch-level feature vectors into slide-level representations and requires only slide-level labels, without pixel-, region-, or patch-level annotations [
36,
37,
38]. Beyond molecular biomarker prediction, deep learning analysis of histopathology has also been investigated for clinically relevant tasks such as prognosis, risk stratification, and treatment response prediction across cancer types [
39].
Although deep learning models have been increasingly developed to facilitate MSI screening, several issues remain unresolved. First, pathology-specific foundation models and conventional vision backbones are often evaluated using different cohorts, preprocessing procedures, downstream classifiers, and validation designs, making it difficult to determine whether observed performance differences reflect the quality of the frozen representations or differences in the pipeline [
31,
32,
33,
34]. Therefore, a controlled comparison under a common frozen feature framework is needed. Second, it remains uncertain whether subsequently developed, more complex models can improve predictive performance and thereby improve clinical relevance when a strong frozen representation is already available [
32,
33,
34]. Third, preserved discrimination in an external cohort does not necessarily mean that predicted probabilities and decision thresholds are also preserved. Histopathology models can be sensitive to differences in image acquisition, staining, institutional characteristics, and other hidden variables, resulting in a domain shift when applied to new cohorts [
40,
41,
42,
43]. Moreover, probability distributions may shift or lose calibration because discrimination and calibration represent distinct dimensions of predictive performance [
44,
45,
46,
47]. This indicates that, even when discrimination is preserved during external validation, predicted probabilities or decision thresholds may still be altered when the model is applied to a completely new dataset.
This study aimed to evaluate the performance of nine frozen image encoders for MSI-H prediction from colorectal cancer histopathology while comparing the task utility of their frozen representations under a common downstream framework. We conducted the study with three main objectives: to compare frozen representations generated by pathology-specific foundation models and conventional vision backbones under a common evaluation framework; to determine whether the tested nonlinear classifiers or attention-based multiple-instance learning (MIL) configuration can provide consistent and practically meaningful improvements in MSI-H discrimination compared with a simple linear probe; and to separately assess external discrimination, probability calibration, and the transportability of development-derived decision thresholds. We hypothesized that pathology-pretrained encoders would generate MSI-H-related representations that were sufficiently linearly separable for regularized linear probes to achieve performance comparable to that of the more complex downstream models tested. In addition, we hypothesized that model discrimination, probability calibration, and development-derived decision thresholds would be preserved when the models were applied to an independent external cohort.
2. Materials and Methods
2.1. Study Design
This study was designed as a frozen-representation benchmark rather than as a task-specific model-optimization study. The primary question was whether pathology-specific foundation models provide more discriminative frozen representations for MSI-H prediction than conventional computer-vision backbones under an identical patient-level evaluation framework. The benchmark pipeline was fixed as follows: H&E image tiles were passed through a frozen encoder, tile embeddings were mean-pooled to slide embeddings, slide embeddings were mean-pooled to a patient embedding, and the patient embedding was evaluated with the same logistic regression linear probe to produce a patient-level MSI-H probability. No encoder was fine-tuned at any stage.
The main comparison was defined at the frozen-representation level while acknowledging that the encoders also differ in architecture, scale, embedding dimensionality, input resolution, pretraining objective, and pretraining data. Encoders were divided into pathology-specific models (CONCH, CONCH v1.5, UNI, Virchow2, and Phikon) and conventional vision backbones (ResNet18, ResNet50, ViT-B/16 pretrained on ImageNet-21k, and ConvNeXt-Tiny). The primary endpoint was a pooled out-of-fold AUROC from the linear probe on the development cohort. The secondary analyses assessed nonlinear classifier dependence, attention–MIL performance, representational similarity, calibration, external transportability, local threshold adaptation, and top-tile interpretability artifacts.
2.2. Development Cohort
Candidate TCGA-COAD/READ patients were drawn from the TCGA-CRC-DX resource using class-balanced sampling from a 616-patient pool: all available MSI-H patients were retained, and MSS patients were randomly sampled at an approximate 1.5:1 MSS:MSI-H ratio (fixed seed 20260709), yielding 209 candidate patients (83 MSI-H, 126 MSS). Of these, 3 patients were subsequently dropped because of incomplete or corrupted tile archive downloads, leaving 206 patients with usable tile data. TCGA-CRC-DX tiles were provided as 512 × 512 pixel tumor tiles at approximately 0.5 μm/pixel, so no additional WSI tiling or tumor filtering was required for the development cohort.
The MSI labels for these 206 patients were derived from a MANTIS-score threshold, with MANTIS > 0.4 labeled MSI-H. The complete 206-patient cohort (82 MSI-H, 124 MSS) was retained as the primary development benchmark. A subsequent label provenance audit cross-referenced this labeling against independent MSIsensor scores and categorical MSI provenance and identified 15 of 206 patients (7.3%) as cross-source discordant. Fourteen patients had MANTIS scores above 0.4 but MSIsensor scores consistent with MSS (MSIsensor 0.00–2.26), while one patient had categorical MSI provenance despite a MANTIS score below 0.4 and a low MSIsensor score. To quantify the influence of this molecular-label discordance without relabeling the primary cohort after observing performance, we report the concordant-label 191-patient cohort (68 MSI-H, 123 MSS) as a secondary label provenance sensitivity analysis (
Section 3.9). The improved sensitivity cohort performance should not be interpreted as evidence that the excluded cases were incorrectly labeled; rather, it illustrates the influence of molecular-label provenance on measured performance. This study therefore reports MSI-H prediction from molecular MSI-derived labels and does not claim IHC-confirmed mismatch-repair deficiency.
2.3. External Cohort
CPTAC-COAD was used as the independent external cohort after the data-readiness review. MSI status was obtained from the cBioPortal CPTAC-COAD study, where the usable MSI call was recorded under MUTATION_PHENOTYPE. The final external cohort included 105 patients, comprising 24 MSI-H and 81 MSS patients across 221 whole-slide images. Patients without a usable MSI label were excluded. The nominal MMR status was unknown for the external cohort; therefore, all external analyses used MSI-H versus MSS as the endpoint. An overlap audit confirmed no shared patient or slide identifiers between the TCGA development cohort and the CPTAC external cohort.
CPTAC-COAD images were downloaded as raw SVS whole-slide images. Unlike TCGA-CRC-DX, CPTAC-COAD required local tiling and tumor tile filtering. Whole-slide images were tiled at a target resolution of 0.5 μm/pixel to match the TCGA tile scale. Foreground tissue tiles were generated as 512 × 512 pixel images after HSV saturation-based background rejection. A separate tumor-filtering step then retained tiles predicted as tumor epithelium by the
tiatoolbox ResNet18-Kather100K tissue classifier. This produced 117,151 tumor tiles for the 105 externally evaluable patients (
Figure 1).
2.4. Frozen Encoder Feature Extraction
Nine encoders were evaluated as frozen feature extractors: CONCH, CONCH v1.5, UNI, Virchow2, Phikon, ResNet18, ResNet50, and ViT-B/16, pretrained on ImageNet-21k, and ConvNeXt-Tiny. For each encoder, tile embeddings were extracted without gradient updates. Patient-level features for linear probing were computed by mean-pooling tile embeddings to slide level and then mean-pooling slide embeddings to patient level, so each slide contributed equally to the patient-level representation, regardless of its tile count. This aggregation choice was intentionally simple, deterministic, and identical across encoders, so that the comparison assessed the downstream task utility of each frozen representation under the same aggregation and classifier protocol rather than isolating representation quality as a causal factor.
Because some pathology foundation models are pretrained on public histopathology resources that may overlap common downstream benchmarks, we performed a source-level pretraining provenance audit using public model cards, repositories, and associated papers (
Table 1). This audit was designed to identify dataset-level overlap risk. It cannot prove patient- or slide-level independence for encoders whose pretraining manifests are not public.
For attention–MIL experiments, per-tile embeddings were retained as patient-level bags. When a patient had multiple slides, all tiles from that patient’s slides were concatenated into one bag. This matched the patient-level unit used by the linear probe analysis and avoided slide-level leakage across cross-validation folds.
2.5. Linear Probe and Nonlinear Classifier Benchmark
Each encoder was evaluated on the TCGA development cohort using patient-level stratified 5-fold cross-validation with shuffling and a fixed random seed of 20260709. For every fold, the training and test patient sets were explicitly checked to be disjointed. Features were standardized within the training fold and transformed into the held-out fold. The primary linear probe was logistic regression with maximum 2000 iterations, balanced class weights, and the same random seed. Three nonlinear comparators were trained on the same patient embeddings: radial-basis-function support vector machine with probability output and balanced class weights, XGBoost with 200 estimators, maximum depth 3, learning rate 0.05, and log-loss evaluation, and 5-nearest-neighbor classification.
Pooled out-of-fold predictions were used to compute AUROC, AUPRC, sensitivity, specificity, accuracy, and F1 score. The default threshold for these internal classifier metrics was 0.5 and was not interpreted as a clinical rule-out threshold. To quantify linear separability, we defined the linear-separability gap for each encoder as:
A near-zero or negative value indicates that the frozen embedding is already linearly separable for the MSI-H task.
2.6. Representation Geometry, Similarity, and Calibration
UMAP projections were generated from patient-level embeddings for qualitative visualization of class geometry, using 15 neighbors and a minimum distance of 0.1. These plots were used as qualitative support only and were not treated as primary statistical evidence.
Pairwise linear centered-kernel alignment (CKA) was computed across all encoders after row-aligning patient embeddings by patient identifier. For column-centered embedding matrices
and
, linear CKA was defined as:
Calibration was assessed on pooled out-of-fold predictions using the Brier score, mean predicted probability, calibration intercept, and calibration slope for every encoder–classifier pair. Calibration curves were generated with quantile binning. Calibration intercept and slope were estimated by fitting a logistic calibration model with the logit-transformed predicted probability as the only covariate.
2.7. Attention–MIL Comparator
To test whether a trained nonlinear tile aggregator improved over simple patient-level linear probing, we trained a gated-attention multiple-instance learning model for each encoder using CLAM_SB. The model used gated attention, the small CLAM configuration, a dropout of 0.25, k_sample = 8, two output classes, and encoder-specific embedding dimensionality. Training used the same patient-level stratified 5-fold split seed as the linear probe analysis. Each fold was trained for 20 epochs with Adam optimization, a learning rate of , and weight decay of . The loss combined bag-level cross-entropy and CLAM instance-level loss with weights of 0.7 and 0.3, respectively; for rare patients with fewer than eight tiles, instance-level supervision was skipped and bag-level loss alone was used. Fold checkpoints were saved separately and pooled out-of-fold predictions were used for evaluation.
2.8. External Validation and Local Threshold Adaptation
For external linear probe evaluation, each encoder’s final logistic regression pipeline was fit on the 206-patient primary development cohort and applied unchanged to CPTAC-COAD patient-level embeddings. A TCGA-locked threshold was selected from the development out-of-fold predictions at target sensitivity 0.93 and then applied to the external cohort without refitting. The 0.93 internal target was selected as an illustrative high-sensitivity operating point for exploratory prescreening analysis, rather than as a guideline-derived or clinically validated requirement. Discrimination metrics were reported independently from threshold metrics, because AUROC can transport even when the probability scale and decision threshold do not.
For the CONCH attention–MIL external experiment, the five fold checkpoints trained on the 206-patient cohort were loaded without weight updates and applied to each CPTAC patient bag. The five fold probabilities were averaged into one ensemble probability per patient. The locked TCGA threshold for the CONCH attention–MIL model was 0.000577, chosen from TCGA out-of-fold predictions at target sensitivity 0.93.
Local threshold adaptation was evaluated as a threshold-only simulation. In each of the 100 repeats, 15 MSI-H and 30 MSS CPTAC patients were drawn without replacement as a site-local threshold-adaptation subset, and the decision threshold was reselected on that subset at a target sensitivity of 0.90. The 0.90 target was also an illustrative high-sensitivity operating point and should not be interpreted as an endorsed clinical sensitivity target. The remaining 9 MSI-H and 51 MSS patients were held out for evaluation. Predicted probabilities and model weights were never updated, so this analysis is not a formal probability recalibration. Because the external cohort contained only 24 MSI-H patients, repeat-level threshold-adaptation subsets necessarily overlapped substantially; therefore, percentile intervals from the 100 repeats were interpreted as exploratory uncertainty summaries rather than independent validation intervals.
2.9. Statistical Analysis
Uncertainty in the primary group-level comparison was estimated with patient-level paired bootstrap resampling. For each of 2,000 bootstrap replicates, the same 206 TCGA patients were sampled with replacement, AUROC was recomputed for every encoder’s logistic regression out-of-fold predictions, the mean AUROC across the five pathology-specific encoders was computed, the mean AUROC across the four conventional encoders was computed, and the difference
was recorded. Because the encoders were selected models rather than random experimental units sampled from broader pathology-specific and conventional model populations, this grouped comparison was interpreted descriptively and not as a formal class-level hypothesis test. Analyses were performed at the patient level.
2.10. Interpretability Artifacts
Top-tile exports were generated to support later morphological review. The tiles were ranked in two complementary ways: attention–MIL ranking by learned attention weight and linear probe ranking by the dot product between per-tile embeddings and the scaled logistic regression coefficient vector. Exported tile images and manifests preserve patient, slide, coordinate, label, encoder, and ranking provenance. These tile exports are interpretability artifacts for pathologist review and are not yet presented as validated biological findings.
4. Discussion
This study provides evidence for the ability to predict MSI-H from colorectal cancer histopathology using nine frozen image encoders. First, the discriminative performance of frozen representations varied substantially across the nine encoders when compared within the same standardized linear probe framework. Several encoders pretrained on large-scale histopathology images, particularly Virchow2 and UNI, generated representations that better discriminated MSI-H from MSS than conventional vision backbones, with AUROCs of 0.861 and 0.855, respectively, whereas other pathology-specific models and conventional vision architectures showed lower discriminative performance. This suggests that pretraining on large-scale histopathology data may enable encoders to learn morphological features related to the glandular architecture, tumor differentiation, lymphocytic infiltration, and tumor microenvironment that are associated with MSI-H and can still be exploited by a simple linear classifier [
53,
54,
55,
56]. However, the advantage of pathology-specific pretraining was not uniform across all models. The CONCH v1.5 encoder achieved an AUROC of 0.780, lower than several conventional backbones such as ResNet18 and ResNet50, whose underlying architectural families were originally developed for general computer vision tasks [
57,
58]. Therefore, the results do not support an absolute conclusion that every pathology foundation model outperforms conventional vision models. Instead, they indicate a descriptive, group-level advantage among the selected encoders, accompanied by substantial variability across individual encoders. Representation quality may depend not only on whether a model was pretrained on pathology data, but also on the scale and composition of the pretraining dataset and the self-supervised objective [
21,
22,
23,
24,
25,
26,
27,
28], as well as model architecture, image resolution, and sampling strategy during pretraining [
29,
30]. Previous studies have demonstrated the utility of pathology-pretrained models designed to learn morphological representations across a variety of downstream tasks [
21,
22,
24,
25,
26,
27,
28,
32,
59]. More recent standardized benchmarks have also shown that the relative ranking of pathology foundation models can vary across datasets, prediction tasks, sample sizes, and downstream adaptation strategies [
31,
32,
33,
34,
35].
Second, a label provenance audit of the development cohort showed that discrimination performance was sensitive to the molecular MSI label definition. The as-collected TCGA-CRC-DX cohort of 206 patients used a single MANTIS-score threshold to assign MSI-H status; cross-referencing this labeling against independent MSIsensor scores and categorical MSI provenance identified 15 of 206 patients (7.3%) with cross-source discordance. We retained the complete MANTIS-defined 206-patient cohort as the primary development benchmark and report the 191-patient concordant-label cohort only as a secondary sensitivity analysis (
Table 11). AUROC for the strongest linear probes increased in the concordant-label cohort: UNI increased from 0.855 to 0.928, Virchow2 from 0.861 to 0.926, Phikon from 0.822 to 0.890, and CONCH from 0.816 to 0.869. This pattern indicates that molecular-label provenance materially affects apparent discrimination. However, it should not be interpreted as proof that the discordant cases were incorrectly labeled or as evidence that the higher-performing concordant subset should replace the originally specified full MANTIS-defined cohort as the primary benchmark.
Third, the results showed that in the primary development cohort of 206 patients, Virchow2 and UNI achieved the strongest discrimination with linear probes, with AUROCs of 0.861 and 0.855, respectively, while nonlinear classifiers did not further improve performance for these two encoders. Specifically, the best nonlinear AUROCs for Virchow2 and UNI were 0.856 and 0.853, corresponding to linear-separability gaps of
and
. In contrast, some encoders, such as CONCH and CONCH v1.5, improved when nonlinear classifiers were used, with AUROC increasing from 0.816 to 0.856 for CONCH and from 0.780 to 0.809 for CONCH v1.5. This suggests that the benefit of nonlinear decision boundaries depends on the characteristics of the representation generated by each encoder rather than reflecting a general advantage of nonlinear models. A similar pattern was observed when comparing linear probes with attention–MIL. Attention–MIL improved performance for some encoders, particularly CONCH and CONCH v1.5, with AUROC increasing from 0.816 to 0.855 and from 0.780 to 0.837, respectively, but did not provide a consistent benefit across the remaining models. For Virchow2 and UNI, the two encoders with the highest linear probe performance, attention–MIL produced lower AUROCs than the linear probes (0.846 vs. 0.861 for Virchow2 and 0.843 vs. 0.855 for UNI). For several other backbones, performance also decreased after the use of attention–MIL, including ResNet18 (0.798 to 0.694), ViT-B/16 (0.763 to 0.738), and ConvNeXt-Tiny (0.738 to 0.711); ConvNeXt represents a modernized convolutional architecture originally developed for general computer vision [
60]. Overall, increasing downstream model complexity through the tested nonlinear classifiers or the tested attention–MIL configuration did not provide consistent improvement over a simple linear probe. Therefore, the results do not support the assumption that using a more complex downstream architecture automatically leads to better MSI-H discrimination.
One possible explanation is that strong frozen representations had already encoded most of the MSI-H-related morphological signal in a manner that could be effectively exploited by a simple linear decision boundary. In this situation, more complex models increase the number of parameters to be estimated, dependence on hyperparameters, variability arising from optimization, and the risk of overfitting, particularly when the patient-level sample size remains limited [
32,
33,
61]. A linear probe may therefore be considered an important and practically useful baseline because of its simple architecture, low computational cost, and favorable reproducibility. However, these results do not demonstrate that linear models are generally superior to attention–MIL, because some encoders, such as CONCH and CONCH v1.5, still benefited from greater downstream complexity. Attention–MIL remains architecturally valuable because it can aggregate variable numbers of tiles across patients and learn different weights for individual tiles rather than treating all tiles as equally important [
36,
37,
38]. Recent benchmarking studies of pathology foundation models have also shown that downstream performance depends not only on the complexity of the aggregator but also strongly on the quality of the frozen representation, the type of task, and the cohort being evaluated [
31,
32,
33,
34]. Some benchmarks have shown that relatively simple aggregation methods can remain competitive with newer and more complex architectures [
32,
33,
34], emphasizing that architectural complexity should not be regarded as a direct proxy for predictive quality.
Fourth, the three components of discrimination, predicted-probability calibration, and the transportability of thresholds derived from the development cohort were not necessarily preserved simultaneously when the model was applied to external data. In the independent CPTAC-COAD dataset, CONCH attention–MIL trained on the primary development cohort still maintained relatively good discrimination, with an AUROC of 0.880. However, the threshold derived from TCGA did not produce a meaningful binary classification when directly applied to CPTAC-COAD, with a specificity equal to 0. This indicates that the ability to rank MSI-H cases above MSS cases was preserved to some extent, whereas the development-derived threshold did not transport to the new cohort. This finding is consistent with established prediction model methodology, showing that discrimination and calibration characterize different aspects of predictive performance [
44,
45,
46,
47]. AUROC primarily reflects the relative ranking ability between individuals with and without an outcome, whereas calibration reflects the agreement between predicted probabilities and observed outcome frequencies. Therefore, a model may maintain relatively good discrimination while failing to preserve the same probability scale or decision threshold when transported to another population.
Differences between cohorts in tissue processing, H&E staining, scanners, image resolution, case mix, and tissue region selection are recognized sources of distributional variation and domain shift in computational pathology [
40,
41,
42,
43]. In this study, the probability distribution of MSS cases in TCGA was concentrated near 0, whereas in CPTAC, it was distributed much more broadly across the probability scale, suggesting a change in probability scale between the two cohorts. We also found no clear evidence of label inversion, extraction of the wrong probability column, application of softmax along the wrong dimension, or checkpoint–loading incompatibility; probabilities were obtained from the correct MSI-H class, and the external AUROCs across folds were directionally consistent rather than suggesting that one fold had been abnormally inverted. These checks reduce the likelihood that the specificity of 0 was simply the consequence of a basic coding error. However, the current results also do not allow the entire threshold failure to be attributed to domain shift, because prediction generation differed between development and external evaluation. The development threshold was derived from out-of-fold predictions, whereas the external prediction was generated by averaging probabilities from five fold-specific models. These two mechanisms do not necessarily produce the same probability distribution, and prediction model methodology emphasizes that calibration and operating characteristics must be evaluated for the actual model implementation used in the target population [
45,
46,
62,
63]. Therefore, part of the threshold failure may result from incompatibility between the prediction-generating procedures in addition to genuine differences between the two cohorts. After the threshold was reselected within the CPTAC threshold-adaptation subsets, the specificity of CONCH attention–MIL improved from 0 to a median of 0.686, while median sensitivity was maintained at 0.889. This suggests that at least part of the failure of the frozen threshold was related to the position of the operating point or a shift in probability scale rather than a complete loss of discrimination. However, this improvement remained insufficient to establish a clinically deployable threshold. Furthermore, this analysis represented threshold-only adaptation rather than formal probability recalibration; reselection of the threshold does not automatically correct the relationship between predicted probability and observed event frequency [
47,
62,
63].
Interpretation of the threshold-adaptation results was also limited by the small external sample size. CPTAC-COAD contained only 24 MSI-H cases, of which each repeat used 15 MSI-H cases for threshold selection and retained only nine MSI-H cases for evaluation. Therefore, a single misclassified case could change sensitivity by more than 11 percentage points. Contemporary methodological work has shown that the required sample size for external validation depends on the precision required for discrimination, calibration, and other performance measures rather than on a single universal event-count rule [
64,
65,
66,
67]. In addition, the same small number of patients was reused across 100 splits; each MSI-H patient appeared in the threshold-adaptation subset an average of approximately 62.5 times, and the overlap between any two splits was approximately 62.4%. Therefore, the percentiles obtained across the 100 repeats primarily reflect the sensitivity of the results to how the data were split rather than reproducibility across independent populations. These repeated splits do not increase the effective external sample size and cannot replace a larger external-validation cohort. More generally, performance in external populations can be affected by differences in case mix as well as by model misspecification [
45,
46,
68]. Overall, the results indicate that discrimination may be better preserved than the transportability of the probability scale and decision threshold when a model is applied to an independent cohort. At the same time, stronger apparent discrimination in the concordant-label sensitivity cohort did not eliminate the problem of external threshold transportability. Therefore, external validation of pathology AI models should not rely solely on AUROC but should separately assess discrimination, calibration, and threshold-specific operating characteristics [
45,
46,
53]. Establishing a deployable threshold should be performed in a sufficiently large calibration cohort and subsequently confirmed in a completely independent test cohort [
64,
65,
66,
67].
From a translational perspective, these findings support further investigation of histopathology-derived MSI-H scores as prescreening or prioritization tools, but they do not support replacement of reference molecular testing [
3,
13,
14]. A proposed prescreening strategy should clearly define its intended use, acceptable false-negative rate, target population, and expected MSI-H prevalence. Measures such as the number of confirmatory tests avoided, the number of MSI-H tumors missed, and net clinical benefit provide more informative guidance for deployment decisions than discrimination alone [
63,
69,
70]. We therefore moved the detailed per-1000 rule-out projection to
Supplementary Table S1 and retained only a brief main-text summary. Using the observed MSI-H prevalence of 22.9% in CPTAC-COAD, the n = 206-trained CONCH attention–MIL-based strategy could avoid approximately 555 confirmatory tests per 1000 patients in the median scenario, but would simultaneously miss approximately 25 MSI-H cases. More importantly, uncertainty remained substantial, with the number of missed MSI-H cases at the 97.5th percentile reaching approximately 102 per 1000 patients. Therefore, the potential benefit of reducing the number of tests must be weighed against the risk of false-negative results, particularly for a strategy intended for rule-out purposes. These findings should be regarded as an exploratory simulation of workflow efficiency rather than evidence of clinical safety. Confirmation of model utility will require a prespecified threshold evaluated in a larger, completely independent external cohort with a sufficient number of MSI-H cases to provide stable estimates of sensitivity and the false-negative rate [
65,
67].
In the future, the study should be validated in larger and multicenter external cohorts with a sufficient number of MSI-H cases to allow precise estimation of discrimination, calibration, and decision threshold stability [
64,
65,
66,
67]. A more rigorous design should separate the cohort used for probability or threshold calibration from a completely independent test cohort used for final evaluation [
45,
62,
63]. In addition, the effects of different tissue-region-selection strategies should be assessed, image-processing procedures should be standardized across cohorts, and blinded evaluation of highly weighted tiles should be performed by pathologists. The importance of controlling technical and institution-specific variation in digital pathology has been repeatedly demonstrated [
40,
41,
42,
43]. Future studies may also explore multimodal strategies that integrate pathology-derived representations with complementary clinical, imaging, or molecular information, as such frameworks have shown potential for improving clinically oriented prognostic modeling [
71]. These steps will help determine whether the MSI-H signal learned from frozen pathology representations is sufficiently stable and clinically valuable to be developed into a prescreening tool for clinical support.