Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding
Abstract
1. Introduction
- We propose a source-only evidence-guided vision–language segmentation framework that constructs image-aware segmentation queries from dense visual features and class-domain textual context.
- We develop an evidence-guided decoding and feature–style regularization strategy for source-domain learning. Positive and negative pathology evidence are incorporated into query-based mask decoding, while text-conditioned style perturbation adds feature-level variation during training.
- Experiments on cross-organ adenocarcinoma segmentation show that the proposed method achieves state-of-the-art performance compared with representative segmentation and domain generalization baselines.
2. Related Work
2.1. Domain Generalization in Computational Pathology
2.2. Vision–Language Foundation Models for Pathology
2.3. Language Descriptions for Visual Recognition and Segmentation
2.4. Generative and Feature–Style Regularization for Domain Generalization
3. Methodology
3.1. Problem Formulation
3.2. Vision–Language Query Construction
3.3. Evidence-Guided Query Decoding
The 180 sample-level records were screened at the record level. A record was eligible when it could be parsed as a JSON object and was neither labeled exclude nor assigned usable_for_training=false. A total of 137 records met these conditions, comprising 127 usable and 10 uncertain records. The remaining 43 records were excluded.“Use the mask only to identify the segmentation target. Do not infer tumour grade, stage, prognosis, or molecular status. Do not use organ identity, stain darkness, scanner brightness, blur, compression, background colour, or image artifacts as positive evidence. Focus on stable morphology that may transfer across hospitals, stains, scanners, and organs. Return strict JSON containing target-region morphology, tumour-stroma interface, gland and lumen structure, cellularity, stable positive evidence, domain-style evidence to ignore, artifact or unreliable evidence, and negative shortcut evidence.”
3.4. Text-Conditioned Feature–Style Regularization
4. Experiments and Results
4.1. Dataset
4.2. Implementation Details
4.3. Evaluation Metrics
4.4. Results
4.5. Visualization and Ablation Study
5. Discussion
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Da, Q.; Wang, C.; Zuo, Y.; Guo, Y.; Jiang, G.; Shen, L.; Luo, X.; Ding, M.; Liu, J.; Hou, X.; et al. Cross-Organ and Cross-Scanner Adenocarcinoma Segmentation Challenge; Zenodo: Geneva, Switzerland, 2024. [Google Scholar] [CrossRef]
- Zhou, K.; Liu, Z.; Qiao, Y.; Xiang, T.; Loy, C.C. Domain generalization: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 4396–4415. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jahanifar, M.; Raza, M.; Xu, K.; Vuong, T.T.L.; Jewsbury, R.; Shephard, A.; Zamanitajeddin, N.; Kwak, J.T.; Raza, S.E.A.; Minhas, F.; et al. Domain generalization in computational pathology: Survey and guidelines. ACM Comput. Surv. 2025, 57, 285. [Google Scholar] [CrossRef] [Scilit]
- Meng, B.; Long, X.; Yang, W.; Liu, R.; Tian, Y.; Zheng, Y.; Liu, J. Advancing cross-organ domain generalization with test-time style transfer and diversity enhancement. In Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), Houston, TX, USA, 14–17 April 2025; pp. 1–5. [Google Scholar]
- Chen, R.J.; Ding, T.; Lu, M.Y.; Williamson, D.F.; Jaume, G.; Song, A.H.; Chen, B.; Zhang, A.; Shao, D.; Shaban, M.; et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 2024, 30, 850–862. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, Z.; Bianchi, F.; Yuksekgonul, M.; Montine, T.J.; Zou, J. A visual–language foundation model for pathology image analysis using medical twitter. Nat. Med. 2023, 29, 2307–2316. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, M.Y.; Chen, B.; Williamson, D.F.; Chen, R.J.; Liang, I.; Ding, T.; Jaume, G.; Odintsov, I.; Le, L.P.; Gerber, G.; et al. A visual-language foundation model for computational pathology. Nat. Med. 2024, 30, 863–874. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tellez, D.; Litjens, G.; Bándi, P.; Bulten, W.; Bokhorst, J.M.; Ciompi, F.; Van Der Laak, J. Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology. Med. Image Anal. 2019, 58, 101544. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hoque, M.Z.; Keskinarkaus, A.; Nyberg, P.; Seppänen, T. Stain normalization methods for histopathology image analysis: A comprehensive review and experimental comparison. Inf. Fusion 2024, 102, 101997. [Google Scholar] [CrossRef] [Scilit]
- Faryna, K.; van der Laak, J.; Litjens, G. Automatic data augmentation to improve generalization of deep learning in H&E stained histopathology. Comput. Biol. Med. 2024, 170, 108018. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, T.H.; Juyal, D.; Li, J.; Prakash, A.; Nofallah, S.; Shah, C.; Gullapally, S.C.; Yu, L.; Griffin, M.; Sampat, A.; et al. ContriMix: Scalable stain color augmentation for domain generalization without domain labels in digital pathology. In Proceedings of the MICCAI Workshop on Computational Pathology, Marrakesh, Morocco, 6–10 October 2024; Ciompi, F., Khalili, N., Studer, L., Poceviciute, M., Khan, A., Veta, M., Jiao, Y., Haj-Hosseini, N., Chen, H., Raza, S., et al., Eds.; Proceedings of Machine Learning Research: Cambridge, MA, USA, 2024; Volume 254, pp. 121–130. [Google Scholar]
- Konwer, A.; Prasanna, P. MetaStain: Stain-Generalizable Meta-learning for Cell Segmentation and Classification with Limited Exemplars. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 307–317. [Google Scholar]
- Aubreville, M.; Stathonikos, N.; Bertram, C.A.; Klopfleisch, R.; Ter Hoeve, N.; Ciompi, F.; Wilm, F.; Marzahl, C.; Donovan, T.A.; Maier, A.; et al. Mitosis domain generalization in histopathology images—The MIDOG challenge. Med. Image Anal. 2023, 84, 102699. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aubreville, M.; Stathonikos, N.; Donovan, T.A.; Klopfleisch, R.; Ammeling, J.; Ganz, J.; Wilm, F.; Veta, M.; Jabari, S.; Eckstein, M.; et al. Domain generalization across tumor types, laboratories, and species—Insights from the 2022 edition of the mitosis domain generalization challenge. Med. Image Anal. 2024, 94, 103155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, H.; Usuyama, N.; Bagga, J.; Zhang, S.; Rao, R.; Naumann, T.; Wong, C.; Gero, Z.; González, J.; Gu, Y.; et al. A whole-slide foundation model for digital pathology from real-world data. Nature 2024, 630, 181–188. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Meng, B.; Yang, W.; Long, X.; Wang, Y.; Dang, K.; Zheng, Y.; Liu, J. PathVLG: A Vision-Language Framework for Domain Generalization in Cross-Organ Adenocarcinoma Segmentation. In Proceedings of the 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Wuhan, China, 15–18 December 2025; pp. 5478–5485. [Google Scholar]
- Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Learning to prompt for vision-language models. Int. J. Comput. Vis. 2022, 130, 2337–2348. [Google Scholar] [CrossRef] [Scilit]
- Menon, S.; Vondrick, C. Visual classification via description from large language models. arXiv 2022, arXiv:2210.07183. [Google Scholar]
- Pratt, S.; Covert, I.; Liu, R.; Farhadi, A. What does a platypus look like? Generating customized prompts for zero-shot image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 15691–15701. [Google Scholar]
- Da, L.; Wang, R.; Xu, X.; Bhatia, P.; Kass-Hout, T.; Wei, H.; Xiao, C. FlanS: A Foundation Model for Free-Form Language-based Segmentation in Medical Images. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Toronto, ON, Canada, 3–7 August 2025; pp. 404–414. [Google Scholar]
- Zhao, Z.; Zhang, Y.; Wu, C.; Zhang, X.; Zhou, X.; Zhang, Y.; Wang, Y.; Xie, W. Large-vocabulary segmentation for medical images with text prompts. npj Digit. Med. 2025, 8, 566. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, K.; Yang, Y.; Hospedales, T.; Xiang, T. Learning to generate novel domains for domain generalization. In Proceedings of the European Conference on Computer Vision, Virtual, 23–28 August 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 561–578. [Google Scholar]
- Zhou, K.; Yang, Y.; Qiao, Y.; Xiang, T. Domain generalization with mixstyle. arXiv 2021, arXiv:2104.02008. [Google Scholar]
- Yang, Y.; Soatto, S. Fda: Fourier domain adaptation for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 4085–4095. [Google Scholar]
- Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 1290–1299. [Google Scholar]
- Yan, F.; Wu, J.; Li, J.; Wang, W.; Chen, Y.; Wei, L.; Lu, J.; Chen, W.; Gao, Z.; Li, J.; et al. Pathorchestra: A comprehensive foundation model for computational pathology with over 100 diverse clinical-grade tasks. npj Digit. Med. 2025, 8, 695. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kerssies, T.; Cavagnero, N.; Hermans, A.; Norouzi, N.; Averta, G.; Leibe, B.; Dubbelman, G.; De Geus, D. Your vit is secretly an image segmentation model. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 25303–25313. [Google Scholar]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision transformers for dense prediction. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Virtual, 11–17 October 2021; pp. 12159–12168. [Google Scholar]
- Taha, A.A.; Hanbury, A. Metrics for evaluating 3D medical image segmentation: Analysis, selection, and tool. BMC Med. Imaging 2015, 15, 29. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar]
- Zhou, Z.; Rahman Siddiquee, M.M.; Tajbakhsh, N.; Liang, J. Unet++: A nested u-net architecture for medical image segmentation. In Proceedings of the International Workshop on Deep Learning in Medical Image Analysis, Granada, Spain, 20 September 2018; Springer: Berlin/Heidelberg, Germany, 2018; pp. 3–11. [Google Scholar]
- Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-unet: Unet-like pure transformer for medical image segmentation. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 205–218. [Google Scholar]
- Benigmim, Y.; Roy, S.; Essid, S.; Kalogeiton, V.; Lathuilière, S. Collaborating foundation models for domain generalized semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 3108–3119. [Google Scholar]
- Hu, S.; Liao, Z.; Zhang, J.; Xia, Y. Domain and content adaptive convolution based multi-source domain generalization for medical image segmentation. IEEE Trans. Med. Imaging 2022, 42, 233–244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Q.; Zhang, K.Y.; Yao, T.; Lu, X.; Ding, S.; Ma, L. Test-time domain generalization for face anti-spoofing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 175–187. [Google Scholar]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. Dinov2: Learning robust visual features without supervision. arXiv 2023, arXiv:2304.07193. [Google Scholar]
- Siméoni, O.; Vo, H.V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al. Dinov3. arXiv 2025, arXiv:2508.10104. [Google Scholar]
- Reid, M.D.; Balci, S.; Ohike, N.; Xue, Y.; Kim, G.E.; Tajiri, T.; Memis, B.; Coban, I.; Dolgun, A.; Krasinskas, A.M.; et al. Ampullary carcinoma is often of mixed or hybrid histologic type: An analysis of reproducibility and clinical relevance of classification as pancreatobiliary versus intestinal in 232 cases. Mod. Pathol. 2016, 29, 1575–1585. [Google Scholar] [CrossRef] [Scilit] [PubMed]


| Model | Seen | Unseen | Overall | |||
|---|---|---|---|---|---|---|
| IoU | Dice | IoU | Dice | IoU | Dice | |
| Conventional segmentation models | ||||||
| Unet [30] | 74.37 | 85.30 | 73.03 | 84.41 | 73.70 | 84.86 |
| Unet++ [31] | 74.09 | 85.12 | 71.79 | 83.58 | 72.94 | 84.35 |
| Swin-UNet [32] | 64.96 | 78.76 | 68.33 | 81.19 | 66.65 | 79.99 |
| Domain generalization methods | ||||||
| CFM [33] | 74.27 | 85.24 | 69.93 | 82.30 | 72.10 | 83.79 |
| DCAC [34] | 75.94 | 86.32 | 76.42 | 86.63 | 76.18 | 86.48 |
| TTDG [35] | 76.44 | 86.65 | 72.38 | 83.98 | 74.41 | 85.33 |
| T3s [4] | 79.25 | 88.42 | 73.92 | 85.01 | 76.59 | 86.74 |
| Foundation model-based baselines | ||||||
| DINOv2 + EoMT [36] | 81.06 | 89.54 | 75.14 | 85.81 | 78.40 | 87.89 |
| DINOv3 + EoMT [37] | 79.27 | 88.44 | 75.08 | 85.77 | 77.50 | 87.32 |
| UNI2-h +Shallow Decoder [5] | 72.56 | 84.10 | 71.66 | 83.49 | 72.57 | 84.11 |
| CONCH + Shallow Decoder [7] | 71.08 | 83.10 | 73.08 | 84.45 | 72.59 | 84.12 |
| PathOrchestra + Shallow Decoder [26] | 81.70 | 89.93 | 76.96 | 86.98 | 79.60 | 88.64 |
| Ours | 82.26 | 90.27 | 78.31 | 87.84 | 80.56 | 89.23 |
| Model | Colorectum | Stomach | Pancreas | |||
|---|---|---|---|---|---|---|
| IoU | Dice | IoU | Dice | IoU | Dice | |
| Conventional segmentation models | ||||||
| Unet [30] | 77.61 | 87.39 | 78.15 | 87.74 | 67.35 | 80.49 |
| Unet++ [31] | 75.87 | 86.28 | 78.90 | 88.21 | 67.49 | 80.59 |
| Swin-UNet [32] | 74.98 | 85.70 | 64.39 | 78.34 | 55.51 | 71.39 |
| Domain generalization methods | ||||||
| CFM [33] | 76.51 | 86.69 | 79.46 | 88.55 | 66.83 | 80.12 |
| DCAC [34] | 78.62 | 88.03 | 80.27 | 89.06 | 68.94 | 81.61 |
| TTDG [35] | 81.73 | 89.95 | 80.02 | 88.90 | 67.56 | 80.64 |
| T3s [4] | 82.52 | 90.42 | 80.31 | 89.08 | 74.92 | 85.66 |
| Foundation model-based baselines | ||||||
| DINOv2 + EoMT [36] | 85.82 | 92.37 | 80.27 | 89.06 | 77.15 | 87.10 |
| DINOv3 + EoMT [37] | 84.10 | 91.36 | 78.20 | 87.77 | 75.56 | 86.08 |
| UNI2-h + Shallow Decoder [5] | 77.70 | 87.45 | 71.15 | 83.14 | 68.83 | 81.54 |
| CONCH + Shallow Decoder [7] | 74.89 | 85.64 | 71.67 | 83.50 | 66.62 | 79.97 |
| PathOrchestra + Shallow Decoder [26] | 84.90 | 91.83 | 83.99 | 91.30 | 76.24 | 86.52 |
| Ours | 84.74 | 91.74 | 84.67 | 91.70 | 77.37 | 87.24 |
| Model | Ampullary | Gallbladder | Intestine | |||
|---|---|---|---|---|---|---|
| IoU | Dice | IoU | Dice | IoU | Dice | |
| Conventional segmentation models | ||||||
| Unet [30] | 60.86 | 75.67 | 78.36 | 87.87 | 79.86 | 88.80 |
| Unet++ [31] | 61.03 | 75.80 | 71.64 | 83.48 | 82.71 | 90.54 |
| Swin-UNet [32] | 49.77 | 66.46 | 74.42 | 85.33 | 80.79 | 89.37 |
| Domain generalization methods | ||||||
| CFM [33] | 62.18 | 76.68 | 67.72 | 80.75 | 79.90 | 88.82 |
| DCAC [34] | 63.15 | 77.41 | 81.57 | 89.85 | 84.54 | 91.62 |
| TTDG [35] | 53.32 | 69.55 | 80.52 | 89.21 | 83.31 | 90.89 |
| T3s [4] | 53.32 | 69.55 | 83.97 | 91.29 | 84.48 | 91.59 |
| Foundation-model-based baselines | ||||||
| DINOv2 + EoMT [36] | 68.44 | 81.26 | 65.19 | 78.93 | 87.29 | 93.21 |
| DINOv3 + EoMT [37] | 66.40 | 79.81 | 72.07 | 83.77 | 83.13 | 90.79 |
| UNI2-h + Shallow Decoder [5] | 63.53 | 77.70 | 65.93 | 79.47 | 80.09 | 88.94 |
| CONCH + Shallow Decoder [7] | 64.06 | 78.09 | 67.98 | 80.94 | 79.78 | 88.75 |
| PathOrchestra + Shallow Decoder [26] | 70.58 | 82.75 | 72.89 | 84.32 | 82.72 | 90.54 |
| Ours | 71.53 | 83.40 | 74.17 | 85.17 | 84.04 | 91.33 |
| Model | Params (M) | FLOPs (G) | Latency (ms) |
|---|---|---|---|
| Conventional segmentation models | |||
| Unet [30] | 31.032 | 167.255 | |
| Unet++ [31] | 36.615 | 422.101 | |
| Swin-UNet [32] | 27.168 | 6.130 | |
| Domain generalization methods | |||
| CFM [33] | 218.0 | 172.424 | |
| DCAC [34] | 15.833 | 37.630 | |
| TTDG [35] | 58.095 | 274.032 | |
| T3s [4] | 80.508 | 274.032 | |
| Foundation model-based baselines | |||
| DINOv2 + EoMT [27,36] | 93.171 | 170.506 | |
| DINOv3 + EoMT [27,37] | 23.273 | 36.883 | |
| UNI2-h + shallow decoder [5] | 683.144 | 311.962 | |
| CONCH + shallow decoder [7] | 87.804 | 84.895 | |
| PathOrchestra + shallow decoder [26] | 306.236 | 371.586 | |
| Ours | 110.727 | 96.427 | |
| Image-Aware Query | Evidence Recal. | Positive Ev. | Negative Ev. | IoU | Dice |
|---|---|---|---|---|---|
| – | – | – | – | 64.49 | 78.41 |
| ✓ | – | – | – | 72.19 | 83.13 |
| ✓ | ✓ | – | ✓ | 73.90 | 84.99 |
| ✓ | ✓ | ✓ | – | 77.18 | 87.12 |
| ✓ | ✓ | ✓ | ✓ | 80.56 | 89.23 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Meng, B.; Wang, J.; Liu, J. Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding. Electronics 2026, 15, 3691. https://doi.org/10.3390/electronics15163691
Meng B, Wang J, Liu J. Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding. Electronics. 2026; 15(16):3691. https://doi.org/10.3390/electronics15163691
Chicago/Turabian StyleMeng, Biwen, Jiahao Wang, and Jingxin Liu. 2026. "Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding" Electronics 15, no. 16: 3691. https://doi.org/10.3390/electronics15163691
APA StyleMeng, B., Wang, J., & Liu, J. (2026). Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding. Electronics, 15(16), 3691. https://doi.org/10.3390/electronics15163691

