Bridging Perception and Reasoning: An Evidence-Based Agentic System for Diagnosis and Treatment Recommendations of Vascular Anomalies
Abstract
1. Introduction
- We construct a high-quality, expert-annotated cohort of 7565 VA cases to conduct a preliminary feasibility study of SOTA open-source MLLMs. This preparatory evaluation exposes the significant limitations of generalist models in specialized diagnostics.
- We introduce HevaDx, a novel modular system that explicitly decouples the clinical workflow into a visual specialist and a reasoning specialist. By combining a lightweight, adaptable visual encoder with a knowledge-augmented LLM, we achieve superior diagnostic accuracy.
- We establish a rigorous pipeline for dataset construction, incorporating strict quality control, Region of Interest (ROI) annotation, and class balancing strategies, mitigating the long-tail distribution problem inherent in clinical data. We also validate that a Retrieval-Augmented Generation (RAG) mechanism enhances clinical safety by transforming opaque model outputs into transparent, evidence-based reasoning chains grounded in medical guidelines.
2. Materials and Methods
2.1. The Large-Scale VA Dataset and Evaluations on Advanced MLLMs
2.1.1. Data Collection and Annotation
2.1.2. Quality Control and Dataset Statistics
2.1.3. Evaluations on Advanced Open-Source MLLMs
2.2. The HevaDx System
2.2.1. The Visual Specialist: Efficient Perception with DINOv2
2.2.2. The Reasoning Specialist: Transparent, Evidence-Based Decision Making
3. Data Preprocessing and Setup
3.1. Dataset Stratification and Balancing
3.2. Data Preprocessing
- Region of interest (ROI) annotation: We manually annotated bounding boxes for all lesion images. This step forces the model to focus its attention on the relevant pathological features, eliminating interference from background factors (e.g., clothing, medical equipment, or unrelated skin areas).
- Quality control: We conducted a secondary review to filter out low-quality samples. Images where the lesion location was ambiguous, or the diagnosis was clinically controversial, were excluded to prevent label noise.
3.3. Implementation Details
- Reasoning Specialist Setup: For the reasoning component, we employed a Retrieval-Augmented Generation (RAG) framework. We constructed a specialized external knowledge base derived from the physician-summarized Guidelines for Diagnosis and Treatment of Hemangiomas and Vascular Malformations (2024 Edition) [5]. This ensures that the LLM’s (Qwen2.5-7B-Instruct) treatment recommendations are grounded in the latest clinical evidence. Detailed specifications regarding the chunking strategy, the rationale for choosing RAG, the knowledge update mechanism, and the prompt are provided in Appendix A.
- Metrics: To comprehensively assess system performance, we report the top-1 and top-3 accuracy for both the diagnosis task and the treatment recommendation task. Given that VA management often involves a hierarchy of valid therapeutic options (e.g., strictly observing a stable lesion is a valid alternative to laser therapy), we interpret these metrics through a nuanced lens. Specifically, top-1 accuracy quantifies the system’s precise alignment with the primary consensus gold standard (the adjudicated “best” option). Complementing this, top-3 accuracy serves as the critical indicator of clinical admissibility and guideline consistency. This metric assesses whether the consensus gold standard is retained within the model’s high-probability candidates, which captures the system’s alignment with the broader therapeutic consensus, accommodating the flexibility of clinical decision-making beyond a rigid single-label prediction. The evaluation was conducted using the independent test set described in Section 3.1. Note that since the reasoning specialist need to receive diagnostic results from the visual specialist to make further actions, the top-1 and top-3 accuracy for treatment recommendations are both based on the top-1 diagnosis. Furthermore, to provide a robust measure of statistical uncertainty, we report the 95% confidence intervals [CIs] for all metrics, estimated via a non-parametric bootstrap procedure with 1000 iterations. we also report class-specific Recall, Precision, and F1-score.
4. Results
4.1. Main Results
4.2. Ablation Study on Data Preprocessing
- Resolution of Clinical Mimicry: The normalized confusion matrices further elucidate how preprocessing mitigates phenotypic confusion. As shown in Figure 3A (Before Preprocessing), the baseline model struggled with clinical mimicry, appearing unable to distinguish intrinsic lesion features from background noise. For instance, in the raw setting, PWS was frequently misclassified as VM (8 out of 17 cases), resulting in a recall of only 0.29. Similarly, VVM was heavily confused with PWS and VH, achieving a recall of just 0.13.In contrast, Figure 3B (after preprocessing) demonstrates strong diagonal dominance, indicating robust correct classification. The rigorous ROI annotation forced the visual encoder to attend to fine-grained texture and boundary features rather than background artifacts. Consequently, the confusion between PWS and VM was drastically reduced (only 2 misclassified), raising the PWS recall to 0.82. Although some confusion persists between the highly similar “verrucous” subtypes (VH and VVM), the overall class separability has been significantly enhanced, confirming that high-quality data curation is a prerequisite for resolving the long-tail distribution in vascular anomaly diagnosis.Figure 3. Resolution of clinical mimicry via data preprocessing. The model trained on preprocessed data demonstrates strong diagonal dominance, indicating robust correct classification.Figure 3. Resolution of clinical mimicry via data preprocessing. The model trained on preprocessed data demonstrates strong diagonal dominance, indicating robust correct classification.
- Enhancement of Discriminative Capability: The quantitative improvement is visualized in Figure 4. The model trained on preprocessed data exhibited dramatic performance gains across all disease categories. Notably, the F1-score for PWS surged from 0.39 to 0.88, and IH improved from 0.50 to 0.88. Even for morphologically complex subtypes like VH, which previously suffered from extremely low recognition (F1 = 0.21), the preprocessing strategy restored the model’s discriminative capability, raising the F1-score to 0.64.
4.3. Qualitative Analysis on Reasoning Specialist
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| VA | Vascular Anomaly |
| PWS | Port-wine stain |
| IH | Infantile Hemangioma |
| VM | Venous Malformation |
| VH | Verrucous Hemangioma |
| VVM | Verrucous Venous Malformation |
| LLM | Large Language Model |
| MLLM | Multimodal Large Language Model |
| RAG | Retrieval-Augmented Generation |
| SOTA | State-of-the-art |
| ICL | In-Context Learning |
| ROI | Region of Interest |
Appendix A. Technical Implementation of the Reasoning Specialist
Appendix A.1. Disease-Centric Chunking


Appendix A.2. Rationale for RAG over Rule-Based Systems
- Decoupling Inference from Knowledge Management: A rule-based approach necessitates maintaining a synchronized code module that explicitly maps diagnostic outputs to document indices. As the system may expand to cover more subtypes in the ISSVA classification, this hard-coded logic would introduce significant deployment complexity and maintenance overhead. RAG effectively decouples the inference engine from the knowledge repository. This allows for the seamless expansion of disease categories by simply ingesting new vectorized chunks into the database, without requiring any modifications to the underlying deployment code or inference logic.
- Adaptability to Heterogeneous Knowledge Sources: Currently, the knowledge base is organized around disease-centric chunks. However, future iterations of the system could incorporate heterogeneous information sources that are not strictly disease-specific, such as symptom-management protocols, cross-disease contraindications, and unstructured expert clinical notes. Rule-based systems, which rely on rigid key-value pairing (i.e., Diagnosis → Guideline), fail to retrieve relevant context when the information is not taxonomically indexed by a specific disease name. In contrast, the semantic retrieval mechanism of RAG is agnostic to the data structure, allowing it to flexibly retrieve and synthesize context from diverse sources based on semantic relevance rather than rigid rules.
Appendix A.3. Knowledge Update Mechanism
Appendix B. Justification of Test Set Size and Statistical Robustness
Appendix B.1. Data Distribution Constraints
Appendix B.2. Statistical Uncertainty Quantification
References
- Kanada, K.N.; Merin, M.R.; Munden, A.; Friedlander, S.F. A prospective study of cutaneous findings in newborns in the United States: Correlation with race, ethnicity, and gestational status using updated classification and nomenclature. J. Pediatr. 2012, 161, 240–245. [Google Scholar] [CrossRef] [Scilit]
- Johnson, A.B.; Richter, G.T. Vascular Anomalies. Clin. Perinatol. 2018, 45, 737–749. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Queisser, A.; Seront, E.; Boon, L.M.; Vikkula, M. Genetic Basis and Therapies for Vascular Anomalies. Circ. Res. 2021, 129, 155–173. [Google Scholar] [CrossRef] [Scilit]
- Kunimoto, K.; Yamamoto, Y.; Jinnin, M. ISSVA Classification of Vascular Anomalies and Molecular Biology. Int. J. Mol. Sci. 2022, 23, 2358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chinese Society for the Study of Vascular Anomalies (CSSVA). Guidelines for the diagnosis and treatment of hemangiomas and vascular malformations (2024 edition). J. Tissue Eng. Reconstr. Surg. 2024, 20, 1–50. [Google Scholar]
- Sebaratnam, D.F.; Rodríguez Bandera, A.L.; Wong, L.C.F.; Wargon, O. Infantile hemangioma. Part 2: Management. J. Am. Acad. Dermatol. 2021, 85, 1395–1404. [Google Scholar] [CrossRef] [Scilit]
- Liu, L.; Li, X.; Zhao, Q.; Yang, L.; Jiang, X. Pathogenesis of Port-Wine Stains: Directions for Future Therapies. Int. J. Mol. Sci. 2022, 23, 12139. [Google Scholar] [CrossRef] [Scilit]
- Greene, A.K.; Alomari, A.I. Management of venous malformations. Clin. Plast. Surg. 2011, 38, 83–93. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Jiang, H.; Liu, Y.H.; Ma, C.Y.; Zhang, X.; Pan, Y.; Liu, M.; Gu, P.; Xia, S.; Li, W.; et al. A Comprehensive Review of Multimodal Large Language Models: Performance and Challenges Across Different Tasks. arXiv 2024, arXiv:2408.01319. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Gan, W.; Chen, Z.; Wan, S.; Yu, P.S. Multimodal large language models: A survey. In Proceedings of the 2023 IEEE International Conference on Big Data (BigData), Sorrento, Italy, 15-18 December 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 2247–2256. [Google Scholar]
- Xu, P.J.; Kan, S.X.; Jin, J.; Zhang, Z.J.; Gu, Y.X.; Zhang, B.; Zhou, Y.L. Multimodal large language models in medical research and clinical practice: Development, applications, challenges and future. Neurocomputing 2025, 660, 131817. [Google Scholar] [CrossRef] [Scilit]
- Ye, J.; Tang, H. Multimodal Large Language Models for Medicine: A Comprehensive Survey. arXiv 2025, arXiv:2504.21051. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Wu, C.; Fan, Y.; Qiu, P.; Zhang, X.; Sun, Y.; Zhou, X.; Zhang, S.; Peng, Y.; Wang, Y.; et al. An Agentic System for Rare Disease Diagnosis with Traceable Reasoning. Nature 2025, 1–23. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Ma, Z.; Wang, Z.; Wu, C.; Chen, W.; Li, X.; Yuan, Y. A Survey of LLM-based Agents in Medicine: How far are we from Baymax? In Proceedings of the Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models from Natural Language Supervision. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021. [Google Scholar]
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. arXiv 2022, arXiv:2203.02155. [Google Scholar] [CrossRef] [Scilit]
- van de Ven, G.M.; Soures, N.; Kudithipudi, D. Continual Learning and Catastrophic Forgetting. arXiv 2024, arXiv:2403.05175. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Ding, L.; Fang, M.; Tao, D. Revisiting Catastrophic Forgetting in Large Language Model Tuning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA, 12–16 November 2024; pp. 4297–4308. [Google Scholar]
- Zhang, Y.; Chen, M.; Chen, S.; Peng, B.; Zhang, Y.; Li, T.; Lu, C. CauSight: Learning to Supersense for Visual Causal Discovery. arXiv 2025, arXiv:2512.01827. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.V.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv 2023, arXiv:2304.07193. [Google Scholar]
- Yang, Q.A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Dong, G.; et al. Qwen2.5 Technical Report. arXiv 2024, arXiv:2412.15115. [Google Scholar] [CrossRef] [Scilit]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Kuttler, H.; Lewis, M.; tau Yih, W.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv 2020, arXiv:2005.11401. [Google Scholar]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Guo, Q.; Wang, M.; et al. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2023, arXiv:2312.10997. [Google Scholar]
- Li, Y.; Zhang, W.; Yang, Y.; Huang, W.C.; Wu, Y.; Luo, J.; Bei, Y.Q.; Zou, H.P.; Luo, X.; Zhao, Y.; et al. Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs. arXiv 2025, arXiv:2507.09477. [Google Scholar] [CrossRef] [Scilit]
- Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. Qwen2.5-VL Technical Report. arXiv 2025, arXiv:2502.13923. [Google Scholar] [CrossRef] [Scilit]
- Du, K.T.A.; Yin, B.; Xing, B.; Qu, B.; Wang, B.; Chen, C.; Zhang, C.; Du, C.; Wei, C.; Wang, C.; et al. Kimi-VL Technical Report. arXiv 2025, arXiv:2504.07491. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual Instruction Tuning. arXiv 2023, arXiv:2304.08485. [Google Scholar]
- Sellergren, A.; Kazemzadeh, S.; Jaroensri, T.; Kiraly, A.P.; Traverse, M.; Kohlberger, T.; Xu, S.; Jamil, F.; Hughes, C.; Lau, C.; et al. MedGemma Technical Report. arXiv 2025, arXiv:2507.05201. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Wong, C.; Zhang, S.; Usuyama, N.; Liu, H.; Yang, J.; Naumann, T.; Poon, H.; Gao, J. LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. arXiv 2023, arXiv:2306.00890. [Google Scholar]
- Li, J.; Li, D.; Savarese, S.; Hoi, S.C.H. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023. [Google Scholar]
- Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; Li, L.; Sui, Z. A Survey on In-context Learning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022. [Google Scholar]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. arXiv 2020, arXiv:2005.14165. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Li, J.; Xiang, Y.; Yan, H.; Gui, L.; He, Y. The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis. In Proceedings of the Conference on Empirical Methods in Natural7 Language Processing, Singapore, 6–10 December 2023. [Google Scholar]
- Alansari, A.; Luqman, H. Large Language Models Hallucination: A Comprehensive Survey. arXiv 2025, arXiv:2510.06265. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Yuan, Y.; Zhang, Z. Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases. arXiv 2024, arXiv:2403.10446. [Google Scholar] [CrossRef] [Scilit]
- Loshchilov, I.; Hutter, F. Fixing Weight Decay Regularization in Adam. arXiv 2017, arXiv:1711.05101. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. CoRR 2014, arXiv:1412.6980. [Google Scholar]




| Characteristic | Value/Details |
|---|---|
| Patients | 7565 Cases |
| Disease | 14 Types |
| Treatment | 8 Options |
| Gender | Male & Female |
| Age | 0–70 (Years) |
| Lesions | 154 Body Sites |
| Model | DiagAcc@1 | DiagAcc@3 | TreatAcc@1 | TreatAcc@3 |
|---|---|---|---|---|
| General-purpose models | ||||
| Qwen2.5-VL-7B-Instruct | 48.8 | 61.9 | 10.6 | 30.6 |
| Qwen2.5-VL-32B-Instruct | 58.1 | 67.5 | 16.3 | 35.6 |
| LLaVA-v1.5-7B | 14.4 | 40.6 | 8.8 | 15.6 |
| Kimi-VL-16B | 43.1 | 62.5 | 15.6 | 38.1 |
| Medical-specialized models | ||||
| MedGemma-4B | 6.3 | 26.9 | 11.9 | 26.9 |
| MedGemma-27B | 35.6 | 66.9 | 13.1 | 36.9 |
| LLaVA-Med-v1.5-Mistral-7B | 22.5 | 53.1 | 10.6 | 28.1 |
| Method | DiagAcc@1 | DiagAcc@3 | TreatAcc@1 | TreatAcc@3 |
|---|---|---|---|---|
| HevaDx | 75.0 [66.7–83.3] | 94.8 [89.6–99.0] | 62.5 [53.1–71.9] | 83.3 [75.0–90.6] |
| Category | Precision | Recall | F1-Score |
|---|---|---|---|
| PWS | 0.583 | 0.824 | 0.683 |
| IH | 0.882 | 0.882 | 0.882 |
| VM | 0.765 | 0.722 | 0.743 |
| VH | 0.700 | 0.583 | 0.636 |
| VVM | 0.692 | 0.600 | 0.643 |
| nevi | 0.933 | 0.824 | 0.875 |
| Macro Average | 0.759 | 0.739 | 0.744 [0.644–0.828] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, Y.; Qiu, Y.; Lin, X. Bridging Perception and Reasoning: An Evidence-Based Agentic System for Diagnosis and Treatment Recommendations of Vascular Anomalies. Diagnostics 2026, 16, 621. https://doi.org/10.3390/diagnostics16040621
Zhang Y, Qiu Y, Lin X. Bridging Perception and Reasoning: An Evidence-Based Agentic System for Diagnosis and Treatment Recommendations of Vascular Anomalies. Diagnostics. 2026; 16(4):621. https://doi.org/10.3390/diagnostics16040621
Chicago/Turabian StyleZhang, Yize, Yajing Qiu, and Xiaoxi Lin. 2026. "Bridging Perception and Reasoning: An Evidence-Based Agentic System for Diagnosis and Treatment Recommendations of Vascular Anomalies" Diagnostics 16, no. 4: 621. https://doi.org/10.3390/diagnostics16040621
APA StyleZhang, Y., Qiu, Y., & Lin, X. (2026). Bridging Perception and Reasoning: An Evidence-Based Agentic System for Diagnosis and Treatment Recommendations of Vascular Anomalies. Diagnostics, 16(4), 621. https://doi.org/10.3390/diagnostics16040621

