VMMedSAM-X: A State-Enhanced Dual-Branch Encoder for Efficient Promptable Medical Image Segmentation
Abstract
1. Introduction
- A state-enhanced encoder for MedSAM: We redesign the MedSAM image encoder using a dual-branch architecture that integrates selective state space modeling and visual long-short-term memory (LSTM), enabling efficient modeling of long-range dependencies while reducing the computational burden compared to standard ViT-based encoders.
- A bidirectional cross-attention fusion mechanism: We propose a cross-attention module that facilitates mutual refinement of structural and contextual representations, thereby enabling dynamic alignment and synergistic fusion of complementary information across the two branches.
- Superior efficiency and competitive accuracy: Experimental results on five public datasets show that the proposed framework cuts the computational cost by 95.3% (from 369.44 GB to 17.36 GB) compared to the ViT-based MedSAM encoder. This is achieved while maintaining or improving segmentation accuracy across different modalities and anatomical structures.
2. Related Work
2.1. CNN-Based Medical Image Segmentation
2.2. Transformer-Based Segmentation Models
2.3. Foundation Models for Medical Image Segmentation
2.4. State Space Models in Vision
2.5. Recurrent Memory Models in Medical Image Segmentation
2.6. Summary
3. Materials and Methods
3.1. Overall Architecture
3.2. Multi-Scale State-Enhanced Encoder
3.3. Dual-Path Cross-Attention Mechanism
3.4. Experimental Setup and Data Preprocessing
3.4.1. Datasets
- segTHOR [43]: 120 thoracic CT scans (80 training, 40 testing) for organ segmentation of five structures including trachea and esophagus.
- AMOS22 [44]: 800 abdominal CT/MRI cases with 15 annotated abdominal organs.
- ACDC [45]: 200 cardiac MRI images for cardiovascular structure segmentation.
- MSD-Lung [46]: 96 thin-slice CT scans of non-small cell lung cancer patients for lung tumor segmentation.
- KiTS23 [47]: 599 CT cases of kidney and tumors for renal and tumor segmentation.
3.4.2. Data Preprocessing
- Label cleaning and small structure removal: 3D connected component analysis is used to remove artifact structures smaller than 1000 pixels, and small regions smaller than 100 pixels are removed at the 2D slice level. Multi-tumor cases are independently instance-labeled to ensure semantic consistency.
- Intensity normalization: For CT data, clinical window width and level (WL = 40, WW = 400) are applied for intensity clipping, followed by unified normalization to [0, 255]. For MRI data, non-background region windowing using the 0.5–99.5% percentile range is applied to suppress extreme noise.
- Effective slice extraction: Slices containing organs/lesions are automatically cropped by detecting non-zero voxel ranges in annotations, removing irrelevant empty slices to improve training efficiency.
- Spatial resolution unification: All slices are interpolated to a unified resolution of using nearest-neighbor interpolation to avoid label boundary smoothing.
- Sample expansion: Each 3D case is transformed into multiple 2D slices with clear anatomical semantics after preprocessing, expanding the training sample scale from case-level to slice-level, effectively alleviating the sample scarcity problem faced by deep learning models in medical imaging scenarios.
3.4.3. Prompting Protocol
3.4.4. Experimental Environment
4. Results
4.1. Quantitative Evaluation on Multiple Datasets
4.1.1. segTHOR Thoracic Organ Segmentation
4.1.2. MSD-Lung Lung Tumor Segmentation
4.1.3. KiTS23 Kidney Tumor Segmentation
4.1.4. AMOS22 Abdominal Multi-Organ Segmentation
4.1.5. ACDC Cardiac MRI Segmentation
4.2. Computational Complexity Analysis
4.3. Ablation Study
4.4. Summary of Experimental Findings
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| ASD | Average Surface Distance |
| CNN | Convolutional Neural Network |
| CT | Computed Tomography |
| DSC | Dice Similarity Coefficient |
| HD95 | 95% Hausdorff Distance |
| IoU | Intersection over Union |
| LSTM | Long Short-Term Memory |
| MedSAM | Medical Segment Anything Model |
| MRI | Magnetic Resonance Imaging |
| MSD | Medical Segmentation Decathlon |
| SAM | Segment Anything Model |
| SS2D | Two-Dimensional Selective Scanning |
| SSM | State Space Model |
| U-Net | U-shaped Convolutional Network |
| ViT | Vision Transformer |
| VM-UNet | Vision Mamba U-Net |
| xLSTM | Extended Long Short-Term Memory |
Appendix A. Per-Class Segmentation Results
| Dataset | Organ/Structure | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|---|
| segTHOR | Heart | 83.87 ± 7.59 | 72.89 ± 10.30 | 4.21 ± 1.52 | 1.78 ± 0.61 |
| Aorta | 95.32 ± 3.15 | 91.22 ± 5.51 | 2.48 ± 0.91 | 0.91 ± 0.33 | |
| Trachea | 91.48 ± 4.30 | 84.57 ± 6.71 | 4.79 ± 1.68 | 1.63 ± 0.57 | |
| Esophagus | 90.73 ± 4.52 | 84.25 ± 6.25 | 3.92 ± 1.41 | 1.28 ± 0.46 | |
| Average (slice-level, reported in Table 1) | 90.8 ± 0.3 | 84.0 ± 0.4 | 3.87 ± 0.12 | 1.34 ± 0.06 | |
| ACDC | Left Ventricle (LV) | 96.8 ± 1.5 | 93.8 ± 2.2 | 0.97 ± 0.35 | 0.32 ± 0.11 |
| Right Ventricle (RV) | 93.5 ± 2.0 | 87.8 ± 3.1 | 1.42 ± 0.51 | 0.51 ± 0.18 | |
| Myocardium (MYO) | 96.2 ± 1.8 | 92.7 ± 2.8 | 1.15 ± 0.41 | 0.40 ± 0.14 | |
| Average (slice-level, reported in Table 4) | 95.7 ± 0.2 | 92.0 ± 0.3 | 1.18 ± 0.05 | 0.41 ± 0.03 | |
| AMOS22 | Spleen | 93.91 ± 6.21 | 89.04 ± 9.03 | 2.45 ± 0.88 | 0.82 ± 0.29 |
| Right Kidney | 94.12 ± 4.11 | 89.14 ± 6.55 | 2.78 ± 1.02 | 0.95 ± 0.34 | |
| Left Kidney | 94.25 ± 3.25 | 89.30 ± 5.53 | 2.82 ± 0.99 | 0.98 ± 0.33 | |
| Gallbladder | 87.49 ± 8.11 | 77.57 ± 11.43 | 6.12 ± 2.21 | 2.15 ± 0.78 | |
| Esophagus | 82.16 ± 10.06 | 70.81 ± 12.88 | 7.45 ± 2.63 | 2.72 ± 0.98 | |
| Liver | 94.09 ± 5.08 | 89.21 ± 7.86 | 2.53 ± 0.91 | 0.88 ± 0.31 | |
| Stomach | 89.22 ± 8.42 | 81.43 ± 11.94 | 4.72 ± 1.68 | 1.62 ± 0.58 | |
| Aorta | 93.37 ± 4.64 | 87.88 ± 7.22 | 2.61 ± 0.94 | 0.91 ± 0.32 | |
| Inferior Vena Cava | 87.30 ± 8.59 | 76.38 ± 12.21 | 3.48 ± 1.25 | 1.22 ± 0.44 | |
| Pancreas | 80.60 ± 11.22 | 68.79 ± 13.84 | 7.18 ± 2.58 | 2.56 ± 0.92 | |
| Right Adrenal Gland | 71.38 ± 14.68 | 57.35 ± 16.38 | 9.45 ± 3.40 | 3.38 ± 1.22 | |
| Left Adrenal Gland | 77.20 ± 8.90 | 62.90 ± 11.20 | 8.92 ± 3.21 | 3.18 ± 1.15 | |
| Duodenum | 79.03 ± 12.03 | 66.78 ± 14.84 | 8.78 ± 3.16 | 3.10 ± 1.12 | |
| Bladder | 90.32 ± 8.88 | 83.33 ± 12.36 | 3.25 ± 1.17 | 1.14 ± 0.41 | |
| Prostate/Uterus | 90.88 ± 5.98 | 83.79 ± 9.24 | 3.08 ± 1.11 | 1.08 ± 0.39 | |
| Average (slice-level, reported in Table 4) | 93.1 ± 0.4 | 87.5 ± 0.5 | 3.27 ± 0.15 | 1.38 ± 0.07 |
References
- Xu, G.; Udupa, J.K.; Luo, J.; Zhao, S.; Yu, Y.; Raymond, S.B.; Peng, H.; Ning, L.; Rathi, Y.; Liu, W.; et al. Is the medical image segmentation problem solved? A survey of current developments and future directions. arXiv 2025, arXiv:2508.20139. [Google Scholar] [CrossRef] [Scilit]
- Litjens, G.; Kooi, T.; Bejnordi, B.E.; Setio, A.A.A.; Ciompi, F.; Ghafoorian, M.; Van Der Laak, J.A.W.M.; Van Ginneken, B.; Sánchez, C.I. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Çiçek, Ö.; Abdulkadir, A.; Lienkamp, S.S.; Brox, T.; Ronneberger, O. 3D U-Net: Learning dense volumetric segmentation from sparse annotation. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Cham, Switzerland, 2016; pp. 424–432. [Google Scholar]
- Kabil, A.; Khoriba, G.; Yousef, M.; Rashed, E.A. Advances in medical image segmentation: A comprehensive survey with a focus on lumbar spine applications. Comput. Biol. Med. 2025, 198, 111171. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 5 October 2021. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Azad, R.; Rauf, Z.; Khan, A.R.; Rathore, S.; Khan, S.H.; Shah, N.S.; Farooq, U.; Asif, H.; Asif, A.; Zahoora, U.; et al. A recent survey of vision transformers for medical image segmentation. arXiv 2024, arXiv:2402.04899. [Google Scholar]
- Kumar, S.S. Advancements in medical image segmentation: A review of transformer models. Comput. Electr. Eng. 2025, 123, 110099. [Google Scholar] [CrossRef] [Scilit]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y.; et al. Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2023; pp. 4015–4026. [Google Scholar]
- Ma, J.; He, Y.; Li, F.; Han, L.; You, C.; Wang, B. Segment anything in medical images. Nat. Commun. 2024, 15, 654. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fan, W.; Li, A.; Xu, M.; Sun, W.; Man, F. Practical application of SAM for breast nodules segmentation. Front. Oncol. 2026, 16, 1756011. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, E.; Liu, H.; Ruan, D. Necessity and impact of specialization of large foundation model for medical segmentation tasks. Med. Phys. 2024, 52, 321–328. [Google Scholar] [CrossRef] [Scilit]
- Awad, M.A.; Mabrouk, M.S.; Elnokrashy, A.F. Medical image segmentation using transformer encoders and prompt-based learning: A systematic review of adaptation strategies, performance, and challenges. In Proceedings of the 2025 Twelfth International Conference on Intelligent Computing and Information Systems (ICICIS); IEEE: New York, NY, USA, 2025; pp. 281–288. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef] [Scilit]
- Zhang, R.; Guo, H.; Tian, K.; Zhou, J.; Yan, M.; Zhang, Z.; Zhao, S. Unified medical image segmentation with state space modeling snake. In Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM); Association for Computing Machinery: New York, NY, USA, 2025; pp. 7825–7834. [Google Scholar]
- Dong, X.; Zhou, B.; Yin, C.; Liao, I.Y.; Jin, Z.; Xu, Z. ÆMMamba: An efficient medical segmentation model with edge enhancement. IEEE J. Biomed. Health Inform. 2026, 30, 1889–1901. [Google Scholar] [CrossRef] [Scilit]
- Tang, F.; Nian, B.; Li, Y.; Jiang, Z.; Yang, J.; Liu, W.; Zhou, K. MambaMIM: Pre-training Mamba with state space token interpolation and its application to medical image segmentation. Med. Image Anal. 2025, 103, 103606. [Google Scholar] [CrossRef] [Scilit]
- Beck, M.; Pöppel, K.; Spanring, M.; Auer, A.; Prudnikova, O.; Kopp, M.; Klambauer, G.; Brandstetter, J.; Hochreiter, S. xLSTM: Extended long short-term memory. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 107547–107603. [Google Scholar]
- Dutta, P.; Bose, S.; Roy, S.K.; Mitra, S. Are Vision-xLSTM-embedded U-Nets better at segmenting medical images? Neural Netw. 2025, 192, 107925. [Google Scholar] [CrossRef] [Scilit]
- Shen, D.; Wu, G.; Suk, H.I. Deep learning in medical image analysis. Annu. Rev. Biomed. Eng. 2017, 19, 221–248. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Siddiquee, M.M.R.; Tajbakhsh, N.; Liang, J. UNet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Trans. Med. Imaging 2018, 39, 1856–1867. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention U-Net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar] [CrossRef] [Scilit]
- Abdellaoui, C.; Belkacem, S.; Messaoudi, N. Deep learning architectures for medical image segmentation: An organized analysis of CNN-based models and uses. Bull. Electr. Eng. Inform. 2026, 15, 424–437. [Google Scholar] [CrossRef] [Scilit]
- Oliveira, J.V.S.; Vieira, D.F.; Silva, M.P.; Fernandes, D.L.; Ribeiro, M.H.F.; Oliveira, H.N. Strategies for deep learning in volumetric medical imaging: A survey. In Proceedings of SIBGRAPI; IEEE: New York, NY, USA, 2025. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021. [Google Scholar]
- Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient transformers: A survey. ACM Comput. Surv. 2022, 55, 109. [Google Scholar] [CrossRef] [Scilit]
- Kim, J.W.; Khan, A.U.; Banerjee, I. Systematic review of hybrid vision transformer architectures for radiological image analysis. J. Imaging Inform. Med. 2025, 38, 3248–3262. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Moglia, A.; Leccardi, M.; Cavicchioli, M.; Maccarini, A.; Marcon, M.; Mainardi, L.; Cerveri, P. Generalist models in medical image segmentation: A survey and performance comparison with task-specific approaches. Inf. Fusion 2026, 127, 103709. [Google Scholar] [CrossRef] [Scilit]
- Noh, S.; Lee, B.-D. A narrative review of foundation models for medical image segmentation: Zero-shot performance evaluation on diverse modalities. Quant. Imaging Med. Surg. 2025, 15, 5825–5858. [Google Scholar] [CrossRef] [Scilit]
- Ayllon, E.M.; Mantegna, M.; Shen, L.; Soda, P.; Guarrasi, V.; Tortora, M. Can foundation models really segment tumors? A benchmarking odyssey in lung CT imaging. In 2025 IEEE 38th International Symposium on Computer-Based Medical Systems (CBMS); IEEE: New York, NY, USA, 2025. [Google Scholar]
- Lee, H.H.; Gu, Y.; Zhao, T.; Xu, Y.; Yang, J.; Usuyama, N.; Wong, C.; Wei, M.; Landman, B.A.; Huo, Y.; et al. Foundation Models for Biomedical Image Segmentation: A Survey. arXiv 2024, arXiv:2401.07654. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Goel, K.; Ré, C. Efficiently modeling long sequences with structured state spaces. Presented at the International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. arXiv 2024, arXiv:2401.10166. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef] [Scilit]
- Sun, Y.; Wang, J.; Yin, R. From channel-spatial attention to state space models: A review of evolving mechanisms in tumour segmentation. Clin. Transl. Discov. 2026, 6, e70127. [Google Scholar] [CrossRef] [Scilit]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
- Qayyum, A.; Mazher, M.; Niederer, S.A. Assessing self-supervised xLSTM-UNet architectures for head and neck tumor segmentation. In Challenge on Head and Neck Tumor Segmentation for MRI-Guided Applications; Springer: Cham, Switzerland, 2024; pp. 166–178. [Google Scholar]
- Novikov, A.A.; Major, D.; Wimmer, M.; Lenis, D.; Bühler, K. Deep sequential segmentation of organs in volumetric medical scans. IEEE Trans. Med. Imaging 2018, 38, 1207–1215. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ter-Sarkisov, A. One shot model for COVID-19 classification and lesions segmentation in chest CT scans using long short-term memory network with attention mechanism. IEEE Intell. Syst. 2022, 37, 54–64. [Google Scholar] [CrossRef] [Scilit]
- Guo, M.-H.; Xu, T.-X.; Liu, J.-J.; Liu, Z.-N.; Jiang, P.-T.; Mu, T.-J.; Zhang, S.-H.; Martin, R.R.; Cheng, M.-M.; Hu, S.-M. Attention mechanisms in computer vision: A survey. Comput. Vis. Media 2022, 8, 331–368. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Zhang, C.; Zhang, L. Vision Mamba: A comprehensive survey and taxonomy. arXiv 2024, arXiv:2405.04404. [Google Scholar] [CrossRef] [Scilit]
- Lambert, Z.; Petitjean, C.; Dubray, B.; Kuan, S. SegTHOR: Segmentation of thoracic organs at risk in CT images. In Proceedings of the 2020 Tenth International Conference on Image Processing Theory, Tools and Applications (IPTA); IEEE: New York, NY, USA, 2020; pp. 1–6. [Google Scholar]
- Ji, Y.; Bai, H.; Ge, C.; Yang, J.; Zhu, Y.; Zhang, R.; Li, Z.; Zhanng, L.; Ma, W.; Wan, X.; et al. AMOS: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation. In Advances in Neural Information Processing Systems (NeurIPS); IEEE: New York, NY, USA, 2022; Volume 35, pp. 36722–36732. [Google Scholar]
- Bernard, O.; Lalande, A.; Zotti, C.; Cervenansky, F.; Yang, X.; Heng, P.A.; Cetin, I.; Lekadir, K.; Camara, O.; Ballester, M.A.G.; et al. Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: Is the problem solved? IEEE Trans. Med. Imaging 2018, 37, 2514–2525. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Antonelli, M.; Reinke, A.; Bakas, S.; Farahani, K.; Kopp-Schneider, A.; Landman, B.A.; Litjens, G.; Menze, B.; Ronneberger, O.; Summers, R.M.; et al. The medical segmentation decathlon. Nat. Commun. 2022, 13, 4128. [Google Scholar] [CrossRef] [Scilit]
- Heller, N.; Isensee, F.; Tejpaul, R.; Wood, A.; Papanikolopoulos, N.; Weight, C. The 2023 Kidney and Kidney Tumor Segmentation Challenge (KiTS23). In Kidney and Kidney Tumor Segmentation: MICCAI 2023 Challenge, KiTS 2023; Springer: Cham, Switzerland, 2024; pp. 1–10. [Google Scholar]








| Method | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|
| U-Net | 79.2 | 66.2 | 6.12 | 2.44 |
| VM-UNet | 86.1 | 75.9 | 4.85 | 1.91 |
| TransUNet | 80.7 | 67.6 | 5.78 | 2.18 |
| SwinUNet | 82.8 | 70.6 | 5.41 | 2.03 |
| MedSAM | 87.3 | 78.5 | 4.27 | 1.58 |
| nnU-Net | 88.9 | 80.1 | 4.02 | 1.47 |
| UNETR | 84.2 | 72.6 | 5.33 | 2.07 |
| Swin-UNETR | 89.6 | 81.2 | 3.94 | 1.41 |
| UX-Net | 90.1 | 82.1 | 3.88 | 1.39 |
| Ours |
| Method | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|
| U-Net | 74.2 | 61.2 | 7.92 | 3.71 |
| VM-UNet | 88.1 | 78.9 | 5.13 | 2.36 |
| TransUNet | 90.1 | 82.1 | 4.67 | 1.95 |
| SwinUNet | 89.2 | 80.6 | 4.98 | 2.01 |
| MedSAM | 92.5 | 87.2 | 4.52 | 2.82 |
| nnU-Net | 92.8 | 87.6 | 3.98 | 1.74 |
| UNETR | 87.4 | 77.7 | 5.91 | 2.65 |
| Swin-UNETR | 93.1 | 88.1 | 3.62 | 1.66 |
| UX-Net | 93.3 | 88.4 | 3.44 | 1.59 |
| Ours |
| Method | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|
| U-Net | 88.6 | 79.7 | 8.02 | 3.46 |
| VM-UNet | 90.3 | 82.5 | 7.44 | 3.11 |
| TransUNet | 90.9 | 83.4 | 7.11 | 3.02 |
| SwinUNet | 91.5 | 84.3 | 6.83 | 2.92 |
| MedSAM | 92.4 | 86.7 | 8.37 | 3.52 |
| nnU-Net | 93.1 | 87.1 | 7.44 | 3.03 |
| UNETR | 90.4 | 82.6 | 8.33 | 3.61 |
| Swin-UNETR | 93.6 | 88.0 | 6.91 | 2.84 |
| UX-Net | 94.0 | 88.7 | 6.52 | 2.76 |
| Ours |
| Method | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|
| U-Net | 84.8 | 73.5 | 6.02 | 2.41 |
| VM-UNet | 88.2 | 79.0 | 4.93 | 1.97 |
| TransUNet | 89.3 | 80.7 | 4.67 | 1.88 |
| SwinUNet | 90.9 | 83.5 | 4.12 | 1.69 |
| MedSAM | 86.9 | 77.8 | 5.50 | 1.86 |
| nnU-Net | 91.2 | 84.0 | 4.21 | 1.71 |
| UNETR | 88.4 | 79.4 | 5.83 | 2.03 |
| Swin-UNETR | 92.4 | 86.5 | 3.96 | 1.63 |
| UX-Net | 92.8 | 87.2 | 3.88 | 1.57 |
| Ours |
| Method | DSC (%) | IoU (%) | HD95 (mm) | ASD (mm) |
|---|---|---|---|---|
| U-Net | 89.6 | 81.2 | 2.93 | 1.22 |
| VM-UNet | 88.5 | 79.5 | 4.33 | 1.49 |
| TransUNet | 91.2 | 83.9 | 2.21 | 0.91 |
| SwinUNet | 93.0 | 87.5 | 1.39 | 0.66 |
| MedSAM | 90.4 | 82.5 | 3.92 | 1.28 |
| nnU-Net | 92.8 | 86.9 | 1.52 | 0.71 |
| UNETR | 90.1 | 82.8 | 2.31 | 0.96 |
| Swin-UNETR | 93.9 | 88.4 | 1.36 | 0.63 |
| UX-Net | 93.3 | 87.7 | 1.41 | 0.69 |
| Ours |
| Method | FLOPs (G) | Params (M) | Inference Time (ms) |
|---|---|---|---|
| U-Net | 28.54 | 31.04 | 18.6 |
| TransUNet | 124.67 | 105.32 | 52.3 |
| SwinUNet | 96.82 | 89.45 | 47.1 |
| MedSAM (ViT-B) | 369.44 | 89.67 | 136.8 |
| VM-UNet | 22.41 | 98.63 | 34.2 |
| Swin-UNETR | 218.53 | 112.36 | 78.5 |
| Ours | 17.36 | 110.87 | 28.6 |
| Model | SS2D | xLSTM | Cross-Attention | DSC (%) |
|---|---|---|---|---|
| Baseline | 79.2 | |||
| #1 | ✓ | 86.8 | ||
| #2 | ✓ | ✓ | 88.9 | |
| VMMedSAM-X | ✓ | ✓ | ✓ | 90.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, H.; Li, W.; Liu, Y. VMMedSAM-X: A State-Enhanced Dual-Branch Encoder for Efficient Promptable Medical Image Segmentation. Appl. Sci. 2026, 16, 4199. https://doi.org/10.3390/app16094199
Zhang H, Li W, Liu Y. VMMedSAM-X: A State-Enhanced Dual-Branch Encoder for Efficient Promptable Medical Image Segmentation. Applied Sciences. 2026; 16(9):4199. https://doi.org/10.3390/app16094199
Chicago/Turabian StyleZhang, Hengwei, Wei Li, and Yazhi Liu. 2026. "VMMedSAM-X: A State-Enhanced Dual-Branch Encoder for Efficient Promptable Medical Image Segmentation" Applied Sciences 16, no. 9: 4199. https://doi.org/10.3390/app16094199
APA StyleZhang, H., Li, W., & Liu, Y. (2026). VMMedSAM-X: A State-Enhanced Dual-Branch Encoder for Efficient Promptable Medical Image Segmentation. Applied Sciences, 16(9), 4199. https://doi.org/10.3390/app16094199

