Self-Prompting Segment Anything Model for Esophageal OCT Images
Abstract
1. Introduction
- We developed a one-shot segmentation framework specifically optimized for esophageal OCT images.
- We designed a self-prompting strategy to effectively adapt SAM for esophageal OCT image segmentation tasks.
- We validated the effectiveness of the proposed method on both a self-collected dataset and a public OCT dataset.
2. Related Works
2.1. Esophageal OCT Image Segmentation Using Deep Learning
2.2. SAM for Medical Image Segmentation
3. Methods
3.1. Overview of the Self-Prompting SAM for Esophageal OCT Images
3.2. The Self-Pretrained Prompt Encoder
3.3. The Segmentation Decoder
4. Experiments
4.1. Dataset
4.2. Implementation Details
4.3. Self-Pretraining of the Prompt Encoder
4.4. Segmentation Results and Comparisons
4.5. Ablation Study
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Huang, D.; Swanson, E.A.; Lin, C.P.; Schuman, J.S.; Stinson, W.G.; Chang, W.; Hee, M.R.; Flotte, T.; Gregory, K.; Puliafito, C.A.; et al. Optical Coherence Tomography. Science 1991, 254, 1178–1181. [Google Scholar] [CrossRef] [PubMed]
- Jiang, H.; Zhang, B.; Xiang, B.; Liu, J.; Ma, Z.; Lv, H.; Yu, Y.; Zhao, Y.; Yang, Y.; Luan, J.; et al. Diopter Measurement of Human Eye Based on Dual-Focus Swept Source Optical Coherence Tomography. Photonics 2025, 12, 856. [Google Scholar] [CrossRef]
- Palmer, L.D.; Thompson, A.C.; Asrani, S. Diagnosing glaucoma progression with optical coherence tomography. Curr. Opin. Ophthalmol. 2025, 36, 130–134. [Google Scholar] [PubMed]
- Mehrjoo, M.; Khamar, P.; Darzi, S.; Verma, S.; Shetty, R.; Arba Mosquera, S. Automated Characterization of Intrastromal Corneal Cuts Induced by Two Femtosecond Laser Systems Using OCT Imaging. Photonics 2024, 11, 1123. [Google Scholar] [CrossRef]
- Wax, A.; Zhang, H.; Jelly, E.T.; Price, H.B.; Sun, T.; Chu, K.K.; Cotton, C.C.; Eluri, S.; Goldblum, J.R.; Shaheen, N.J. Prospective Identification of Dysplasia in Barrett’s Esophagus with Combined Optical Coherence Tomography and Light Scattering Measurements. J. Biophotonics 2026, 19, e202500380. [Google Scholar] [PubMed]
- Sterkenburg, A.; Dijkhuis, T.; Danskin, T.; Vahrmijer, A.; De Boer, J.; Nagengast, W. Assessing the safety and feasibility of optical coherence tomography and near-infrared fluorescence capsule endoscopy in Barrett’s Esophagus patients. Endoscopy 2025, 57, eP299. [Google Scholar] [CrossRef]
- Zhang, J.L.; Yuan, W.; Liang, W.X.; Yu, S.Y.; Liang, Y.M.; Xu, Z.Y.; Wei, Y.X.; Li, X.D. Automatic and robust segmentation of endoscopic OCT images and optical staining. Biomed. Opt. Express 2017, 8, 2697–2708. [Google Scholar] [CrossRef] [PubMed]
- Qi, X.; Pan, Y.S.; Sivak, M.V.; Willis, J.E.; Isenberg, G.; Rollins, A.M. Image analysis for classification of dysplasia in Barrett’s esophagus using endoscopic optical coherence tomography. Biomed. Opt. Express 2010, 1, 825–847. [Google Scholar] [CrossRef] [PubMed]
- Yang, Z.Y.; Soltanian-Zadeh, S.; Chu, K.K.; Zhang, H.R.; Moussa, L.; Watts, A.E.; Shaheen, N.J.; Wax, A.; Farsiu, S. Connectivity-based deep learning approach for segmentation of the epithelium in in vivo human esophageal OCT images. Biomed. Opt. Express 2021, 12, 6326–6340. [Google Scholar] [CrossRef] [PubMed]
- Wang, C.; Gan, M. Wavelet attention network for the segmentation of layer structures on OCT images. Biomed. Opt. Express 2022, 13, 6167–6181. [Google Scholar] [CrossRef] [PubMed]
- Jeihouni, P.; Dehzangi, O.; Amireskandari, A.; Rezai, A.; Nasrabadi, N.M. MultiSDGAN: Translation of OCT Images to Superresolved Segmentation Labels Using Multi-Discriminators in Multi-Stages. IEEE J. Biomed. Health Inform. 2022, 26, 1614–1627. [Google Scholar] [CrossRef] [PubMed]
- Moradi, M.; Du, X.; Huan, T.; Chen, Y. Feasibility of the soft attention-based models for automatic segmentation of OCT kidney images. Biomed. Opt. Express 2022, 13, 2728–2738. [Google Scholar] [CrossRef] [PubMed]
- Moradi, M.; Chen, Y.; Du, X.; Seddon, J.M. Deep ensemble learning for automated non-advanced AMD classification using optimized retinal layer segmentation and SD-OCT scans. Comput. Biol. Med. 2023, 154, 106512. [Google Scholar] [CrossRef] [PubMed]
- Gan, M.; Wang, C.; Yang, T.; Yang, N.; Zhang, M.; Yuan, W.; Li, X.D.; Wang, L.R. Robust layer segmentation of esophageal OCT images based on graph search using edge-enhanced weights. Biomed. Opt. Express 2018, 9, 4481–4495. [Google Scholar] [CrossRef] [PubMed]
- Li, D.W.; Wu, J.M.; He, Y.F.; Yao, X.W.; Yuan, W.; Chen, D.F.; Park, H.C.; Yu, S.Y.; Prince, J.L.; Li, X.D. Parallel deep neural networks for endoscopic OCT image segmentation. Biomed. Opt. Express 2019, 10, 1126–1135. [Google Scholar] [CrossRef] [PubMed]
- Wang, C.; Gan, M.; Zhang, M.; Li, D.Y. Adversarial convolutional network for esophageal tissue segmentation on OCT images. Biomed. Opt. Express 2020, 11, 3095–3110. [Google Scholar] [CrossRef] [PubMed]
- Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef] [PubMed]
- Awais, M.; Naseer, M.; Khan, S.; Anwer, R.M.; Cholakkal, H.; Shah, M.; Yang, M.H.; Khan, F.S. Foundation models defining a new era in vision: A survey and outlook. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 2245–2264. [Google Scholar] [CrossRef] [PubMed]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 3992–4003. [Google Scholar]
- Hu, X.; Xu, X.; Shi, Y. How to efficiently adapt large segmentation model (sam) to medical images. arXiv 2023, arXiv:2306.13731. [Google Scholar]
- Ma, J.; He, Y.; Li, F.; Han, L.; You, C.; Wang, B. Segment anything in medical images. Nat. Commun. 2024, 15, 654. [Google Scholar] [CrossRef] [PubMed]
- Wang, H.; Guo, S.; Ye, J.; Deng, Z.; Cheng, J.; Li, T.; Chen, J.; Su, Y.; Huang, Z.; Shen, Y.; et al. SAM-Med3D: A vision foundation model for general-purpose segmentation on volumetric medical images. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 1024. [Google Scholar] [CrossRef]
- Ravi, N.; Gabeur, V.; Hu, Y.T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L.; et al. SAM 2: Segment Anything in Images and Videos. arXiv 2024, arXiv:2408.00714. [Google Scholar]
- Li, D.; Cheng, Y.; Guo, Y.; Wang, L. Esophageal tissue segmentation on OCT images with hybrid attention network. Multimed. Tools Appl. 2024, 83, 42609–42628. [Google Scholar]
- Wang, C.; Gan, M. Few-shot segmentation for esophageal OCT images based on self-supervised vision transformer. Int. J. Imaging Syst. Technol. 2024, 34, e23006. [Google Scholar]
- Ali, M.; Wu, T.; Hu, H.; Luo, Q.; Xu, D.; Zheng, W.; Jin, N.; Yang, C.; Yao, J. A review of the segment anything model (sam) for medical image analysis: Accomplishments and perspectives. Comput. Med Imaging Graph. 2025, 119, 102473. [Google Scholar] [CrossRef] [PubMed]
- Wu, Q.; Zhang, Y.; Elbatel, M. Self-prompting Large Vision Models for Few-Shot Medical Image Segmentation. In Proceedings of the Domain Adaptation and Representation Transfer; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2024; Volume 14293, pp. 156–167. [Google Scholar] [CrossRef]
- Miao, J.; Chen, C.; Zhang, K.; Chuai, J.; Li, Q.; Heng, P.A. Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2024; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2024; Volume 15011, pp. 167–177. [Google Scholar] [CrossRef]
- Chen, T.; Zhu, L.; Ding, C.; Cao, R.; Wang, Y.; Zhang, S.; Li, Z.; Sun, L.; Zang, Y.; Mao, P. SAM-Adapter: Adapting Segment Anything in Underperformed Scenes. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France, 2–6 October 2023; pp. 3359–3367. [Google Scholar] [CrossRef]
- Gao, Y.; Xia, W.; Hu, D.; Zhou, S.K. Segment anything in medical images with adapter. Sci. Rep. 2024, 14, 12461. [Google Scholar] [CrossRef]
- Wu, J.; Fu, R.; Fang, H.; Liu, Y.; Wang, Z.; Xu, Y.; Jin, Y.; Arbel, T. Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation. arXiv 2023, arXiv:2304.12620. [Google Scholar]
- Fazekas, B.; Morano, J.; Lachinov, D.; Aresta, G.; Bogunovic, H. SAMedOCT: Adapting Segment Anything Model (SAM) for Retinal OCT. arXiv 2023, arXiv:2308.09331. [Google Scholar]
- Gu, H.; Dong, H.; Yang, J.; Mazurowski, M.A. How to Build the Best Medical Image Segmentation Algorithm Using Foundation Models: A Comprehensive Empirical Study with Segment Anything Model. Mach. Learn. Biomed. Imaging 2025, 3, 88–120. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked Autoencoders Are Scalable Vision Learners. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 15979–15988. [Google Scholar]
- Bao, H.; Dong, L.; Piao, S.; Wei, F. BEiT: BERT Pre-Training of Image Transformers. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Mishra, P.; Sarawadekar, K. Polynomial learning rate policy with warm restart for deep neural network. In Proceedings of the TENCON 2019—2019 IEEE Region 10 Conference (TENCON), Kochi, India, 17–20 October 2019; pp. 2087–2092. [Google Scholar]
- Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; Elsen, E.; Garcia, D.; Ginsburg, B.; Houston, M.; Kuchaiev, O.; Venkatesh, G.; et al. Mixed Precision Training. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015, Proceedings of the 18th International Conference 2015, Munich, Germany, 5–9 October 2015; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; Volume 9351, pp. 234–241. [Google Scholar]
- Feng, S.; Zhao, H.; Shi, F.; Cheng, X.; Wang, M.; Ma, Y.; Xiang, D.; Zhu, W.; Chen, X. CPFNet: Context pyramid fusion network for medical image segmentation. IEEE Trans. Med Imaging 2020, 39, 3008–3018. [Google Scholar] [CrossRef] [PubMed]
- Hatamizadeh, A.; Tang, Y.; Nath, V.; Yang, D.; Myronenko, A.; Landman, B.; Roth, H.R.; Xu, D. UNETR: Transformers for 3D Medical Image Segmentation. In Proceedings of the 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 4–8 January 2022; pp. 1748–1758. [Google Scholar]
- He, K.M.; Zhang, X.Y.; Ren, S.Q.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Müller, D.; Soto-Rey, I.; Kramer, F. Towards a guideline for evaluation metrics in medical image segmentation. BMC Res. Notes 2022, 15, 210. [Google Scholar] [CrossRef] [PubMed]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2999–3007. [Google Scholar]
- Abraham, N.; Khan, N.M. A novel focal Tversky loss function with improved attention U-Net for lesion segmentation. In Proceedings of the 2019 IEEE 16th International Symposium on Biomedical Imaging (ISBI 2019), Venice, Italy, 8–11 April 2019; pp. 683–687. [Google Scholar]





| Characteristic | Dataset #1 | Dataset #2 |
|---|---|---|
| Data source | Self-collected mouse esophageal OCT data | Public human esophageal OCT data [9] |
| Cohort size | 5 mice; 100 B-scans | 30 subjects; 784 B-scans |
| Clinical composition | All healthy | 23 healthy subjects and 7 subjects with Barrett’s esophagus |
| Segmentation target | Four tissue classes | Epithelium |
| Image dimensions (pixels) | Width: 506–1020; height: 310–648 | |
| Cross-validation | 5-fold subject-level cross-validation | 10-fold subject-level cross-validation |
| Fold composition | 4 subjects for training and 1 subject for testing | 27 subjects for training and 3 subjects for testing |
| One-shot protocol | One randomly selected labeled B-scan in each training fold | One randomly selected labeled B-scan in each training fold |
| Method | Precision (%) | Recall (%) | DSC (%) | Accuracy (%) |
|---|---|---|---|---|
| UNet [40] | ||||
| CPFNet [41] | ||||
| UNETR [42] | ||||
| AutoSAM [20] | ||||
| Self-prompting SAM (Ours) | ||||
| Ours (w/o SAM) | ||||
| Ours (w/o self-prompting) |
| Method | Precision (%) | Recall (%) | DSC (%) | Accuracy (%) |
|---|---|---|---|---|
| UNet [40] | ||||
| CPFNet [41] | ||||
| UNETR [42] | ||||
| AutoSAM [20] | ||||
| Self-prompting SAM (Ours) | ||||
| Ours (w/o SAM) | ||||
| Ours (w/o self-prompting) |
| Method | Parameters (M) | Inference Time (ms/Image) | Peak GPU Memory (GB) |
|---|---|---|---|
| UNet (ResNet-50) | 32.4 | 41 | 2.4 |
| CPFNet | 44.3 | 49 | 2.8 |
| UNETR (ViT-B) | 92.6 | 128 | 6.2 |
| AutoSAM (ViT-B) | 96.8 | 93 | 4.7 |
| Self-prompting SAM (Ours) | 187.6 | 190 | 8.5 |
| Analysis | Dataset | Setting/Model | DSC (%) |
|---|---|---|---|
| Sample sensitivity | Dataset #1 | Primary one-shot selection | |
| Dataset #1 | Reselected one-shot sample | ||
| Dataset #2 | Primary one-shot selection | ||
| Dataset #2 | Reselected one-shot sample | ||
| Model comparison | Dataset #2 | SAM with identical box prompts | |
| Dataset #2 | SAM2 with identical box prompts |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, C.; Gan, M. Self-Prompting Segment Anything Model for Esophageal OCT Images. Photonics 2026, 13, 664. https://doi.org/10.3390/photonics13070664
Wang C, Gan M. Self-Prompting Segment Anything Model for Esophageal OCT Images. Photonics. 2026; 13(7):664. https://doi.org/10.3390/photonics13070664
Chicago/Turabian StyleWang, Cong, and Meng Gan. 2026. "Self-Prompting Segment Anything Model for Esophageal OCT Images" Photonics 13, no. 7: 664. https://doi.org/10.3390/photonics13070664
APA StyleWang, C., & Gan, M. (2026). Self-Prompting Segment Anything Model for Esophageal OCT Images. Photonics, 13(7), 664. https://doi.org/10.3390/photonics13070664

