Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement
Abstract
1. Introduction
- We propose an anatomical landmark detection framework that uses a single labeled template image without requiring additional unlabeled training images.
- We design MLMF to adaptively fuse multi-layer and multi-facet DINO features, allowing the model to exploit complementary key and value projections from different ViT depths.
- We introduce MLCA to inject long-range anatomical context into fused patch descriptors using efficient Mamba-based bidirectional sequence modeling.
- We incorporate TCGR to refine independently matched landmarks through topology-constrained graph reasoning.
- We conduct extensive experiments, ablation studies, foundation-model comparisons, statistical testing, graph-design sensitivity analysis, and qualitative visualization to validate the effectiveness and limitations of the proposed framework.
2. Related Work
2.1. Anatomical Landmark Detection
2.2. Vision Foundation Models
2.3. State Space Models and Mamba in Visual Representation Learning
2.4. Graph-Based Structural Reasoning for Landmarks
3. Methodology
3.1. Problem Formulation
3.2. Multi-Layer Multi-Facet Adaptive Fusion (MLMF)
3.3. Mamba-Based Long-Range Context Aggregation (MLCA)
3.4. Two-Stage Global-to-Local Coarse-to-Fine Matching
3.5. Topology-Constrained Graph Refinement (TCGR)
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Comparison with State of the Art
4.4. Ablation Study
4.5. Qualitative Comparison
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wang, C.W.; Huang, C.T.; Lee, J.H.; Li, C.H.; Chang, S.W.; Siao, M.J.; Lai, T.M.; Ibragimov, B.; Vrtovec, T.; Ronneberger, O.; et al. A benchmark for comparison of dental radiography analysis algorithms. Med. Image Anal. 2016, 31, 63–76. [Google Scholar] [CrossRef] [PubMed]
- Payer, C.; Štern, D.; Bischof, H.; Urschler, M. Integrating spatial configuration into heatmap regression based CNNs for landmark localization. Med. Image Anal. 2019, 54, 207–219. [Google Scholar] [CrossRef] [PubMed]
- Urschler, M.; Ebner, T.; Štern, D. Integrating geometric configuration and appearance information into a unified framework for anatomical landmark localization. Med. Image Anal. 2018, 43, 23–36. [Google Scholar] [CrossRef] [PubMed]
- Yao, Q.; Quan, Q.; Xiao, L.; Zhou, S.K. One-shot medical landmark detection. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer International Publishing: Cham, Switzerland, 2021; pp. 177–188. [Google Scholar]
- Chen, R.; Ma, Y.; Chen, N.; Lee, D.; Wang, W. Cephalometric landmark detection by attentive feature pyramid fusion and regression-voting. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer International Publishing: Cham, Switzerland, 2019; pp. 873–881. [Google Scholar]
- Zhang, J.; Liu, M.; Shen, D. Detecting anatomical landmarks from limited medical imaging data using two-stage task-oriented deep neural networks. IEEE Trans. Image Process. 2017, 26, 4753–4764. [Google Scholar] [CrossRef] [PubMed]
- Lee, J.H.; Yu, H.J.; Kim, M.J.; Kim, J.W.; Choi, J. Automated cephalometric landmark detection with confidence regions using Bayesian convolutional neural networks. BMC Oral Health 2020, 20, 270. [Google Scholar] [CrossRef] [PubMed]
- Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 9650–9660. [Google Scholar]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.V.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning robust visual features without supervision. arXiv 2023, arXiv:2304.07193. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Amir, S.; Gandelsman, Y.; Bagon, S.; Dekel, T. Deep ViT features as dense visual descriptors. In Proceedings of the ECCV Workshops, Tel Aviv, Israel, 23–27 October 2022. [Google Scholar]
- Miao, J.; Chen, C.; Zhang, K.; Chuai, J.; Li, Q.; Heng, P.A. FM-OSD: Foundation Model-Enabled One-Shot Detection of Anatomical Landmarks. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2024; Springer Nature: Cham, Switzerland, 2024; Volume 15011, pp. 297–307. [Google Scholar] [CrossRef]
- Hamilton, M.; Zhang, Z.; Hariharan, B.; Snavely, N.; Freeman, W.T. Unsupervised semantic segmentation by distilling feature correspondences. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Cootes, T.F.; Edwards, G.J.; Taylor, C.J. Active appearance models. IEEE Trans. Pattern Anal. Mach. Intell. 2001, 23, 681–685. [Google Scholar] [CrossRef]
- Bertinetto, L.; Valmadre, J.; Henriques, J.F.; Vedaldi, A.; Torr, P.H.S. Fully-convolutional siamese networks for object tracking. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Amsterdam, The Netherlands, 8–10, 15–16 October 2016; pp. 850–865. [Google Scholar]
- Sung, F.; Yang, Y.; Zhang, L.; Xiang, T.; Torr, P.H.S.; Hospedales, T.M. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 1199–1208. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 4015–4026. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 16000–16009. [Google Scholar]
- Truong, P.; Danelljan, M.; Timofte, R.; Van Gool, L. PDC-Net+: Enhanced probabilistic dense correspondence network. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10247–10266. [Google Scholar] [CrossRef] [PubMed]
- Xu, Y.; Zhang, J.; Zhang, Q.; Tao, D. ViTPose: Simple vision transformer baselines for human pose estimation. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 38571–38584. [Google Scholar]
- Hayat, M.; Izhar, R.; Nadeem, M.; Anjum, H.; Muhammad, A.; Bhattacharjee, S. HAMSRNet-Hybrid Attention Multiscale Super-Resolution Network for Endoscopic Images. In Proceedings of the 2025 5th International Conference on Digital Futures and Transformative Technologies (ICoDT2); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Wang, X.; Huo, Y.; Liu, Y.; Guo, X.; Yan, F.; Zhao, G. Multimodal feature-guided audio-driven emotional talking face generation. Electronics 2025, 14, 2684. [Google Scholar] [CrossRef]
- Dong, G.; Schultz, L.; Hassanpour, N.; Gao, C. RePack: Representation Packing of Vision Foundation Model Features Enhances Diffusion Transformer. arXiv 2025, arXiv:2512.12083. [Google Scholar]
- Tian, Y.; Wang, Z.; Xu, S.; Guo, L. KAN-OSD: Dino-Based Encoder and KAN-Based Decoders with Dual Contrastive Learning for One-Shot Anatomical Landmark Detection. In Proceedings of the 2026 IEEE 23rd International Symposium on Biomedical Imaging (ISBI); IEEE: Piscataway, NJ, USA, 2026; pp. 1–4. [Google Scholar]
- Wang, A.; Elbatel, M.; Liu, K.; Lin, L.; Lan, M.; Yang, Y.; Li, X. Geometric-Guided Few-Shot Dental Landmark Detection with Human-Centric Foundation Model. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer Nature: Cham, Switzerland, 2025; pp. 197–207. [Google Scholar]
- Gu, A.; Goel, K.; Ré, C. Efficiently modeling long sequences with structured state spaces. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. In Proceedings of the Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; Volume 37, pp. 103031–103063. [Google Scholar] [CrossRef]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient visual representation learning with bidirectional state space model. In Proceedings of the International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024; pp. 62429–62442. [Google Scholar]
- Xing, Z.; Ye, T.; Yang, Y.; Liu, G.; Zhu, L. SegMamba: Long-range sequential modeling Mamba for 3D medical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; pp. 578–588. [Google Scholar]
- Ruan, J.; Li, J.; Xiang, S. VM-UNet: Vision Mamba UNet for medical image segmentation. arXiv 2024, arXiv:2402.02491. [Google Scholar] [CrossRef]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph attention networks. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Zhao, L.; Peng, X.; Tian, Y.; Kapadia, M.; Metaxas, D.N. Semantic graph convolutional networks for 3D human pose regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 3420–3430. [Google Scholar]
- Cai, Y.; Ge, L.; Liu, J.; Cai, J.; Cham, T.J.; Yuan, J.; Magnenat Thalmann, N. Exploiting spatial-temporal relationships for 3D pose estimation via graph convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 2272–2281. [Google Scholar]
- Lu, G.; Zhang, Y.; Kong, Y.; Zhang, C.; Coatrieux, J.L.; Shu, H. Landmark localization for cephalometric analysis using multiscale image patch-based graph convolutional networks. IEEE J. Biomed. Health Inform. 2022, 26, 3015–3024. [Google Scholar] [CrossRef]
- Chen, J.; Yu, B.; Lei, B.; Feng, R.; Chen, D.Z.; Wu, J. Doctor imitator: A graph-based bone age assessment framework using hand radiographs. In Proceedings of the International Conference on Medical Image Computing and Computer Assisted Intervention, Virtual, 4–8 October 2020; pp. 764–774. [Google Scholar]
- Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar] [CrossRef]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Yan, K.; Cai, J.; Jin, D.; Miao, S.; Guo, D.; Harrison, A.P.; Tang, Y.; Xiao, J.; Lu, J.; Lu, L. SAM: Self-supervised learning of pixel-wise anatomical embeddings in radiological images. IEEE Trans. Med. Imaging 2022, 41, 2658–2669. [Google Scholar] [CrossRef]
- Yin, Z.; Gong, P.; Wang, C.; Yu, Y.; Wang, Y. One-shot medical landmark localization by edge-guided transform and noisy landmark refinement. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 473–489. [Google Scholar]
- Zhu, H.; Quan, Q.; Yao, Q.; Liu, Z.; Zhou, S.K. UOD: Universal one-shot detection of anatomical landmarks. In Proceedings of the International Conference on Medical Image Computing and Computer Assisted Intervention, Vancouver, BC, Canada, 8–12 October 2023; pp. 24–34. [Google Scholar]
- Franceschi, L.; Niepert, M.; Pontil, M.; He, X. Learning discrete structures for graph neural networks. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; pp. 1972–1982. [Google Scholar]



| Method | Cephalometric (Head) Dataset | Hand X-Ray Dataset | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| L/U | MRE ↓ | SDR@2 mm ↑ | SDR@2.5 mm ↑ | SDR@3 mm ↑ | SDR@4 mm↑ | L/U | MRE ↓ | SDR@2 mm ↑ | SDR@4 mm ↑ | |
| CC2D | 1/149 | 2.36 | 51.81 | 63.13 | 73.66 | 86.25 | 1/608 | 2.65 | 51.19 | 82.56 |
| SAEM | 1/149 | 2.58 | 54.34 | 64.51 | 70.82 | 80.76 | 1/608 | 1.69 | 76.61 | 92.52 |
| EGTNLR | 1/149 | 2.27 | 49.45 | 63.07 | 74.70 | 88.91 | 1/608 | 1.81 | 64.62 | 95.03 |
| UOD | 3/985 | 2.43 | 51.14 | 62.37 | 74.40 | 86.49 | 3/985 | 2.52 | 53.37 | 84.27 |
| FM-OSD | 1/0 | 1.82 | 67.35 | 77.92 | 84.59 | 91.92 | 1/0 | 1.41 | 86.66 | 96.66 |
| Ours | 1/0 | 1.72 | 71.31 | 80.14 | 86.42 | 94.08 | 1/0 | 1.65 | 80.81 | 95.93 |
| # | Configuration | MRE ↓ | ΔMRE | SDR@2 mm ↑ | SDR@3 mm ↑ | SDR@4 mm ↑ |
|---|---|---|---|---|---|---|
| Component-wise addition (each row adds one module to the previous) | ||||||
| 1 | Baseline: Layer-9 key only, no MLCA, no TCGR | 1.82 | +0.10 | 67.40 | 82.31 | 91.89 |
| 2 | +MLMF (Layers 5/8/11, key + value, adaptive fusion) | 1.78 | +0.06 | 68.93 | 83.46 | 92.58 |
| 3 | +MLCA (2-direction scan, , residual gate) | 1.76 | +0.04 | 69.53 | 84.17 | 93.09 |
| 4 | +TCGR (3 GAT layers + topology loss)—Full Method | 1.72 | — | 71.31 | 86.42 | 94.08 |
| MLMF design variants (MLCA + TCGR fixed at full-method settings) | ||||||
| 5 | Uniform fusion (equal weights, no learned ) | 1.73 | +0.01 | 70.87 | 85.94 | 93.71 |
| 6 | Key facet only (value projections removed) | 1.75 | +0.03 | 70.34 | 85.29 | 93.34 |
| 7 | Single Layer {8} only (multi-layer disabled) | 1.77 | +0.05 | 69.91 | 84.73 | 92.87 |
| MLCA design variants (MLMF + TCGR fixed at full-method settings) | ||||||
| 8 | No MLCA (MLMF output used directly) | 1.76 | +0.04 | 69.18 | 83.84 | 92.76 |
| 9 | 1-direction scan (left-to-right only) | 1.74 | +0.02 | 70.48 | 85.63 | 93.51 |
| 10 | No residual gate ( fixed to 1) | 1.74 | +0.02 | 70.21 | 85.40 | 93.31 |
| TCGR design variants (MLMF + MLCA fixed at full-method settings) | ||||||
| 11 | No TCGR (matching output used directly) | 1.76 | +0.04 | 69.67 | 84.39 | 93.22 |
| 12 | 1 GAT layer | 1.74 | +0.02 | 70.19 | 85.33 | 93.29 |
| 13 | 2 GAT layers | 1.73 | +0.01 | 70.84 | 85.89 | 93.67 |
| 14 | No topology loss () | 1.75 | +0.03 | 69.95 | 85.15 | 93.12 |
| 15 | GCN mean aggregation (no attention) | 1.73 | +0.01 | 70.96 | 86.08 | 93.64 |
| Backbone | Cephalometric (Head) Dataset | Hand X-Ray Dataset | ||||||
|---|---|---|---|---|---|---|---|---|
| MRE ↓ | SDR@2 mm ↑ | SDR@2.5 mm ↑ | SDR@3 mm ↑ | SDR@4 mm ↑ | MRE ↓ | SDR@2 mm ↑ | SDR@4 mm ↑ | |
| SAM ViT-B | 1.83 | 67.52 | 77.86 | 84.13 | 91.94 | 1.93 | 76.34 | 92.17 |
| DINO ViT-S/8 (ours) | 1.72 | 71.31 | 80.14 | 86.42 | 94.08 | 1.65 | 80.81 | 95.93 |
| DINO ViT-B/8 | 1.69 | 72.18 | 81.07 | 87.33 | 94.61 | 1.61 | 80.27 | 96.34 |
| Graph Design | MRE ↓ | SDR@2 mm ↑ | SDR@3 mm ↑ | SDR@4 mm ↑ |
|---|---|---|---|---|
| No TCGR | 1.76 | 69.67 | 84.39 | 93.22 |
| Sparse random graph with matched edge number | 1.75 | 70.02 | 84.91 | 93.37 |
| k-nearest-neighbor graph () | 1.74 | 70.42 | 85.57 | 93.58 |
| Fully connected graph | 1.73 | 70.88 | 85.94 | 93.71 |
| Default anatomy-guided graph | 1.72 | 71.31 | 86.42 | 94.08 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tian, Y.; Wang, Z.; Guo, L. Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement. Electronics 2026, 15, 2414. https://doi.org/10.3390/electronics15112414
Tian Y, Wang Z, Guo L. Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement. Electronics. 2026; 15(11):2414. https://doi.org/10.3390/electronics15112414
Chicago/Turabian StyleTian, Yinbing, Ziyang Wang, and Li Guo. 2026. "Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement" Electronics 15, no. 11: 2414. https://doi.org/10.3390/electronics15112414
APA StyleTian, Y., Wang, Z., & Guo, L. (2026). Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement. Electronics, 15(11), 2414. https://doi.org/10.3390/electronics15112414

