FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching
Abstract
1. Introduction
- We propose a novel student–teacher framework for keypoint detection, where SuperPoint provides structured supervision to a lightweight student network. To increase stability under homographic transformations, we introduce a consistency loss that can be framed as either a regression or classification objective.
- We design an efficient architecture based on MobileNetV3 enhanced with a Feature Pyramid Network (FPN) to improve multi-scale spatial representation while maintaining low computational cost.
- We introduce a two-stage training strategy that first learns strong feature representations, then refines them using label smoothing and Gaussian-filtered masks to improve spatial consistency.
- We demonstrate that our approach produces consistent and high-quality keypoint heatmaps and achieves competitive results on benchmark datasets while remaining efficient and scalable.
- The code is available at https://github.com/ionut-grigore99/FPC-Net, accessed on 20 April 2025.
2. Related Work
3. Method
4. Experiments
4.1. Keypoint Repeatability on HPatches
4.2. Homography Estimation Accuracy on HPatches
4.3. Pose Estimation Accuracy
4.4. Ablation Studies: FPN, Training Strategy, and Loss Design
4.5. Other Experiments
4.6. Applicability, Edge-Readiness, Scale, FOV, Failure Modes, and Mitigation
5. Implementation Details
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Sattler, T.; Maddern, W.; Toft, C.; Torii, A.; Hammarstrand, L.; Stenborg, E.; Safari, D.; Okutomi, M.; Pollefeys, M.; Sivic, J.; et al. Benchmarking 6DOF outdoor visual localization in changing conditions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 8601–8610. [Google Scholar]
- Schönberger, J.L.; Frahm, J.-M. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 4104–4113. [Google Scholar]
- Nistér, D.; Naroditsky, O.; Bergen, J. Visual odometry for ground vehicle applications. J. Field Robot. 2006, 23, 3–20. [Google Scholar] [CrossRef] [Scilit]
- Lowe, D.G. Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef] [Scilit]
- Bay, H.; Tuytelaars, T.; Van Gool, L. SURF: Speeded up robust features. In Proceedings of the European Conference on Computer Vision (ECCV), Graz, Austria, 7–13 May 2006; pp. 404–417. [Google Scholar]
- Rublee, E.; Rabaud, V.; Konolige, K.; Bradski, G. ORB: An efficient alternative to SIFT or SURF. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain, 6–13 November 2011; pp. 2564–2571. [Google Scholar]
- Yi, K.M.; Trulls, E.; Lepetit, V.; Fua, P. LIFT: Learned invariant feature transform. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 467–483. [Google Scholar]
- DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperPoint: Self-supervised interest point detection and description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshop (CVPRW), Salt Lake City, UT, USA, 18–22 June 2018; pp. 224–236. [Google Scholar]
- Revaud, J.; De Souza, C.; Humenberger, M.; Weinzaepfel, P. R2D2: Reliable and repeatable detector and descriptor. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Volume 32. [Google Scholar]
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for MobileNetV3. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Lindenberger, P.; Sarlin, P.-E.; Pollefeys, M. LightGlue: Local feature matching at light speed. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 17627–17638. [Google Scholar]
- Dusmanu, M.; Rocco, I.; Pajdla, T.; Pollefeys, M.; Sivic, J.; Torii, A.; Sattler, T. D2-Net: A trainable CNN for joint description and detection of local features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 16–20 June 2019; pp. 8092–8101. [Google Scholar]
- Taira, H.; Okutomi, M.; Sattler, T.; Cimpoi, M.; Pollefeys, M.; Sivic, J.; Pajdla, T.; Torii, A. InLoc: Indoor visual localization with dense matching and view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7199–7209. [Google Scholar]
- Zhou, H.; Sattler, T.; Jacobs, D.W. Evaluating local features for day-night matching. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Amsterdam, The Netherlands, 8–16 October 2016; pp. 724–736. [Google Scholar]
- Ono, Y.; Trulls, E.; Fua, P.; Yi, K.M. LF-Net: Learning local features from images. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montréal, QC, Canada, 2–8 December 2018; Volume 31. [Google Scholar]
- Sun, J.; Shen, Z.; Wang, Y.; Bao, H.; Zhou, X. LoFTR: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 19–25 June 2021; pp. 8922–8931. [Google Scholar]
- Farnebäck, G. Two-frame motion estimation based on polynomial expansion. In Proceedings of the Scandinavian Conference on Image Analysis (SCIA), Halmstad, Sweden, 29 June–2 July 2003; pp. 363–370. [Google Scholar]
- Horn, B.K.P.; Schunck, B.G. Determining optical flow. Artif. Intell. 1981, 17, 185–203. [Google Scholar] [CrossRef] [Scilit]
- Rocco, I.; Cimpoi, M.; Arandjelović, R.; Torii, A.; Pajdla, T.; Sivic, J. Neighbourhood consensus networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montréal, QC, Canada, 3–8 December 2018; Volume 31. [Google Scholar]
- Rocco, I.; Arandjelović, R.; Sivic, J. Efficient neighbourhood consensus networks via submanifold sparse convolutions. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 605–621. [Google Scholar]
- Li, X.; Han, K.; Li, S.; Prisacariu, V. Dual-resolution correspondence networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual Event, 6–12 December 2020; Volume 33, pp. 17346–17357. [Google Scholar]
- Chen, H.; Luo, Z.; Zhou, L.; Tian, Y.; Zhen, M.; Fang, T.; McKinnon, D.; Tsin, Y.; Quan, L. Aspanformer: Detector-free image matching with adaptive span transformer. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 20–36. [Google Scholar]
- Wang, Q.; Zhang, J.; Yang, K.; Peng, K.; Stiefelhagen, R. MatchFormer: Interleaving attention in transformers for feature matching. In Proceedings of the Asian Conference on Computer Vision (ACCV), Macau, China, 4–8 December 2022; pp. 2746–2762. [Google Scholar]
- Choy, C.B.; Gwak, J.; Savarese, S.; Chandraker, M. Universal correspondence network. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Barcelona, Spain, 5–10 December 2016; Volume 29. [Google Scholar]
- Fathy, M.E.; Tran, Q.-H.; Zia, M.Z.; Vernaza, P.; Chandraker, M. Hierarchical metric learning and matching for 2D and 3D geometric correspondences. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 803–819. [Google Scholar]
- Savinov, N.; Ladicky, L.; Pollefeys, M. Matching neural paths: Transfer from recognition to correspondence search. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Schönberger, J.L.; Pollefeys, M.; Geiger, A.; Sattler, T. Semantic visual localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6896–6906. [Google Scholar]
- Shi, J. Good features to track. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 21–23 June 1994; pp. 593–600. [Google Scholar]
- Harris, C.; Stephens, M. A combined corner and edge detector. In Proceedings of the Alvey Vision Conference, Manchester, UK, 31 August–2 September 1988; Volume 15, pp. 10–5244. [Google Scholar]
- Rosten, E.; Drummond, T. Machine learning for high-speed corner detection. In Proceedings of the European Conference on Computer Vision (ECCV), Graz, Austria, 7–13 May 2006; pp. 430–443. [Google Scholar]
- Dalal, N.; Triggs, B. Histograms of oriented gradients for human detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA, 20–25 June 2005; Volume 1, pp. 886–893. [Google Scholar]
- Calonder, M.; Lepetit, V.; Ozuysal, M.; Trzcinski, T.; Strecha, C.; Fua, P. BRIEF: Computing a local binary descriptor very fast. IEEE Trans. Pattern Anal. Mach. Intell. 2011, 34, 1281–1298. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Leutenegger, S.; Chli, M.; Siegwart, R.Y. BRISK: Binary robust invariant scalable keypoints. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain, 6–13 November 2011; pp. 2548–2555. [Google Scholar]
- Mikolajczyk, K.; Schmid, C. Scale & affine invariant interest point detectors. Int. J. Comput. Vis. 2004, 60, 63–86. [Google Scholar] [CrossRef] [Scilit]
- Lenc, K.; Vedaldi, A. Learning covariant feature detectors. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Amsterdam, The Netherlands, 8–16 October 2016; pp. 100–117. [Google Scholar]
- Zhang, L.; Rusinkiewicz, S. Learning to detect features in texture images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6325–6333. [Google Scholar]
- Savinov, N.; Seki, A.; Ladicky, L.; Sattler, T.; Pollefeys, M. Quad-networks: Unsupervised learning to rank for interest point detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1822–1830. [Google Scholar]
- Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperGlue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020; pp. 4938–4947. [Google Scholar]
- Lynen, S.; Sattler, T.; Bosse, M.; Hesch, J.A.; Pollefeys, M.; Siegwart, R. Get out of my lab: Large-scale, real-time visual-inertial localization. In Proceedings of the Robotics: Science and Systems (RSS), Rome, Italy, 13–17 July 2015; Volume 1. [Google Scholar]
- Tardioli, D.; Montijano, E.; Mosteo, A.R. Visual data association in narrow-bandwidth networks. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany, 28 September–2 October 2015; pp. 2572–2577. [Google Scholar]
- Wang, C.; Zhang, G.; Cheng, Z.; Zhou, W. Rethinking low-level features for interest point detection and description. In Proceedings of the Asian Conference on Computer Vision (ACCV), Macau, China, 4–8 December 2022; pp. 2059–2074. [Google Scholar]
- Guo, Y.; Xu, Y.; Niu, H.; Li, Z.; Jiao, X.; Li, S. Vision-based full-field panorama generation by UAV using GPS data and feature points filtering. Smart Struct. Syst. 2020, 25, 631–641. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Balntas, V.; Lenc, K.; Vedaldi, A.; Mikolajczyk, K. HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 5173–5182. [Google Scholar]
- Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision meets robotics: The KITTI dataset. Int. J. Robot. Res. 2013, 32, 1231–1237. [Google Scholar] [CrossRef] [Scilit]
- Burri, M.; Nikolic, J.; Gohl, P.; Schneider, T.; Rehder, J.; Omari, S.; Achtelik, M.W.; Siegwart, R. The EuRoC micro aerial vehicle datasets. Int. J. Robot. Res. 2016, 35, 1157–1163. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Zhou, X.; Hariharan, B.; Snavely, N. Learning feature descriptors using camera pose supervision. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 757–774. [Google Scholar]
- Gao, X.-S.; Hou, X.-R.; Tang, J.; Cheng, H.-F. Complete solution classification for the perspective-three-point problem. IEEE Trans. Pattern Anal. Mach. Intell. 2003, 25, 930–943. [Google Scholar]
- Fischler, M.A.; Bolles, R.C. Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 1981, 24, 381–395. [Google Scholar] [CrossRef] [Scilit]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.P.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L. An imperative style, high-performance deep learning library. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Volume 32, p. 8026. [Google Scholar]
- Buslaev, A.; Iglovikov, V.I.; Khvedchenya, E.; Parinov, A.; Druzhinin, M.; Kalinin, A.A. Albumentations: Fast and flexible image augmentations. Information 2020, 11, 125. [Google Scholar] [CrossRef] [Scilit]





| Method | All | i | v | ||||||
|---|---|---|---|---|---|---|---|---|---|
| = 1 | = 3 | = 8 | = 1 | = 3 | = 8 | = 1 | = 3 | = 8 | |
| FPC-Net (ours) | 0.46 | 0.59 | 0.67 | 0.43 | 0.51 | 0.55 | 0.48 | 0.67 | 0.79 |
| SuperPoint [8] | 0.31 | 0.53 | 0.65 | 0.33 | 0.51 | 0.62 | 0.29 | 0.55 | 0.68 |
| Shi [29] | 0.27 | 0.44 | 0.59 | 0.28 | 0.41 | 0.56 | 0.25 | 0.47 | 0.62 |
| Harris [30] | 0.45 | 0.59 | 0.68 | 0.40 | 0.48 | 0.57 | 0.50 | 0.70 | 0.79 |
| FAST [31] | 0.31 | 0.55 | 0.74 | 0.32 | 0.48 | 0.68 | 0.29 | 0.61 | 0.80 |
| SIFT [4] | 0.27 | 0.46 | 0.70 | 0.27 | 0.41 | 0.63 | 0.27 | 0.52 | 0.77 |
| Method | Time (ms) | Descriptor Size (MB/Pair) | All | i | v | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| = 1 | = 3 | = 8 | = 1 | = 3 | = 8 | = 1 | = 3 | = 8 | |||
| FPC-Net (ours) | 8 | 0 | 0.54 | 0.74 | 0.84 | 0.63 | 0.88 | 0.97 | 0.44 | 0.60 | 0.70 |
| SuperPoint [8] | 200 | 614 | 0.36 | 0.75 | 0.93 | 0.46 | 0.84 | 0.97 | 0.26 | 0.66 | 0.89 |
| BRISK [34] | 78 | 153 | 0.31 | 0.64 | 0.78 | 0.38 | 0.67 | 0.76 | 0.24 | 0.62 | 0.80 |
| SIFT [4] | 40 | 307.2 | 0.44 | 0.78 | 0.89 | 0.51 | 0.81 | 0.88 | 0.37 | 0.74 | 0.89 |
| ORB [6] | 20 | 76.8 | 0.17 | 0.43 | 0.58 | 0.28 | 0.47 | 0.57 | 0.07 | 0.41 | 0.60 |
| CAPS [48] | 5 | 614 | 0.22 | 0.55 | 0.8 | 0.35 | 0.71 | 0.88 | 0.08 | 0.41 | 0.66 |
| LANet [42] | 6 | 614 | 0.51 | 0.86 | 0.94 | 0.66 | 0.95 | 0.99 | 0.36 | 0.76 | 0.89 |
| Method | #Params (M) | FLOPs (G) @ 480 × 640 |
|---|---|---|
| BRISK [34] | n/a | n/a |
| ORB [6] | n/a | n/a |
| SIFT [4] | n/a | n/a |
| SuperPoint [8] | 1.3 | 52 |
| CAPS [48] | 20 | 223 |
| LANet [42] | 2.8 | 77.6 |
| FPC-Net (ours) | 1.6 | 8.4 |
| Dataset | K (Matches) | Rot (Ours) | Trans (Ours) | Rot (SIFT) | Trans (SIFT) |
|---|---|---|---|---|---|
| KITTI | 10 | 0.20 | 0.35 | 0.40 | 0.55 |
| 20 | 0.08 | 0.09 | 0.12 | 0.16 | |
| 30 | 0.05 | 0.06 | 0.08 | 0.07 | |
| 40 | 0.015 | 0.035 | 0.060 | 0.060 | |
| 50 | 0.004 | 0.020 | 0.045 | 0.060 | |
| 60 | 0.008 | 0.015 | 0.030 | 0.050 | |
| EuRoC | 10 | 0.38 | 0.29 | 0.47 | 0.33 |
| 20 | 0.25 | 0.19 | 0.36 | 0.34 | |
| 30 | 0.22 | 0.15 | 0.25 | 0.18 | |
| 40 | 0.20 | 0.15 | 0.25 | 0.14 | |
| 50 | 0.16 | 0.10 | 0.22 | 0.28 | |
| 60 | 0.17 | 0.08 | 0.22 | 0.15 |
| Variant | Repeatability ↑ | Homography Acc. ↑ | ||||
|---|---|---|---|---|---|---|
| @1 | @3 | @8 | @1 | @3 | @8 | |
| Full (Stage 2 + Huber + peak-separation) | 0.46 | 0.59 | 0.67 | 0.54 | 0.74 | 0.84 |
| Stage-2 (+Huber consistency) | 0.42 | 0.56 | 0.63 | 0.52 | 0.71 | 0.83 |
| Stage-2 (+KL consistency) | 0.41 | 0.55 | 0.62 | 0.51 | 0.70 | 0.82 |
| Stage-2 (no consistency) | 0.22 | 0.35 | 0.45 | 0.39 | 0.49 | 0.51 |
| Stage-1 only (no Stage 2) | 0.15 | 0.24 | 0.33 | 0.36 | 0.45 | 0.47 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Grigore-Atimuț, I.-O.; Leoveanu-Condrei, C.; Popa, C.-A. FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching. Appl. Sci. 2026, 16, 1223. https://doi.org/10.3390/app16031223
Grigore-Atimuț I-O, Leoveanu-Condrei C, Popa C-A. FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching. Applied Sciences. 2026; 16(3):1223. https://doi.org/10.3390/app16031223
Chicago/Turabian StyleGrigore-Atimuț, Ionuț-Orlando, Claudiu Leoveanu-Condrei, and Călin-Adrian Popa. 2026. "FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching" Applied Sciences 16, no. 3: 1223. https://doi.org/10.3390/app16031223
APA StyleGrigore-Atimuț, I.-O., Leoveanu-Condrei, C., & Popa, C.-A. (2026). FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching. Applied Sciences, 16(3), 1223. https://doi.org/10.3390/app16031223

