Next Article in Journal
Decision-Making Tools for Large Vessel Collisions with Marine Megafauna Species: Research Gaps and Proposed Application
Next Article in Special Issue
Comparative Evaluation of Fusion Strategies Using Multi-Pretrained Deep Learning Fusion-Based (MPDLF) Model for Histopathology Image Classification
Previous Article in Journal
STC-SORT: A Dynamic Spatio-Temporal Consistency Framework for Multi-Object Tracking in UAV Videos
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Robust Intraoral Image Stitching via Deep Feature Matching: Framework Development and Acquisition Parameter Optimization

Division of Artificial Intelligence Convergence Engineering, Sahmyook University, Seoul 01795, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(2), 1064; https://doi.org/10.3390/app16021064
Submission received: 30 December 2025 / Revised: 13 January 2026 / Accepted: 18 January 2026 / Published: 20 January 2026
(This article belongs to the Special Issue AI for Medical Systems: Algorithms, Applications, and Challenges)

Featured Application

This study contributes to cost-effective digital dental diagnostics by enabling high-quality 2D panoramic imaging using standard handheld RGB intraoral cameras.

Abstract

Low-cost RGB intraoral cameras are accessible alternatives to intraoral scanners; however, generating panoramic images is challenging due to narrow fields of view, textureless surfaces, and specular highlights. This study proposes a robust stitching framework and identifies optimal acquisition parameters to overcome these limitations. All experiments were conducted exclusively on a mandibular dental phantom model. Geometric consistency was further validated using repeated physical measurements of mandibular arch dimensions as ground-truth references. We employed a deep learning-based approach using SuperPoint and SuperGlue to extract and match features in texture-poor environments, enhanced by a central-reference stitching strategy to minimize cumulative drift errors. To validate the feasibility in a controlled setting, we conducted experiments on dental phantoms varying working distances (1.5–3.0 cm) and overlap ratios. The proposed method detected approximately 19–20 times more valid inliers than SIFT, significantly improving matching stability. Experimental results indicated that a working distance of 2.5 cm offers the optimal balance between stitching success rate and image detail for handheld operation, while a 1/3 overlap ratio yielded superior geometric integrity. This system demonstrates that robust 2D dental mapping is achievable with consumer-grade sensors when combined with advanced deep feature matching and optimized acquisition protocols.

1. Introduction

Digital dentistry has revolutionized clinical workflows by enabling precise 3D modeling of the oral cavity for diagnostics, prosthetics, and orthodontic treatment planning [1]. Intraoral scanners (IOS) have become the gold standard in this domain, offering high accuracy and real-time visualization capabilities [2]. However, despite their clinical benefits, the widespread adoption of high-performance IOS systems is often hindered by their high acquisition costs, complex maintenance requirements, and steep learning curves, particularly in primary care settings or developing regions [3]. Consequently, there is a growing interest in leveraging cost-effective RGB intraoral cameras (IOC) as accessible alternatives for dental documentation and telemedicine [4].
While RGB intraoral cameras are affordable and user-friendly, they suffer from a significant technical limitation: a narrow Field of View (FOV) [5]. A single image captures only a few teeth, necessitating the application of image stitching techniques to generate a comprehensive panoramic view of the dental arch [6]. However, the oral cavity presents a uniquely challenging environment for traditional computer vision algorithms. First, the tooth surface (enamel) is inherently textureless and features repetitive patterns, which often causes the failure of hand-crafted feature detectors such as Scale-Invariant Feature Transform (SIFT) or Oriented FAST and Rotated BRIEF (ORB) due to a lack of distinct keypoints [7]. Second, the oral environment is wet, leading to severe specular highlights caused by saliva and illumination sources. These reflections shift with camera movement, creating false features (outliers) that degrade the accuracy of image registration [8]. Recent studies on low-texture scene stitching have consistently highlighted the inadequacy of traditional methods in such adverse conditions [9]. In particular, sequential alignment along a curved mandibular arch often induces cumulative curvature distortion (the so-called banana effect), which degrades the geometric stability of 2D stitched panoramas.
The advent of deep learning has provided new avenues for overcoming these matching challenges. Learned local features, such as those extracted by Convolutional Neural Networks (CNNs), have demonstrated improved robustness against illumination changes and textureless regions compared to traditional descriptors [10]. In particular, the combination of SuperPoint [11] for self-supervised interest point detection and SuperGlue [12] for graph-based feature matching has set new benchmarks in geometric computer vision. Recent applications of these models in challenging environments, such as low-light remote sensing [13] and repetitive structural defect detection [14], suggest their potential applicability to the texture-poor and repetitive nature of dental imagery.
Indeed, recent advancements have further substantiated this potential in the dental domain. A systematic review [15] analyzed deep learning methodologies in dental image analysis, while other studies [16] demonstrated the automated processing of intraoral clinical photographs. Furthermore, deep learning techniques have been applied to point cloud patching in intraoral scanning to address data incompleteness [17]. In the context of geometric matching, the robustness of SuperPoint and SuperGlue has been validated in stereo vision SLAM [18], reinforcing the suitability of these descriptors for challenging estimation tasks.
However, robust algorithms alone cannot guarantee successful clinical implementation. The quality of handheld scanning is heavily influenced by acquisition parameters such as working distance and overlap ratio. While extensive studies have optimized scanning protocols for 3D IOS devices—identifying specific distance ranges (e.g., 5–10 mm) to minimize trueness errors [19,20]—similar guidelines for 2D handheld panoramic stitching are notably absent. Capturing images at an improper distance or with insufficient overlap can lead to geometric distortions, drift errors, and ultimately, stitching failure.
To address these gaps, this study proposes a robust stitching framework for RGB intraoral cameras and identifies the optimal acquisition parameters for handheld operation. We introduce a deep learning-based pipeline utilizing SuperPoint and SuperGlue to ensure robust feature matching in the presence of specular highlights and textureless surfaces. Furthermore, we experimentally investigate the impact of working distance (1.5 cm to 3.0 cm) and overlap ratio on stitching success rates and geometric integrity. This work aims to validate the feasibility of generating reliable 2D dental panoramas using consumer-grade sensors and to establish a standardized acquisition protocol for their clinical application. This study focuses exclusively on the mandibular arch using a physical dental phantom to provide a controlled and repeatable evaluation setting. The main contributions are as follows: (1) a deep feature matching-based stitching framework combined with a central-reference strategy to suppress cumulative curvature distortion, (2) experimental optimization of acquisition parameters (working distance and overlap) for low-cost handheld RGB intraoral cameras, and (3) quantitative geometric validation using repeated physical measurements of mandibular arch dimensions as ground-truth references.

2. Materials and Methods

2.1. System Configuration and Dataset Acquisition

The overall framework of the proposed intraoral image stitching system, including the deep feature matching pipeline and the acquisition parameter optimization strategy, is illustrated in Figure 1.
The imaging system utilized in this study consists of a commercial handheld RGB intraoral camera (EzCam; Vatech, Hwaseong, Republic of Korea) equipped with a CMOS sensor, delivering video sequences at a resolution of 1920 × 1080 pixels. To simulate clinical environments while controlling variables, a standard dental typodont (phantom) model Dental Model; Dubiduba, Beijing, China was used.
The dataset D consists of sequential frames captured by moving the camera along the dental arch:
D = I 1 , I 2 , , I N
where N represents the total number of frames in a single scanning sequence acquired under fixed acquisition parameters. The acquisition protocol was strictly controlled by varying the working distance ( d   { 1.5 ,   2.0 ,   2.5 ,   3.0 } cm) and the frame overlap ratio ( ρ { 1 / 3 ,   1 / 2 } ) to quantitatively analyze the impact of acquisition geometry on stitching performance. All sequences were acquired exclusively from the mandibular phantom. For the working-distance study, three independent scan sequences (v1–v3) were acquired at each distance (1.5, 2.0, 2.5, and 3.0 cm), and the frame counts per sequence are reported in the Appendix A (Table A1). For the overlap-ratio study, additional single scan sequences were acquired under approximately 1/3 and 1/2 overlap conditions at each distance; these sequences constitute the experimental trials used to compute the SSR in Section 2.6, and their frame counts are summarized in the experimental results section.
To ensure consistent data acquisition and minimize operator variability, a custom-developed software interface was utilized. As shown in the experimental setup (Figure 2), the software provides real-time visual guides (bounding boxes) on the live feed, consisting of a larger box for molars and a smaller box for anterior teeth. These guides assist the operator in maintaining a constant working distance and appropriate framing during the handheld scan, simulating a guided clinical workflow.

2.2. Pre-Processing: Contrast Limited Adaptive Histogram Equalization (CLAHE)

Intraoral images frequently suffer from uneven illumination and specular highlights caused by the wet surface of teeth. To enhance local contrast and suppress noise amplification in homogeneous regions (textureless enamel), we applied Contrast Limited Adaptive Histogram Equalization (CLAHE).
The image is divided into non-overlapping tiles of size M × M . For each tile, the histogram is computed and clipped at a predefined limit β to prevent over-amplification of noise. The redistribution of clipped pixels is governed by:
h c l i p i = β ,                                     i f   h i β h i +   h i β L ,   i f   h i < β
where h ( i ) is the histogram count for intensity level i , and L is the number of gray levels. Bilinear interpolation is then performed to eliminate artifacts at tile boundaries, yielding the pre-processed frame I t .

2.3. Deep Feature Extraction and Matching

Conventional handcrafted descriptors (e.g., SIFT, ORB) fail to extract sufficient keypoints in textureless dental surfaces. We employed a deep learning-based pipeline comprising SuperPoint for detection and SuperGlue for matching. We used publicly available pre-trained weights for SuperPoint and SuperGlue. No dental-domain fine-tuning was performed in this study, in order to evaluate out-of-the-box robustness under controlled phantom conditions.

2.3.1. SuperPoint Network

SuperPoint serves as a shared encoder–decoder architecture. The encoder maps the input image I R H × W to a feature map. Two decoder heads then operate in parallel:
Interest Point Decoder: Computes a probability map for keypoint locations using a softmax activation. The loss function for keypoint detection L d e t is defined as a cross-entropy loss between the predicted logits and the pseudo-ground truth labels.
Descriptor Decoder: Outputs a dense descriptor map D R H / 8 × W / 8 × 256 . The descriptor loss L d e s c is a hinge loss that enforces matching descriptors to be close and non-matching ones to be far apart:
L d e s c d i , d j , s i j = s i j max 0 , m p | | d i d j | | 2 + 1 s i j max 0 , | | d i d j | | 2 m n
where s i j = 1 if keypoints i and j correspond, and 0 otherwise. m p and m n are the positive and negative margins, respectively.

2.3.2. SuperGlue Matching with Graph Neural Networks

To handle the large displacement and repetitive patterns of teeth, we utilized SuperGlue, an attentional Graph Neural Network (GNN). Given two sets of local features p A , d A and p B , d B from images A and B , SuperGlue constructs a complete graph where nodes are keypoints.
The matching process is formulated as an optimal transport problem. The assignment matrix P 0 , 1 M × N is computed to maximize the total score i , j P i , j S i , j , where S i , j is the affinity score matrix derived from the GNN. The Sinkhorn algorithm is iteratively applied to enforce the doubly stochastic constraints:
P = Sinkhorn S
This method allows for robust rejection of outliers (mismatched teeth) by assigning them to a “dustbin” node, demonstrating substantially improved robustness compared to nearest neighbor search in specular environments.

2.4. Stitching Strategy: Central-Reference Homography and TPS Deformation

A major challenge in scanning the dental arch is the accumulation of drift error and local misalignments caused by the non-planar geometry of teeth. This recursive transformation typically results in the “banana effect” or severe distortion at the sequence end, as illustrated in Figure 3a. To address this, we implemented a two-stage alignment strategy combining global Central-Reference Homography and local Thin-Plate Spline (TPS) Deformation. First, to mitigate the drift error, we employed a central-reference strategy (Figure 3b). Instead of propagating transforms from the first frame, we select the central frame I r e f as the anchor. The global homography H k r e f is computed bidirectionally, halving the error accumulation path:
H k r e f = j = k r e f 1 H j , j + 1 for   k < r e f
However, homography is a rigid planar transformation that cannot perfectly align the curved 3D surfaces of teeth. To correct residual local misalignments, we applied Thin-Plate Spline (TPS) deformation. TPS creates a smooth interpolation surface that maps the matched keypoints of the source image exactly to the target coordinates while minimizing the “bending energy” of the transformation. The objective function E for the TPS mapping function f x , y is defined as:
E = R 2 2 f x 2 2 + 2 2 f x y 2 + 2 f y 2 2 d x d y
By minimizing this energy, the algorithm ensures that the image is warped naturally to align overlapping tooth structures without introducing sharp artifacts. This non-rigid deformation significantly improves the seamlessness of the final panorama.

2.5. Multi-Band Blending

Simple linear blending or seam cuts often result in visible artifacts due to exposure differences and moving specular highlights on the tooth surface. We adopted a Multi-band Blending approach.
The image is decomposed into a Laplacian pyramid. Let L l I be the l -th level of the Laplacian pyramid for image I , and G l W be the Gaussian pyramid for the weight mask W . The blended image B is reconstructed from the blended pyramid L l B :
L l B = k G l W k L l I k
where k iterates over the overlapping images. The final panoramic image is obtained by summing the upsampled levels of L l B . This technique effectively merges high-frequency details (features) while smoothing out low-frequency discrepancies (lighting variations).

2.6. Evaluation Metrics

To quantitatively evaluate the proposed framework and the impact of acquisition parameters, we defined the following metrics:
Number of Inliers ( N i n ): The count of valid feature matches after RANSAC geometric verification. Higher N i n indicates a robust registration.
Stitching Success Rate (SSR): The ratio of (i.e., no visible breaks or major misalignments according to predefined evaluation guidelines) to the total number of experimental trials.
S S R = Number   of   Successful   Panoramas Total   Trials × 100 %
Here, one experimental trial corresponds to one complete scan sequence acquired under a fixed working-distance and overlap configuration.
Visual Quality Score (VQS): A qualitative assessment of geometric integrity and clinical interpretability of stitched panoramas. Five non-dentistry evaluators (two imaging researchers and three faculty members) independently rated each stitched result using a predefined rubric with four criteria: (i) continuity of the mandibular arch, (ii) absence of visible seams or ghosting, (iii) absence of missing-tooth or duplicated-tooth artifacts, and (iv) preservation of local tooth geometry. Each panorama was categorized into High, Medium, or Low based on the number of criteria satisfied. Inter-observer agreement was quantified using overall percent agreement and Fleiss’ kappa, and the resulting values are reported to support reproducibility. As the evaluators were not dental specialists, a clinician-based user study will be conducted in future work to further validate clinical relevance.
Computational Performance (Runtime): We measured the end-to-end processing time per stitched panorama, defined as the total time from completing a scan sequence acquisition to generating the final blended panorama (including feature extraction/matching, warping, and multi-band blending). Measurements were performed on a workstation equipped with an NVIDIA GeForce RTX 4080 GPU (NVIDIA, Santa Clara, CA, USA).

3. Results

3.1. Evaluation of Deep Feature Matching Performance

To validate the effectiveness of the proposed deep learning-based pipeline, we compared it against the standard handcrafted feature detector, SIFT (Scale-Invariant Feature Transform). The evaluation was performed on a dataset of 50 image pairs captured from the textureless buccal surface of the dental typodont.
Qualitative comparisons are presented in Figure 4. The SIFT algorithm predominantly detected features on the gingival margin or high-contrast artifacts, failing to extract reliable keypoints on the smooth enamel surface. This resulted in sparse and often mismatched correspondence lines (Figure 4a). In contrast, the proposed method (SuperPoint + SuperGlue) successfully detected keypoints even on homogeneous tooth surfaces and established dense, parallel matching lines (Figure 4b), demonstrating robustness against specular highlights.
Quantitatively, Table 1 summarizes the average number of detected keypoints and valid inliers ( N i n ) after RANSAC geometric verification. The baseline SIFT method yielded an average of only 9.2 inliers per pair, which is close to the minimum requirement (4 points) for homography estimation, leading to frequent registration failures. The proposed method achieved an average of 179.5 inliers, representing an approximately 19.5-fold increase in matching density compared to SIFT. Consequently, the matching success rate improved from 20% to 98%.

3.2. Impact of Working Distance on Stitching Stability

We investigated the influence of working distance ( d ) on the quality of panoramic stitching. The camera was operated handheld at distances of 1.5 cm, 2.0 cm, 2.5 cm, and 3.0 cm. Figure 5 illustrates representative stitching results for each distance.
At 1.5 cm (Figure 5a): The stitching success rate was 0% (Low). Due to the extremely narrow Field of View (FOV), the number of teeth captured in a single frame was often less than two, providing insufficient geometric context for the homography estimation. Furthermore, slight hand tremors caused severe defocus blur, making feature detection impossible.
At 2.0 cm (Figure 5b): The success rate improved to 33%. While the images retained high detail, the system occasionally failed to register adjacent frames correctly when the camera moved too fast, leading to “missing tooth” artifacts or discontinuities.
At 2.5 cm (Figure 5c): This distance yielded the highest stability with a success rate of 66% (High) in handheld conditions. The FOV was sufficient to capture 2–3 teeth simultaneously, providing robust overlap for the central-reference strategy while maintaining adequate spatial resolution for clinical visualization.
At 3.0 cm (Figure 5d): While the success rate was 100% due to the wide FOV, the resulting panoramic image suffered from low resolution (Pixels Per Degree), reducing its diagnostic value for identifying fine cracks or caries.
Based on these results, we identified 2.5 cm as the optimal working distance for handheld RGB intraoral scanning.

3.3. Influence of Overlap Ratio

The effect of the frame overlap ratio ( ρ ) was analyzed by comparing sequences captured with approximately 1/3 and 1/2 overlap. Table 2 summarizes the outcomes. Table 3 provides a distance-aggregated qualitative summary of VQS outcomes and inter-observer agreement across all stitched panoramas.
Contrary to conventional photogrammetry where higher overlap is generally preferred, our results indicated that an approximately 1/3 overlap ratio produced superior geometric integrity in handheld intraoral stitching. Notably, the recorded frame counts did not consistently increase with higher overlap in our dataset (Table 2), suggesting that operator motion speed and stopping criteria influenced the number of captured frames. Therefore, the observed differences in geometric integrity are better explained by the overlap-dependent matching behavior rather than by frame count alone. Excessive overlap can produce highly redundant adjacent views of texture-poor enamel surfaces, which may increase symmetric or ambiguous correspondences and occasionally destabilize graph-based matching. In contrast, a moderate overlap of approximately 1/3 provides sufficient continuity while preserving an effective baseline for robust registration. Accordingly, an overlap ratio around 1/3 is recommended to balance geometric stability and practical handheld operation.

3.4. Quantitative Geometric Accuracy Using Mandibular Phantom Measurements

To address geometric accuracy beyond visual inspection and SSR, we performed a quantitative evaluation using repeated physical measurements of the mandibular phantom as ground-truth references. Two mandibular dimensions were considered: (i) inter-molar width, defined as the distance between the left and right second molars, and (ii) vertical arch depth, defined as the distance from the inter-molar line to the most anterior point of the mandibular arch. Ground-truth values were obtained by repeated manual measurements (five trials) and are reported as mean ± standard deviation.
On each stitched panorama, the same two distances were measured in pixel units using a straight-line tool. Because absolute pixel-to-millimeter scaling is not inherently preserved in 2D stitching, we report a scale-invariant ratio error as well as leave-one-out millimeter errors. Specifically, the ratio r = depth/width was computed for both the ground truth and the stitched panorama. We explicitly separated the conversion factors for width and depth to avoid circularity in calibration. This separation is essential because 2D stitching of a curved dental arch does not guarantee the preservation of the global scale isotropically; therefore, deriving independent scaling factors for orthogonal directions mitigates measurement errors caused by non-uniform geometric distortions. Additionally, for millimeter-based error reporting without circularity, we used the ground-truth width to convert the stitched depth to millimeters ( D ^ m m W ), and independently used the ground-truth depth to convert the stitched width to millimeters ( W ^ m m D ). The resulting absolute errors and the geometric RMS error are summarized in Table 4. On the RTX 4080 workstation, the end-to-end processing time per panorama ranged from 4 to 6 s across scan sequences (mean 5.0 s; half-range ± 1.0 s), depending on the number of frames in each scan sequence.

4. Discussion

4.1. Geometric Limitations of 2D Manifold Projection

While the proposed deep learning-based framework successfully addressed the local matching challenges on textureless and specular dental surfaces, our results revealed an inherent geometric limitation common to 2D stitching approaches: the “Arch Straightening Effect.”
As observed in the final panoramic images, the naturally U-shaped dental arch tends to flatten into a linear strip. This phenomenon occurs because the dental arch is fundamentally a 3-dimensional manifold with high curvature. When a sequence of images capturing this 3D structure is projected onto a 2D plane using pairwise homography transformations without 3D depth information, the algorithm minimizes the reprojection error between adjacent frames locally. However, lacking a global 3D constraint, these local optimizations accumulate, causing the global curvature to unroll.
Mathematically, homography H creates a projective transformation that preserves collinearity but not necessarily metric properties like angles or distances over long sequences. Consequently, the molar regions (distal ends of the arch) exhibit significant stretching distortion, appearing wider than their actual anatomical proportions. This suggests that while 2D stitching is highly effective for visualization and documentation (e.g., creating a “dental map”), it cannot fully replace 3D intraoral scanners for applications requiring high metric accuracy, such as prosthetic fabrication or orthodontic aligner planning. To overcome this, future studies must integrate Simultaneous Localization and Mapping (SLAM) or Neural Radiance Fields (NeRF) to recover the underlying 3D geometry.

4.2. Acquisition Parameter Optimization: The Trade-Off Between Stability and Resolution

Our experimental results highlight that the quality of handheld intraoral scanning is not solely dependent on the software algorithm but is critically governed by the acquisition parameters, specifically working distance and overlap ratio.
The identification of 2.5 cm as the optimal working distance can be explained by the trade-off between Feature Density and Spatial Resolution. At a distance of 1.5 cm, the Field of View (FOV) is extremely narrow, capturing less than one complete tooth per frame. This lack of distinct geometric context leads to the “aperture problem,” where the algorithm cannot distinguish the relative motion of the camera, resulting in a 0% success rate. Conversely, at 3.0 cm, while stability is maximized due to the inclusion of multiple teeth and gingival landmarks, the effective resolution (pixels per millimeter) drops significantly, potentially obscuring fine clinical details such as early enamel cracks or marginal discrepancies. Furthermore, a critical hardware limitation was observed regarding the autofocus mechanism. Commercial RGB cameras typically rely on contrast-based autofocus, which is prone to “focus hunting” behaviors, particularly at unstable working distances (e.g., 1.5 cm handheld) or during rapid movement. This instability creates a variable depth of field within a single sequence, causing intermittent blur that degrades the feature descriptor quality. Therefore, the superior stability observed at 2.5 cm is attributed not only to the optimal field of view but also to the maintenance of a consistent focal plane compared to the macro range. Regarding the overlap ratio, our finding that an approximately 1/3 overlap outperforms 1/2 overlap contradicts the common intuition in photogrammetry, which typically advocates higher overlap. In the context of pairwise homography-based stitching on texture-poor dental enamel, higher overlap does not necessarily improve correspondence quality because adjacent views can become overly redundant and ambiguous, especially under specular highlights and repetitive tooth patterns. As summarized in Table 2, the recorded frame counts were not a monotonic function of overlap, indicating that practical handheld scanning dynamics also affect acquisition. Therefore, we attribute the superior stability at approximately 1/3 overlap to a more favorable balance between continuity and discriminative motion baseline for robust matching, rather than to frame count alone. Because the recorded frame counts were not a monotonic function of overlap in our handheld acquisition (Table 2), we do not attribute the observed geometric differences to frame count alone. Instead, we interpret the superior stability at approximately 1/3 overlap as a more favorable balance between continuity and discriminative baseline for robust matching.

4.3. Clinical Feasibility and Practical Applications

The proposed system, using a standard RGB camera, demonstrates clear clinical utility as a cost-effective screening and documentation tool. Unlike traditional intraoral scanners that require significant investment and training, our approach utilizes equipment already available in many dental practices.
The generated panoramic images serve as excellent “Patient Communication Tools.” They allow clinicians to visualize the entire quadrant or arch in a single view, facilitating the explanation of hygiene status, multiple carious lesions, or gingival health to patients. Furthermore, this 2D mapping technology is particularly valuable for Tele-dentistry in remote areas, where transmitting a lightweight 2D panorama is far more bandwidth-efficient than sending large 3D mesh files.
However, clinicians must be aware of the metrological limitations. Due to the aforementioned straightening and stretching distortions, measurements taken directly from the stitched panorama (e.g., measuring the edentulous space for an implant) may not be accurate. Therefore, this technology should be positioned as a diagnostic aid rather than a measurement device.

4.4. Limitations and Future Directions

Despite the significant improvements in matching robustness, this study has limitations. First, the evaluation was primarily conducted on dental phantoms. All experiments were conducted on a mandibular phantom only; therefore, dynamic intraoral factors (tongue motion, saliva, patient movement) were not present in the evaluation setup. In future in vivo deployments, frame-quality filtering and motion-robust acquisition protocols will be incorporated to mitigate these effects. Although phantoms simulate the geometry of teeth, they lack the physiological micro-movements of the patient (e.g., tongue movement, breathing) and the variable saliva flow found in real clinical settings. Future in vivo studies are necessary to validate the robustness of the SuperPoint/SuperGlue pipeline against dynamic biological artifacts.
Second, the current processing time for deep feature matching is relatively high compared to real-time SIFT implementations. To enable live feedback for the clinician during scanning, model quantization or distillation techniques should be explored to accelerate the inference speed on edge devices.
Finally, the “straightening” artifact remains a critical hurdle. We propose that future research focuses on Hybrid Approaches that combine 2D deep stitching with lightweight depth estimation networks (Monocular Depth Estimation) to correct the manifold curvature, bridging the gap between 2D panoramas and true 3D reconstruction.

5. Conclusions

This study successfully demonstrated the feasibility of generating reliable 2D dental panoramas using consumer-grade RGB intraoral cameras, overcoming the inherent challenges of textureless enamel surfaces and severe specular highlights. By integrating a deep learning-based pipeline utilizing SuperPoint and SuperGlue, we achieved robust feature matching with approximately 20 times more valid inliers compared to traditional SIFT-based methods, thereby enhancing registration stability in handheld scanning scenarios.
Furthermore, our systematic experiments on acquisition parameters revealed that the quality of panoramic stitching is highly sensitive to the geometric conditions of data capture. We identified that a working distance of 2.5 cm provides the optimal balance between stitching success rate and spatial resolution, while a 1/3 overlap ratio minimizes cumulative drift errors better than higher overlap settings. These findings establish a practical guideline for clinicians using handheld devices.
Although the proposed 2D framework inherently involves geometric distortions such as arch straightening, limiting its application in precision metric tasks like prosthetics, it offers substantial value as a cost-effective tool for dental screening, patient education, and telemedicine. Future research will focus on extending this framework to 3D reconstruction techniques, such as Neural Radiance Fields (NeRF) or 3D Gaussian Splatting, to recover accurate depth information and address the geometric limitations of 2D projection.

Author Contributions

Conceptualization, J.-S.J. and S.W.C.; methodology, J.-S.J. and S.W.C.; software, J.-S.J. and D.-J.S.; validation, D.-J.S.; formal analysis, J.-S.J.; investigation, D.-J.S.; resources, S.W.C.; data curation, J.-S.J.; writing—review and editing, S.W.C.; visualization, J.-S.J. and D.-J.S.; supervision, S.W.C.; project administration, S.W.C.; funding acquisition, S.W.C. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by National Research Foundation of Korea (NRF) grants funded by the Korean government (MSIT) (RS-2024-00445552).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Dataset available on request from the authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
IOSIntraoral Scanner
IOCIntraoral Camera
RGBRed Green Blue
CMOSComplementary Metal-Oxide-Semiconductor
FOVField of View
SIFTScale-Invariant Feature Transform
ORBOriented FAST and Rotated BRIEF
CLAHEContrast Limited Adaptive Histogram Equalization
CNNConvolutional Neural Network
GNNGraph Neural Network
RANSACRandom Sample Consensus
SSRStitching Success Rate
SLAMSimultaneous Localization and Mapping
NeRFNeural Radiance Fields

Appendix A

Table A1. Distribution of frame counts for independent scan sequences at each experimental working distance.
Table A1. Distribution of frame counts for independent scan sequences at each experimental working distance.
Working Distance (cm)v1 (Frames)v2 (Frames)v3 (Frames)Mean (Frames)SD (Frames)Range (Min–Max)
1.524202222.02.020–24
2.022242323.01.022–24
2.522212021.01.020–22
3.018161416.02.014–18

References

  1. Revilla-León, M.; Meyer, M.J.; Özcan, M. Metal Additive Manufacturing Technologies: Literature Review of Current Status and Prosthodontic Applications. Int. J. Comput. Dent. 2019, 22, 55–67. [Google Scholar] [PubMed]
  2. Mangano, F.; Gandolfi, A.; Luongo, G.; Logozzo, S. Intraoral Scanners in Dentistry: A Review of the Current Literature. BMC Oral Health 2017, 17, 149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Azhari, A.A.; Ahmed, W.M.; Khider, T.; Almaghrabi, R.; Alharbi, R.; Merdad, Y.; Bukhari, S.; Lahiq, A.; Azhari, A.A.; Ahmed, W.M.; et al. Comparison of Digital Intraoral Scanning and Conventional Techniques for Post Space Capture. Prosthesis 2025, 7, 87. [Google Scholar] [CrossRef] [Scilit]
  4. Ehrensperger, C.; Körner, P.; Svellenti, L.; Attin, T.; Sahrmann, P.; Ehrensperger, C.; Körner, P.; Svellenti, L.; Attin, T.; Sahrmann, P. Evaluation of an Intraoral Camera with an AI-Based Application for the Detection of Gingivitis. J. Clin. Med. 2025, 14, 5580. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Frame Stitching in Human Oral Cavity Environment Using Intraoral Camera|IEEE Conference Publication|IEEE Xplore. Available online: https://ieeexplore.ieee.org/abstract/document/8803071 (accessed on 17 December 2025).
  6. Liu, J.; Li, X.; Shen, S.; Jiang, X.; Chen, W.; Li, Z.; Liu, J.; Li, X.; Shen, S.; Jiang, X.; et al. Research on Panoramic Stitching Algorithm of Lateral Cranial Sequence Images in Dental Multifunctional Cone Beam Computed Tomography. Sensors 2021, 21, 2200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Kang, L.; Wei, Y.; Jiang, J.; Xie, Y.; Kang, L.; Wei, Y.; Jiang, J.; Xie, Y. Robust Cylindrical Panorama Stitching for Low-Texture Scenes Based on Image Alignment Using Deep Learning and Iterative Optimization. Sensors 2019, 19, 5310. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Islam, M.N.; Tahtali, M.; Pickering, M.; Islam, M.N.; Tahtali, M.; Pickering, M. Specular Reflection Detection and Inpainting in Transparent Object through MSPLFI. Remote Sens. 2021, 13, 455. [Google Scholar] [CrossRef] [Scilit]
  9. Tchinda, E.N.; Panoff, M.K.; Kwadjo, D.T.; Bobda, C.; Tchinda, E.N.; Panoff, M.K.; Kwadjo, D.T.; Bobda, C. Semi-Supervised Image Stitching from Unstructured Camera Arrays. Sensors 2023, 23, 9481. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Yang, M.; Wu, R.; Yang, Y.; Tao, L.; Zhang, Y.; Xie, Y.; Reddy, G.P.R.D.; Yang, M.; Wu, R.; Yang, Y.; et al. Image Matching: Foundations, State of the Art, and Future Directions. J. Imaging 2025, 11, 329. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperPoint: Self-Supervised Interest Point Detection and Description. arXiv 2018, arXiv:1712.07629. [Google Scholar]
  12. Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; Rabinovich, A. SuperGlue: Learning Feature Matching With Graph Neural Networks. arXiv 2020, arXiv:1911.11763. [Google Scholar] [CrossRef] [Scilit]
  13. Li, F.; Chen, Y.; Shi, Q.; Shi, G.; Yang, H.; Na, J.; Li, F.; Chen, Y.; Shi, Q.; Shi, G.; et al. Improved Low-Light Image Feature Matching Algorithm Based on the SuperGlue Net Model. Remote Sens. 2025, 17, 905. [Google Scholar] [CrossRef] [Scilit]
  14. Gan, Z.; Teng, L.; Chang, Y.; Feng, X.; Gao, M.; Gao, X.; Gan, Z.; Teng, L.; Chang, Y.; Feng, X.; et al. A Multi-Information Fusion Method for Repetitive Tunnel Disease Detection. Sustainability 2024, 16, 4285. [Google Scholar] [CrossRef] [Scilit]
  15. Zhou, Z.; Zhu, J.; Zhang, Y.; Guan, X.; Wang, P.; Li, T. Deep Learning in Dental Image Analysis: A Systematic Review of Datasets, Methodologies, and Emerging Challenges. arXiv 2025, arXiv:2510.20634. [Google Scholar] [CrossRef] [Scilit]
  16. Eum, J.; Seo, H.; Jeong, T.; Lee, E.; Park, S.; Shin, J. Automated Processing of Intraoral Clinical Photographs Using Deep Learning Techniques. J. Clin. Pediatr. Dent. 2026, 50, 114–125. [Google Scholar]
  17. Zheng, Q.; Wang, Y.; Zhou, M.; Wu, Y.; Chen, J.; Wang, X.; Gao, L.; Kang, T.; Chen, X.; Zhang, W. Automatic Point Cloud Patching of Intraoral Three-Dimensional Scanning Based on Deep Learning. Int. Dent. J. 2025, 75, 100911. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Yoon, S.-W.; Park, S.-Y. Stereo Vision SLAM with SuperPoint and SuperGlue. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2024, XLVIII–4–W11–2024, 183–188. [Google Scholar] [CrossRef] [Scilit]
  19. Hokayem, P.; Bourgi, R.; Cuevas-Suárez, C.E.; Fernández-Barrera, M.Á.; Zamarripa-Calderón, J.E.; Tohme, H.; Saleh, A.; Nassar, N.; Lukomska-Szymanska, M.; Hardan, L. Optimization of Scanning Distance for Three Intraoral Scanners from Different Manufacturers: An In Vitro Accuracy Analysis. Prosthesis 2025, 7, 88. [Google Scholar] [CrossRef] [Scilit]
  20. Rotar, R.N.; Faur, A.B.; Pop, D.; Jivanescu, A. Scanning Distance Influence on the Intraoral Scanning Accuracy—An In Vitro Study. Materials 2022, 15, 3061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Overview of the proposed intraoral image stitching framework. The system addresses the challenges of textureless and specular dental surfaces by employing a deep learning-based pipeline (SuperPoint and SuperGlue). The stitching process integrates a central-reference strategy with Thin-Plate Spline (TPS) deformation to minimize geometric distortion, followed by multi-band blending to generate a seamless 2D panoramic image.
Figure 1. Overview of the proposed intraoral image stitching framework. The system addresses the challenges of textureless and specular dental surfaces by employing a deep learning-based pipeline (SuperPoint and SuperGlue). The stitching process integrates a central-reference strategy with Thin-Plate Spline (TPS) deformation to minimize geometric distortion, followed by multi-band blending to generate a seamless 2D panoramic image.
Applsci 16 01064 g001
Figure 2. Screenshot of the custom-developed software interface. The real-time visual guides (red bounding boxes) assist the operator in maintaining consistent framing. (Note: The inner box targets anterior teeth, and the outer box targets molars).
Figure 2. Screenshot of the custom-developed software interface. The real-time visual guides (red bounding boxes) assist the operator in maintaining consistent framing. (Note: The inner box targets anterior teeth, and the outer box targets molars).
Applsci 16 01064 g002
Figure 3. Schematic comparison of stitching strategies. (a) Sequential stitching propagates transformation errors unidirectionally ( 1 N ), leading to significant drift and curvature distortion, often referred to as the “banana effect.” (b) The proposed central-reference strategy distributes the transformation path bidirectionally from the center frame ( I r e f ), effectively halving the cumulative error path and preserving the linear geometry of the dental arch. All illustrated stitching results are obtained from mandibular phantom scan sequences.
Figure 3. Schematic comparison of stitching strategies. (a) Sequential stitching propagates transformation errors unidirectionally ( 1 N ), leading to significant drift and curvature distortion, often referred to as the “banana effect.” (b) The proposed central-reference strategy distributes the transformation path bidirectionally from the center frame ( I r e f ), effectively halving the cumulative error path and preserving the linear geometry of the dental arch. All illustrated stitching results are obtained from mandibular phantom scan sequences.
Applsci 16 01064 g003
Figure 4. Qualitative comparison of feature matching performance on textureless dental surfaces. (a) The baseline method (SIFT) fails to extract sufficient valid keypoints due to specular highlights, resulting in sparse and mismatched correspondence lines (outliers). (b) The proposed method (SuperPoint + SuperGlue) demonstrates robust performance, establishing dense and parallel matching lines even in regions with repetitive patterns and varying illumination. Note: The text overlays (e.g., Keypoints, Matches) in (b) represent real-time output logs from the matching pipeline, provided to ensure consistency with the quantitative data in Table 1. Additionally, the multicolored correspondence lines are utilized to visually distinguish individual feature matches for improved clarity.
Figure 4. Qualitative comparison of feature matching performance on textureless dental surfaces. (a) The baseline method (SIFT) fails to extract sufficient valid keypoints due to specular highlights, resulting in sparse and mismatched correspondence lines (outliers). (b) The proposed method (SuperPoint + SuperGlue) demonstrates robust performance, establishing dense and parallel matching lines even in regions with repetitive patterns and varying illumination. Note: The text overlays (e.g., Keypoints, Matches) in (b) represent real-time output logs from the matching pipeline, provided to ensure consistency with the quantitative data in Table 1. Additionally, the multicolored correspondence lines are utilized to visually distinguish individual feature matches for improved clarity.
Applsci 16 01064 g004
Figure 5. Representative stitching results on the mandibular dental phantom according to working distance: (a) 1.5 cm: Stitching failed due to the narrow Field of View (FOV) and unstable autofocus; (b) 2.0 cm: Successful stitching, but prone to minor misalignments during rapid movement; (c) 2.5 cm: Optimal results showing a balance between geometric stability and image resolution; (d) 3.0 cm: Stable stitching but with reduced spatial resolution (low pixels per degree), potentially obscuring fine clinical details. All examples were obtained from controlled phantom experiments under handheld operation.
Figure 5. Representative stitching results on the mandibular dental phantom according to working distance: (a) 1.5 cm: Stitching failed due to the narrow Field of View (FOV) and unstable autofocus; (b) 2.0 cm: Successful stitching, but prone to minor misalignments during rapid movement; (c) 2.5 cm: Optimal results showing a balance between geometric stability and image resolution; (d) 3.0 cm: Stable stitching but with reduced spatial resolution (low pixels per degree), potentially obscuring fine clinical details. All examples were obtained from controlled phantom experiments under handheld operation.
Applsci 16 01064 g005
Table 1. Performance comparison of feature detection and matching between SIFT and the proposed framework.
Table 1. Performance comparison of feature detection and matching between SIFT and the proposed framework.
MethodAvg. Detected KeypointsAvg. Valid Inliers (Nin)Matching Success Rate (%)
SIFT + FLANN 204   ± 52 9.2   ± 4.120.0
Proposed (SuperPoint + SuperGlue)1024 (Fixed) 179.5   ± 23.498.0
Table 2. Impact of working distance and frame overlap ratio on stitching success rate (SSR), geometric integrity, and clinical feasibility.
Table 2. Impact of working distance and frame overlap ratio on stitching success rate (SSR), geometric integrity, and clinical feasibility.
Working Distance (cm)Overlap RatioStitching Success RateGeometric Integrity# Sequences (Trials)Frames/SequenceClinical Feasibility
1.51/3Low Poor123Infeasible
1/2LowPoor117Infeasible
21/3Medium Medium120Moderate
1/2MediumMedium114Moderate
2.51/3HighHigh122Recommended
1/2HighHigh114Recommended
31/3High High120Low
1/2HighHigh115Low
Table 3. Qualitative summary of Visual Quality Score (VQS) distribution and inter-observer agreement among five independent raters.
Table 3. Qualitative summary of Visual Quality Score (VQS) distribution and inter-observer agreement among five independent raters.
GroupMetricValue
VQS distribution (final label)High (n, %)9 (45.0)
Medium (n, %)7 (35.0)
Low (n, %)4 (20.0)
Total (N, %)20 (100.0)
Summary score (final label)Weighted mean VQS2.25
Inter-observer agreement (5 raters)Complete agreement (5/5 exact match, %)20.0
Mean pairwise agreement (%)57.5
Fleiss’ kappa (κ)0.352
InterpretationFair
Table 4. Quantitative geometric accuracy evaluation based on repeated physical measurements of the mandibular phantom as ground-truth references.
Table 4. Quantitative geometric accuracy evaluation based on repeated physical measurements of the mandibular phantom as ground-truth references.
MetricGround Truth (Mean ± SD, mm)Stitched Measurement (Pixels)Scale-Invariant Ratio/Derived ValuesError
Inter-molar width ( W )46.8 ± 0.45 W p x = 562.60 ± 13.94 W ^ m m D = W p x × (41.8/ D p x )| W ^ m m D − 46.8| = 0.396 ± 0.257 mm
Vertical arch depth (D)41.8 ± 0.45 D p x = 506.76 ± 11.24 D ^ m m W = D p x × (46.8/ W p x )| D ^ m m W − 41.8| = 0.358 ± 0.233 mm
Ratio (r = D/W) r G T = 0.893 (±0.013) r s t i t c h e d = D p x / W p x
= 0.901 ± 0.005
ratio error = | r s t i t c h e d r G T |0.00765 ± 0.00498
Geometric RMS error--RMS = e r r W 2 + e r r D 2 / 2 0.377 ± 0.245 mm
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Jeong, J.-S.; Seong, D.-J.; Choi, S.W. Robust Intraoral Image Stitching via Deep Feature Matching: Framework Development and Acquisition Parameter Optimization. Appl. Sci. 2026, 16, 1064. https://doi.org/10.3390/app16021064

AMA Style

Jeong J-S, Seong D-J, Choi SW. Robust Intraoral Image Stitching via Deep Feature Matching: Framework Development and Acquisition Parameter Optimization. Applied Sciences. 2026; 16(2):1064. https://doi.org/10.3390/app16021064

Chicago/Turabian Style

Jeong, Jae-Seung, Dong-Jun Seong, and Seong Wook Choi. 2026. "Robust Intraoral Image Stitching via Deep Feature Matching: Framework Development and Acquisition Parameter Optimization" Applied Sciences 16, no. 2: 1064. https://doi.org/10.3390/app16021064

APA Style

Jeong, J.-S., Seong, D.-J., & Choi, S. W. (2026). Robust Intraoral Image Stitching via Deep Feature Matching: Framework Development and Acquisition Parameter Optimization. Applied Sciences, 16(2), 1064. https://doi.org/10.3390/app16021064

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop