Skip to Content
InformationInformation
  • Article
  • Open Access

28 January 2026

21 Pages

Defect Generation and Detection Strategy for Tempered Glass in Sample-Scarce Scenarios

,
,
,
,
,
and
1
Shandong Key Laboratory of CNC Machine Tool Functional Components, School of Mechanical Engineering, Qilu University of Technology (Shandong Academy of Sciences), Jinan 250353, China
2
Taian City Taishan Huijin Intelligent Technology Co., Ltd., Taian 271000, China
*
Authors to whom correspondence should be addressed.
This article belongs to the Section Artificial Intelligence

Abstract

To address the challenge of defect detection in tempered glass panel production rising from sample scarcity, this paper proposes a few-shot detection methodology that integrates an enhanced Stable Diffusion model with Mask R-CNN. Specifically, the approach utilizes a Mask Encoder to optimize the Stable Diffusion architecture, employing the Structural Similarity Index Measure (SSIM) to evaluate sample quality. This process generates high-fidelity virtual samples to construct a hybrid dataset for training data augmentation. Furthermore, a resource isolation strategy is adopted to facilitate online detection using an improved semi-supervised Mask R-CNN framework. Experimental results demonstrate that the proposed scheme effectively resolves detection difficulties for eight defect types, including edge chipping and scratches. The method achieves an mAP50 of 81.5%, representing a nearly 47% improvement over baseline methods relying solely on real samples, thereby realizing high-precision and high-efficiency industrial defect detection.

1. Introduction

Hollow glass is widely used in industries such as construction and home decoration due to its excellent fire-resistant, thermal-insulating, and heat-insulating properties. The production process mainly involves: cutting the glass sheets, surface coating, re-cutting, edge polishing, cleaning, de-coating, defect detection, TPSS (Thermo Plastic Space Sealant for Insulating Glass Unit), laminating, sealing with the glue line, and finally producing the finished hollow glass. Within this process, defect detection of tempered glass is a critical step in ensuring the quality of the finished hollow glass. With rapid economic development and growing consumer demands, the demand for tempered glass panels is increasing. However, during production, tempered glass panels are often affected by factors such as production processes, equipment, and the environment, leading to surface defects like chipping, delamination, oxidation, and staining. Some defects compromise the physical structure of the glass, reducing its quality and safety, such as chipping, while others affect its aesthetic appearance, like stains. In the early stages, traditional defect detection for tempered glass mainly relied on manual visual inspection, which was time-consuming, highly dependent on experience, and difficult to detect smaller defects. This method was not only inefficient but also lacked consistency in detection quality.
In recent years, with the rapid advancement of machine vision and computer technology, these technologies have shown immense potential in areas like target recognition speed, accuracy, and automation. They can be broadly categorized into traditional methods and deep learning-based methods, especially in large-scale and high-precision tasks. Traditional object detection algorithms typically use a combination of manually designed features, sliding windows, and classifiers. The core idea is to “exhaustively search candidate regions ➡ extract features ➡ classify/regress.” This approach has limited feature representation capability, struggles to adapt to complex lighting conditions, and often generates redundant candidate regions due to the sliding window method. This method, especially when processing large image datasets, consumes significant computational power and time, leading to slow detection speed and limited accuracy. With the introduction of the AlexNet model, deep learning experienced explosive growth, overcoming challenges related to data scale and generalization. It enhanced the modeling capability for nonlinear and high-dimensional data, improving feature representation through nonlinear transformations, and the same network architecture can adapt to various learning tasks. Although deep learning is widely applied today, traditional deep learning object recognition networks still require large datasets for model training. When the training dataset is insufficient, there is a high risk of overfitting.

3. Method

To ensure the production quality of hollow glass, a common automated production process is as follows: glass sheet cutting → surface coating → cutting → edge polishing → cleaning → de-coating → defect detection → TPSS → laminating → sealing with the glue line → finished hollow glass. A flowchart of this process is shown in Figure 1. The main contribution of this paper is the use of industrial cameras for defect detection on tempered glass sheets, and the proposed detection method and workflow are outlined in Figure 2. To address the challenge of data scarcity in few-shot scenarios, we employed a hybrid augmentation strategy combining traditional techniques (e.g., rotation) with an enhanced Stable Diffusion model. As shown in Figure 3, this synthetic dataset was utilized to train an optimized Mask R-CNN model. The generative model reconstructs defects by learning key feature representations from the original data. Qualitative analysis demonstrates that the synthetic samples retain essential characteristics—specifically the irregular edge curvature of the collected samples—while introducing distinct morphological variations to ensure heterogeneity. Consequently, the proposed approach effectively enhances dataset diversity, thereby improving the robustness of few-shot defect detection.
Figure 1. Automated Production Process of Insulated Glass.
Figure 2. Presents the defect detection method and workflow proposed herein.
Figure 3. Defect generation samples (Oxidation, Film Removal, Edge Chipping).
During the offline training phase, the diversity of the original defect dataset is first expanded using traditional augmentation techniques, such as rotation. This augmented dataset then serves as the basis for generating new defect images via the enhanced Stable Diffusion model. Approximately 4000 synthetic images are generated for each defect category. These are combined with the original collected images to construct a comprehensive dataset, which is then utilized to train the semi-supervised Mask R-CNN model.

3.1. Offline Training

During the offline training phase, traditional data augmentation techniques combined with the optimized stable diffusion strategy are employed to enhance sample diversity. To ensure the diversity of generated images, after collecting the original images, sample diversity is initially increased through techniques such as rotation. After traditional data augmentation, the stable diffusion strategy is applied to generate images, further increasing the diversity of the sample images. The structure of the stable diffusion model is shown in Figure 4.
Figure 4. Schematic diagram of the stable diffusion model network.
In simple terms, the diffusion model works by iteratively adding Gaussian noise to “corrupt” the training data, and then learning how to remove the noise to recover the data. The recovered data is not the original training data but new training data generated from the original data. The process consists of two main iterative stages: forward diffusion and reverse diffusion. In the forward diffusion phase, noise is gradually introduced to degrade the image until it becomes pure random noise, as shown in Figure 5.
Figure 5. Forward Diffusion (until it becomes pure noise).
It is described by the formula as follows [20]:
q x t x 0 = N ( x t ; 1 − β t x t − 1 , β t I )
Here, t represents the time frame (ranging from 0 to T), x t is a data sample drawn from the true data distribution q ( x ) (for example, x 0 ~ q ( x ) ), and β T is the variance schedule, which lies between 0 and 1, with β 0 being smaller and β T being larger.
In the reverse diffusion phase, the predicted noise is removed, and the data is recovered from the Gaussian noise, as shown in Figure 6.
Figure 6. Reverse diffusion (removing predicted noise and generating new image).
It is described by the formula as follows [20]:
q x t − 1 x t = N ( x t − 1 ; u ~ t ( x t , x 0 ) , β t ~ I )
Since q x t − 1 x t   cannot be calculated, it is necessary to train p θ x t − 1 x t to approximate q x t − 1 x t [20]:
p θ x t − 1 x t = ( x t − 1 ; μ θ x t , t ,   ∑ θ x t , t )
Here, p θ x t − 1 x t follows a normal distribution, and its mean and variance need to satisfy [20]:
μ θ x t , t   :   =   u ~ t ( x t , x 0 ) ∑ θ x t , t   :   =   β t ~ I
Traditional Stable Diffusion models typically rely on the CLIP encoder to process text prompts. However, linguistic descriptions lack the granularity required to precisely control the morphology and spatial localization of industrial defects, such as elongated scratches or micro-bubbles. To address this, we propose a conditional guidance mechanism based on a Mask Encoder. By replacing the conventional CLIP text encoder with a single-channel mask input, we achieve pixel-level precision in defining defect shape and position. The network architecture of the Mask Encoder is detailed in Figure 7. Furthermore, we constructed a glass panel defect generation pipeline based on the Latent Diffusion Model (LDM). Utilizing a Variational Autoencoder (VAE) for perceptual compression and a U-Net denoising network, this architecture ensures the natural fusion of defects while preserving the high-frequency textural details of the glass panel. The overall model structure is illustrated in Figure 8.
Figure 7. Mask Encoder.
Figure 8. Improved model structure.
The Mask Encoder module accepts a single-channel binary mask as input and employs a multi-layer Convolutional Neural Network (CNN) to extract spatial features. The resulting feature map is flattened and transposed into a (Batch, Sequence Length, Dimension) format—with a dimension of 768—to align with the requirements of the U-Net’s cross-attention mechanism. A key advantage of this architecture is that it compels the model to attend specifically to the masked regions. By encoding the shape information into a high-dimensional vector, the mask effectively functions as a conditioning “prompt” for defect generation.
To ensure the seamless integration of defect features with the glass background, we utilize a U-Net architecture incorporating cross-attention mechanisms. At each timestep of the diffusion process, the U-Net receives noisy latent variables and timestep embeddings as inputs. Simultaneously, the encoded mask features are injected into the network via the cross-attention layers:
A t t e n t i o n Q , K , V = s o f t m a x Q K T d V
Specifically, Q ( Q u e r y ) is derived from the intermediate features of the U-Net (representing the glass background), while K K e y and V ( V a l u e ) correspond to the context features output by the Mask Encoder (encoding the defect morphology).
Through this mechanism, the enhanced generative model effectively “inpaints” defects at precise spatial locations dictated by the mask. Crucially, it maintains the continuity of the surrounding background, thereby preventing artifacts or structural distortions at the defect boundaries.
The proposed generation pipeline consists of the following stages:
  • Conditional Encoding: A binary mask representing the defect morphology is input into the Mask Encoder to generate the conditioning vector.
  • Background Encoding: The original glass image is encoded into an initial latent variable via the VAE Encoder.
  • Reverse Diffusion: During each denoising step, classifier-free guidance is employed to enhance the semantic consistency between the generated defect and the input mask. To ensure that non-defect regions remain identical to the original image, the model-predicted latent variable is fused with the noisy original latent variable after each sampling step, as defined by:
z t − 1 = M ⨀ z p r e d + ( 1 − M ) ⨀ z o r i g i n a l _ n o i s y
where M denotes the downsampled mask.
To validate the plausibility of the synthetic samples, the Structural Similarity Index Measure (SSIM) is introduced as a quality evaluation metric. The assessment protocol is defined as follows: a threshold of 0.9 is established; each generated sample is compared against the original image. Samples yielding an SSIM score below 0.9 are discarded as non-compliant, while those meeting the threshold are retained. The threshold of 0.9 was determined through pilot experiments aimed at balancing sample fidelity and diversity. Our empirical observations indicated that generated samples with SSIM scores below 0.9 frequently exhibited noticeable background artifacts or structural distortions inconsistent with real industrial environments. Conversely, forcing an SSIM score significantly higher than 0.9 (e.g., >0.95) excessively constrained the morphological variations in the defects, thereby diminishing the effectiveness of data augmentation. Therefore, 0.9 was selected as the optimal cutoff. This high threshold serves to maintain the structural integrity of the dominant background, while the heterogeneity of the defect features is independently driven by the random noise injection and the specific guidance of the Mask Encoder. This process iterates until the target sample quantity is achieved. The SSIM formula is defined as:
S S I M x , y = ( 2 μ x μ y + C 1 ) ( 2 σ x y + C 2 ) ( μ x 2 + μ y 2 + C 1 ) ( σ x 2 + σ y 2 + C 2 )
Here, μ x  and μ y represent the mean pixel intensities of the original and generated image windows, respectively, which correspond to luminance. The terms σ x 2 and σ y 2 denote the variances of the respective windows, reflecting image contrast. The term σ x y signifies the covariance between the original and generated images, indicating their structural correlation. C 1 and C 2 are small constants introduced to ensure numerical stability by preventing division by zero. They are typically defined as C 1 = ( K 1 L ) 2 and C 2 = ( K 2 L ) 2 , where K 1 = 0.01 , K 2 = 0.03 , and L represents the dynamic range of the pixel values.
Furthermore, we optimized the defect generation framework to enhance computational efficiency:
Batch Range Control: Supports batch processing within user-specified indices.
Quantity Control: Accommodates multi-sample generation requirements for data augmentation scenarios.
Strength Regulation: A strength parameter modulates noise injection to balance structural retention with stylistic variation, fine-tuned based on SSIM feedback.
A dual-loop architecture designated as “Original Image Traversal → Single-Image Multi-Sample Generation” was designed:
Outer Loop: Traverses all original images within the control range, employing dynamic path construction to support the batch loading of sequentially named files;
Inner Loop: Generates a set number of samples for each image. The seed parameter controls diversity, ensuring that distinct noise initializations produce differentiated results for the same base image.
Standard implementations often repeatedly initialize the model and tokenizer within the generation function, resulting in redundant resource loading and increased latency. The optimized model utilizes a modular parameter-passing mechanism to achieve “write-once initialization, multiple invocations,” significantly reducing redundant computation. This improvement yields substantial time efficiency gains, particularly in large-scale batch processing (e.g., generating 5880 images from 196 originals). Additionally, multi-level error handling and preprocessing mechanisms were integrated to significantly enhance system robustness.

3.2. Online Inspection

The glass defect detection task requires simultaneous object localization, classification, and pixel-level segmentation. To meet this requirement, the online detection phase is implemented using the MASK R-CNN model.
MASK R-CNN overcomes technical bottlenecks in object detection and instance segmentation through its multi-task architecture, RoIAlign technology [21], flexible backbone network adaptability, and extensive application scalability, making it a milestone model in the field of computer vision.
Baseline Model Selection: We employ a Mask R-CNN architecture initialized with a ResNet-50-FPN backbone pre-trained on the COCO dataset. This transfer learning strategy is utilized to significantly accelerate model convergence. To address the challenge of annotating large-scale datasets, we implement a semi-supervised learning framework incorporating a pseudo-label generation mechanism. This approach mitigates the reliance on fully annotated data. The specific workflow for generating these pseudo-labels is illustrated in Figure 9.
Figure 9. Pseudo-label generation process.
Specifically, the Teacher model adopts a ResNet-50-FPN as the backbone network and is initialized with pre-trained weights from the COCO dataset. Before generating pseudo-labels, the Teacher model was first fine-tuned on a small number of real annotated samples (100 images). We specify the training hyperparameters as follows: an SGD optimizer was used with an initial learning rate of 0.002, a momentum of 0.9, and a weight decay of 0.0001.
To enhance the generalization ability of Mask R-CNN, a hybrid training strategy of “generated samples + real samples” is employed:
Dataset partitioning: The generated sample library is split into an 80:20 ratio for the generated training and validation sets. It is noted that given the homogeneity of the glass background and the randomized nature of the generated defects, this random split ensures consistent background distribution across sets without compromising detection rigor. A small number of real annotated samples (accounting for 10% to 20% of the total training samples) are introduced and combined with the generated training set to form the final training dataset. The real samples help mitigate the “domain shift between generated samples and real-world scenarios.”
Data balancing: To address the issue of class imbalance in industrial defect scenarios (e.g., many “scratch” samples but few “dent” samples), generated samples are used to supplement the underrepresented categories. The proportion of samples in each defect category is controlled within a 1:1.5 ratio, preventing the model from being biased towards the majority class.

4. Experiment

4.1. System Architecture Setup

To facilitate the complete workflow of defect data acquisition, generation, and identification, we designed a comprehensive detection system. As illustrated in Figure 10, the hardware architecture consists of three core subsystems: a mechanical conveyor assembly, an image acquisition module, and an industrial control unit (IPC).
Figure 10. Experimental platform.
To ensure optical robustness across different production lines and mitigate the impact of fluctuating ambient light, a custom-designed light shielding hood was integrated into the image acquisition module. This structure physically isolates the camera and light source from the external environment, creating a standardized optical domain. Consequently, the system maintains consistent imaging conditions regardless of the specific illumination layout of the factory floor, thereby minimizing the domain shift caused by environmental changes.
To facilitate the replication of our experiments and benchmark the inference speed, we specify the computational environment of the Industrial Control Unit (IPC) as follows: The system is powered by a 13th Gen Intel(R) Core(TM) i7-13700H CPU @ 2400 MHz and 16 GB of DDR5 RAM, equipped with an NVIDIA GeForce RTX 4060 GPU (8 GB) for accelerated processing. The software stack runs on the Windows 11 operating system with CUDA 12.7, and the deep learning models were implemented using the TensorFlow framework.
To ensure the reproducibility of this study and further clarify the specific experimental configuration, we have organized a detailed dataset description and hardware specifications, as presented in Table 1. This table summarizes key information ranging from image acquisition hardware (including camera resolution and custom lighting) to dataset partitioning (specific sample counts for training and test sets). In particular, we employed a combined illumination scheme of coaxial light and backlight within a custom shielding hood to eliminate environmental interference and enhance the contrast of minute defects. Furthermore, all raw images underwent unified preprocessing to align with the model input requirements.
Table 1. Detailed Configuration of Data Acquisition and Dataset Distribution.

4.2. Offline Generation

During the offline sample generation phase, we introduced the Structural Similarity Index Measure (SSIM) as a novel metric to assess the validity of the synthetic samples. This evaluation was applied across eight distinct defect categories, including edge chipping. The quantitative results are presented in Figure 11.
Figure 11. SSIM comparison results (The x-axis represents the serial numbers of the images randomly sampled from 10 samples).
The comparative analysis demonstrates that optimizing the Stable Diffusion parameters enabled the generated images to retain the primary characteristics of the original defects while introducing sufficient variance to ensure dataset diversity. Crucially, the synthetic images preserved the core morphological features of the original samples, achieving a mean similarity score exceeding 0.91. This confirms that the generated data meets the fundamental requirements for feature integrity in defect detection tasks.
To further validate the robustness of the selected SSIM threshold (0.9), we conducted a sensitivity analysis comparing different threshold settings. As shown in Table 2, a lower threshold (0.85) resulted in noticeable background artifacts, whereas a higher threshold (0.95) significantly constrained the morphological diversity of the defects. The threshold of 0.9 achieved the optimal balance between image fidelity and detection performance.
Table 2. Sensitivity analysis of SSIM thresholds on generation quality and detection performance.

Computational Cost and Feasibility Analysis

To evaluate the deployment cost of this method in industrial scenarios, we conducted a quantitative analysis of the computational overhead. All sample generation tasks were performed on a standard industrial control computer equipped with an Intel Core i7-13700H CPU and an NVIDIA GeForce RTX 4060 GPU (8 GB).
Although the diffusion-based generation process is computationally intensive, we optimized the inference pipeline using a “write-once initialization, multiple invocations” batching strategy, which significantly reduced VRAM usage and I/O latency. Crucially, this generation process is designed strictly as an offline operation. In practical industrial applications, sample augmentation needs to be executed only once before deployment or during overnight downtime. Once the hybrid dataset is constructed, the online phase runs only the lightweight Mask R-CNN model. Consequently, the computational cost of the generation phase does not translate into online detection latency, ensuring that the system maintains a high-speed real-time detection of 145 ms, fully meeting the takt time requirements of the industrial production line.

4.3. Online Inspection

It is worth noting that while GAN-based methods are common in data augmentation, we excluded them as a baseline in this study. This is because GANs typically struggle with training instability and mode collapse in extremely few-shot scenarios (e.g., <100 total samples), whereas the pre-trained Stable Diffusion model offers superior robustness and diversity for the specific industrial constraints addressed here.
In light of the three main challenges in the industrial inspection field (sample scarcity, limitations of traditional data augmentation, and resource waste from model coupling), the experimental design consists of three comparison schemes to validate the proposed method. The scheme design is as follows:
Scheme 1 (Only Real Samples): Train Mask R-CNN using only 100 real samples.
Scheme 2 (Real Samples + Traditional Data Augmentation): Train Mask R-CNN using 100 real samples and 1000 generated samples (Note: Preliminary experiments showed that increasing traditional augmentation beyond this quantity led to overfitting and performance saturation due to the limited diversity of the initial few-shot samples.) from traditional data augmentation techniques (e.g., flipping, cropping).
Scheme 3 (Proposed Method: Real Samples + Generated Samples): Train Mask R-CNN using 100 real samples and 4000 generated samples, with the sample generation and detection phases executed independently.
Scheme 4 (Control group): A coupled model without resource isolation.
Evaluation metrics: mAP50, Recall, Mask IoU, and Detection of inference time. The comparison results are shown in Table 3.
Table 3. Comparison of Overall Evaluation Indicators.
As detailed in Table 3, the inference latencies for the first three experimental schemes are consistent, averaging approximately 145 ms. This performance represents a substantial improvement compared to the control group, thereby validating the necessity of the resource isolation strategy in mitigating the computational overhead associated with model coupling. Notably, Scheme 3 demonstrates a marked performance gain over Scheme 1, increasing mAP50 by nearly 47 percentage points. This improvement indicates that the generative model produces samples with superior semantic diversity, effectively capturing complex texture and background variations that traditional geometric transformations fail to simulate.
To validate the necessity of the “Resource Isolation Strategy,” we constructed a coupled inference architecture for Scheme 4 (Control Group). In this setup, the Stable Diffusion generative model and the Mask R-CNN detection model are simultaneously loaded and integrated within the same online inference pipeline, simulating a scenario of “on-the-fly data augmentation” or “generative reconstruction-based detection.”
This architecture forces the generation and detection tasks to synchronously compete for GPU computational resources (VRAM and CUDA cores). As shown in Table 2, this resource contention causes the inference latency to surge from 145 ms to 300 ms, significantly impeding real-time performance on the production line. This counter-validates that completely offloading the generation process to the offline phase (i.e., the Resource Isolation proposed in this study) is a critical prerequisite for ensuring industrial-grade detection efficiency.
The comparative results in Table 4 demonstrate the significant improvement achieved by the proposed method in detecting “subtle defects.” For defects characterized by a low signal-to-noise ratio (SNR), such as scratches and watermarks, Scheme 1 proved largely ineffective (mAP < 20%). This failure is primarily attributed to the scarcity of real training samples, which hinders the Mask R-CNN from extracting discriminative features against the complex glass background. Traditional augmentation techniques merely rotate existing defect geometries, causing the network to overfit to specific, memorized morphologies. In contrast, the proposed method generates synthetic samples containing previously unseen defect morphologies, thereby significantly enhancing the model’s generalization capabilities. By specifically generating a vast array of scratch and watermark samples under varying illumination conditions, the model is compelled to learn robust textural representations rather than relying on simple geometric shapes. Consequently, these categories exhibited the most substantial performance gain, approaching 60%.
Table 4. Comparison of mAP50 for Eight Defects.
As indicated by the comparative results in Table 5, the proposed method enhances segmentation performance across three primary dimensions:
Table 5. Comparison of Segmentation Performance for Eight Types of Defects.
  • Breakthroughs in Segmenting Complex Defects
Scratches:
Limitations of Scheme 1: Due to the scarcity of real training samples, the Region Proposal Network (RPN) anchors often capture only fragmented segments of scratches. Consequently, the resulting segmentation masks appear as discontinuous, point-like artifacts rather than coherent lines;
Advantages of Scheme 3: The generative model synthesizes scratches with varying orientations and lengths. This forces the mask branch to learn prior knowledge regarding “slender continuity,” enabling the prediction of masks that completely cover the entire scratch.
Watermarks:
Limitations of Scheme 1: Watermarks exhibit minimal pixel intensity differences compared to the glass background, making boundary delineation difficult. Scheme 1 tends to either under-segment (capturing only the distinct center) or over-segment (spilling into the background), resulting in low Intersection over Union (IoU) scores;
Advantages of Scheme 3: By training on a vast array of synthetic watermarks under diverse lighting conditions, the model learns to focus on high-frequency textural variations rather than simple color discrepancies. This allows for the precise delineation of blurred watermark edges.
2.
Adaptability to Morphologically Variable Defects
Edge Chipping:
Edge chipping typically occurs at the glass periphery. Traditional flipping augmentation fails to alter the background context, leading the model to misclassify background features as defects. Scheme 3 generates samples with simulated, diverse edge backgrounds, enabling the model to focus specifically on the jagged chipping patterns, thereby raising the IoU to 82.1%;
Oxidation and Film Removal:
These defects typically manifest as patchy areas with non-uniform internal textures. Scheme 1 often results in a “hollowing” phenomenon, where the central portion of the defect is missed. The extensive dataset provided by Scheme 3 enables the model to understand the internal consistency of these defects, effectively filling voids within the predicted masks.
3.
Steady Improvements for Simple Defects
For black spots and bubbles, Scheme 1 achieves acceptable baseline performance due to their distinct features (high contrast and closed shapes). The improvements in Scheme 3 are primarily observed in boundary smoothness; the predicted masks adhere more closely to the true contours, significantly reducing edge aliasing.

Quantification of Annotation Cost Reduction

To explicitly evaluate the industrial viability, we quantified the reduction in manual annotation effort. In our semi-supervised framework, only 100 real samples per defect category were manually annotated, totaling 800 images. In contrast, training a fully supervised Mask R-CNN with comparable performance typically requires at least 500–1000 annotated samples per category to prevent overfitting. Consequently, our method reduces the manual annotation workload by approximately 80% to 90%. The remaining training data (generated samples) are utilized via the pseudo-labeling mechanism, incurring zero additional manual cost.

4.4. Quantitative Error Analysis and Robustness

To rigorously evaluate the model’s reliability in industrial scenarios and address concerns regarding experimental variance, we conducted a quantitative error analysis on a large-scale test set comprising 6560 images (820 per class, consisting of a 20% split of real and generated samples). Figure 12 illustrates the confusion matrix for the proposed method (Scheme 3) across eight defect categories.
Figure 12. Confusion matrix of the proposed method (Scheme 3) on the test set (820 samples per class). The values on the diagonal represent the number of correctly classified samples (True Positives), while off-diagonal values indicate misclassifications.
Key insights from the confusion matrix include:
  • High Reliability for High-Contrast Defects: For defects with distinct morphological features such as “Black dot” and “Edge Chipping,” the model achieved exceptionally high recall rates (95.5% and 91.6%, respectively) with minimal misclassification. This confirms that the Mask R-CNN, augmented by Stable Diffusion, effectively captures clear geometric features.
  • Physical Confusion in Low-Contrast Defects: “Watermark” exhibited the lowest recall (73.9%), with the primary error source being misclassification as “Dirty” (128 samples). This is consistent with physical expectations, as light watermarks and oil stains exhibit extremely similar grayscale values in industrial imaging, representing “hard samples” that are difficult to distinguish optically.
  • Class Interference due to Similar Textures: A degree of mutual confusion was observed between “Oxidation” and “Film Removal” (85 and 103 misclassified samples, respectively). This is primarily because both manifest visually as irregular, patchy discoloration. Nevertheless, thanks to the diverse training provided by the generated samples, the overall recognition rates for both categories remained above 80%, significantly outperforming the baseline.
Overall, the high values along the diagonal of the confusion matrix confirm the robustness of the proposed method in handling complex industrial defects, while the primary off-diagonal errors accurately reflect the physical challenges of the domain rather than systemic model failures.

4.5. Ablation Study

To rigorously evaluate the individual contributions of the core modules within the proposed framework, we conducted a series of ablation studies. Specifically, based on the full proposed method (Scheme 3), we constructed three model variants by sequentially removing or replacing key components to analyze the impact of the Mask Encoder, SSIM filtering strategy, and Resource Isolation mechanism on detection accuracy (mAP50) and inference efficiency (Inference Time). The detailed comparison results are presented in Table 6.
Table 6. Ablation study of key components on detection performance and efficiency.
Impact of Mask Encoder: When the Mask Encoder was removed and the model reverted to standard text-guided generation (using only the CLIP text encoder), a significant drop in mAP50 (approximately 12%) was observed. This is primarily because linguistic descriptions lack the granularity required to precisely describe the complex morphology of industrial defects (such as the specific orientation and curvature of scratches). Consequently, while the generated samples were semantically correct, they failed to effectively augment the training data in terms of spatial features.
Impact of SSIM Filtering Strategy: Removing the SSIM quality control step (i.e., retaining all generated samples for training) resulted in a decrease of approximately 5% in mAP50. This result indicates that unfiltered low-quality samples (SSIM < 0.9) often contain background artifacts or structural distortions. These noisy data interfere with the convergence of the Mask R-CNN feature extraction network, thereby impairing the final detection performance.
Impact of Resource Isolation Mechanism: To verify the necessity of resource isolation, we coupled the generation module during the online detection phase (similar to the Control Group in Table 1). The results showed that while this modification had a negligible effect on mAP, the inference time per image surged from 145 ms to 300 ms. This confirms that decoupling the computationally intensive generation process from the real-time detection process is critical for achieving industrial-grade deployment efficiency.
In summary, the Mask Encoder provides pixel-level feature guidance, establishing the system’s detection accuracy; the SSIM filtering guarantees the quality of data augmentation; and the resource isolation mechanism resolves the bottleneck of computational efficiency. The synergistic operation of these three components enables the system to achieve an optimal balance between accuracy and speed.

4.6. Generalization Discussion

Although the experimental validation in this study is primarily based on tempered glass datasets, the proposed “Mask Encoder + Stable Diffusion” framework is inherently material-agnostic. Tempered glass, due to its transparency and high reflectivity, is often regarded as a “hard case” in industrial vision inspection. The superior performance demonstrated by our method, particularly on low-contrast defects (e.g., watermarks and shallow scratches), suggests robust adaptability to complex optical environments. Algorithmically, the Mask Encoder effectively decouples defect “morphology” from “texture.” This implies that for cracks on metal surfaces or holes in textiles, the framework can be transferred by simply adjusting the text prompts and providing a small set of reference images. Consequently, we posit that this method can be effectively extended to other inspection scenarios, such as metal components and injection-molded products. Verification on multi-domain datasets will be a key focus of our future work.

Limitations

Despite the significant advantages demonstrated in few-shot defect detection, this study has certain limitations. First, regarding computational efficiency, while the resource isolation strategy ensures real-time online detection, the offline training and sample generation processes based on Stable Diffusion remain computationally intensive, incurring a high one-time temporal cost. Second, regarding physical detection limits, as shown in the error analysis in Section 4.4, the model still exhibits confusion between low-contrast defects such as “Watermarks” and “Dirty” spots. This indicates that when the optical characteristics (grayscale values) of defects highly overlap with the background or noise, generative data augmentation alone cannot fully overcome the physical signal-to-noise ratio bottlenecks of the imaging system.

5. Conclusions

The proposed “Stable Diffusion-Mask R-CNN Phased Collaboration” method efficiently coordinates sample generation and defect detection by clearly delineating the boundaries between these tasks, facilitating effective collaboration between sample augmentation and model training/inference. The core value of this approach lies in addressing the issue of sample scarcity in industrial defect detection, eliminating the resource waste caused by dual-model coupling, and improving the generalization ability of the detection model through the use of a mixed training dataset. While the method realizes a ‘low-cost, high-efficiency, high-accuracy’ paradigm for industrial deployment, it is worth acknowledging that the offline training of the Stable Diffusion model entails significant computational overhead. However, thanks to the proposed resource isolation strategy, this is a one-time investment that does not burden the online inference phase, thereby preserving the system’s real-time operational efficiency, which is expected to be widely applicable in industries such as automotive manufacturing, electronic component quality inspection, and aerospace parts testing.

Author Contributions

Conceptualization, K.H.; methodology, J.-F.Y., P.Z. and F.W.; software, K.H.; validation, K.H.; formal analysis, R.-Z.F. and X.-F.L.; investigation, K.H.; resources, J.-F.Y. and P.Z.; data curation, P.Z.; writing—original draft preparation, K.H.; writing—review and editing, J.-F.Y., P.Z. and F.W.; visualization, R.-Z.F. and X.-F.L.; supervision, G.-C.X. and F.W.; project administration, G.-C.X.; funding acquisition, J.-F.Y. and F.W. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Key Research and Development Program of Shandong Province (project numbers: 2024CXGC010208, 2024TZXD065), the Key Research and Development Project of Rizhao City (project number: 2025ZDYF0105), and the Major Innovation Project of Qilu University of Technology (Shandong Academy of Sciences) (project number: 2025ZDZX03).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The data that support the findings of this study are subject to an embargo due to commercial restrictions.

Acknowledgments

We declare that none of the work contained in this manuscript is published in any language or currently under consideration at any other journal, and there are no conflicts of interest to declare. Our manuscript has also been edited by a native English-speaking expert to ensure its English is good enough for publication.

Conflicts of Interest

Authors Run-Ze Fan and Xiang-Feng Liu were employed by the Taian City Taishan Huijin Intelligent Technology Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Li, F.; Fergus, R.; Perona, P. One-shot learning of object categories. IEEE Trans. Pattern Anal. Mach. Intell. 2006, 28, 594–611. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Yanyan, S.; Dianxi, S.; Ziteng, Q.; Yi, Z.; Yangyang, L.; Shaowu, Y. A Survey on Recent Advances in Few-Shot Object Detection. Chin. J. Comput. 2023, 46, 1753–1780. [Google Scholar]
  3. Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the Advances in Neural Information Processing Systems 27 (NIPS 2014), Montreal, QC, Canada, 8–13 December 2014; Volume 27, pp. 2672–2680. [Google Scholar]
  4. Zhao, C.; Xue, W.; Fu, W.-P.; Li, Z.-Q.; Fang, X. Defect Sample Image Generation Method Based on GANs in Diamond Tool Defect Detection. IEEE Trans. Instrum. Meas. 2023, 72, 2519009. [Google Scholar] [CrossRef] [Scilit]
  5. Xu, C.; Li, W.; Cui, X.; Wang, Z.; Zheng, F.; Zhang, X.; Chen, B. Scarcity-GAN: Scarce data augmentation for defect detection via generative adversarial nets. Neurocomputing 2024, 566, 127061. [Google Scholar] [CrossRef] [Scilit]
  6. Guo, Y.; Zhong, L.; Qiu, Y.; Wang, H.; Gao, F.; Wen, Z.; Zhan, C. Using ISU-GAN for unsupervised small sample defect detection. Sci. Rep. 2022, 12, 11604. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Finn, C.; Abbeel, P.; Levine, S. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 1126–1135. [Google Scholar]
  8. Dong, C.; Zhang, K.; Xie, Z.; Wang, J.; Guo, X.; Shi, C.; Xiao, Y. Transmission Line Key Components and Defects Detection Based on Meta-Learning. IEEE Trans. Instrum. Meas. 2024, 73, 5022213. [Google Scholar] [CrossRef] [Scilit]
  9. Xie, T.; Huang, X.; Choi, S.-K. Metric-based Meta-Learning for Cross-Domain Few-Shot Identification of Welding Defect. J. Comput. Inf. Sci. Eng. 2023, 23, 030902. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, S.; Chen, H.; Liu, K.; Zhou, Y.; Feng, H. Meta-FSDet: A meta-learning based detector for few-shot defects of photovoltaic modules. J. Intell. Manuf. 2022, 34, 3413–3427. [Google Scholar] [CrossRef] [Scilit]
  11. Qiao, L.; Zhang, Y.; Wang, Q. Fault detection in wind turbine generators using a meta-learning-based convolutional neural network. Mech. Syst. Signal Process. 2023, 200, 110528. [Google Scholar] [CrossRef] [Scilit]
  12. Pan, S.J.; Fellow, Q.Y. A Survey on Transfer Learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, J.; Guo, F.; Gao, H.; Li, M.; Zhang, Y.; Zhou, H. Defect detection of injection molding products on small datasets using transfer learning. J. Manuf. Process. 2021, 70, 400–413. [Google Scholar] [CrossRef] [Scilit]
  14. Mishra, G.; Gupta, P.; Tanwar, R. Target Recognition Using Pre-Trained Convolutional Neural Networks and Transfer Learning. Procedia Comput. Sci. 2024, 235, 1445–1454. [Google Scholar] [CrossRef] [Scilit]
  15. Zhu, Q.-X.; Zhang, H.-T.; Tian, Y.; Zhang, N.; Xu, Y.; He, Y.-L. Co-training based virtual sample generation for solving the small sample size problem in process industry. ISA Trans. 2023, 134, 290–301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Ruan, J.; He, J.; Tong, Y.; Wang, Y.; Fang, Y.; Qu, L. Knowledge Embedding Relation Network for Small Data Defect Detection. Appl. Sci. 2024, 14, 7922. [Google Scholar] [CrossRef] [Scilit]
  17. Lin, X.; Li, Z.; Liu, L.; Wu, J.; Zhang, L.; Zhou, X.-D. Irecut+MM: Data Generalization and Metric Improvement for Few-shot Learning. In Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME), Brisbane, Australia, 10–14 July 2023; pp. 2915–2920. [Google Scholar] [CrossRef] [Scilit]
  18. Chen, Z.; Mao, Y.; Qian, Y.; Pan, Z.; Xu, S. FRDet: Few-shot object detection via feature reconstruction. IET Image Process. 2023, 17, 3599–3615. [Google Scholar] [CrossRef] [Scilit]
  19. Ouyang, Y.; Wang, X.-Q.; Hu, R.-Z.; Xu, H.-H. Few-shot object detection based on positive-sample improvement. Def. Technol. 2023, 28, 74–86. [Google Scholar] [CrossRef] [Scilit]
  20. Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. NeurIPS 2020, 33, 6840–6851. [Google Scholar] [CrossRef] [Scilit]
  21. He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.