1. Introduction
Wood, as a renewable natural material, is widely used in fields such as building structures, furniture manufacturing, and interior decoration [
1]. Glued Laminated Timber (GLT), formed by bonding together multiple layers of sawn timber to create large-scale structural components [
2], effectively overcomes the limitations of natural wood’s size constraints and discrete mechanical properties. It has become a core load-bearing material in modern timber structures [
3]. However, surface defects such as knots, cracks, and mineral streaks, which occur during wood growth and processing, can significantly reduce key mechanical properties like bending strength and modulus of elasticity of the GLT, directly impacting its quality grade and safety for engineering applications [
4]. Therefore, precise detection and segmentation of wood surface defects is a core prerequisite for GLT quality grading.
According to the national standard GB/T26899-2022, “Structural glued laminated timber” [
5], the core criteria for GLT quality grading include the diameter and area of individual defects, the concentrated knot diameter ratio (the ratio of defect diameter to board width), and the number and density of defects per unit area. All these indicators are based on pixel-level precise segmentation results of the defects, rather than merely relying on rectangular bounding boxes output by detection stages. Especially for small-sized defects (diameter < 3 mm), bounding box regression suffers from significant localization errors, making it difficult to accurately reflect the true size and morphology of the defects [
6]. In GLT production practice, defects are categorized into three types by size: large defects (diameter ≥ 10 mm), medium defects (3 mm ≤ diameter < 10 mm), and small-sized defects (diameter < 3 mm) [
7]. Small-sized defects mainly include fine cracks (width < 2 mm), small knots (diameter < 3 mm), and fine mineral streaks (width < 3 mm). Although individually small, when their density exceeds 3 per 100 cm
2, they can significantly lower the GLT grade [
8]. Existing detection methods have achieved a mean Average Precision (mAP) of over 85% for large-sized wood defects, mainly including one-stage detectors such as YOLOv3, YOLOv5, and SSD [
9,
10,
11], as well as two-stage detectors such as Faster R-CNN [
10]. However, the detection accuracy for small-sized defects with diameter < 3 mm is only about 70–75% [
9], which becomes the main bottleneck restricting the overall performance of the grading system.
Semantic segmentation technology, by classifying each pixel in an image, can output the precise boundary contours of defects, providing a basis for defect size quantification. In recent years, semantic segmentation models have developed rapidly: the Fully Convolutional Network (FCN) [
12] first introduced a fully convolutional structure to semantic segmentation, overcoming the size limitations of traditional classification models. The encoder–decoder architecture and skip connections proposed by U-Net [
13] have become classic paradigms in medical and industrial defect segmentation. The DeepLab series [
14] expands the receptive field using atrous convolutions, effectively capturing multi-scale contextual information. Transformer-based segmentation models, with their global modeling capabilities, have become a current research hotspot: SegFormer [
15] employs a lightweight Transformer encoder and hierarchical feature fusion strategy, achieving excellent performance on multiple benchmark datasets. Mask2Former [
16], with its mask classification paradigm, unifies semantic and instance segmentation. Unlike traditional pixel-wise classification methods (e.g., FCN [
12]), Mask2Former formulates segmentation as a set of binary mask classifications. Its masked attention mechanism restricts each query’s focus to its corresponding mask region, and the Transformer architecture provides superior global modeling capability, leading to significantly better segmentation accuracy in general scenarios than traditional models. However, most of these general semantic segmentation models are designed for general scenes (e.g., medical images, urban scenes) and lack targeted optimization for wood small-sized defects. Specifically, they fail to address the core issues of weak feature response of small defects, severe interference from wood growth ring textures, and inaccurate boundary localization required for industrial measurement, leading to low segmentation accuracy and high miss rates in wood defect detection [
10,
11].
To address the issue of weak feature responses for small-sized objects, researchers have proposed various feature enhancement strategies. The Feature Pyramid Network (FPN) proposed by Lin et al. [
17] constructs a multi-scale feature pyramid through top-down feature propagation and lateral connections, increasing the AP for small object detection by about 8 percentage points on the COCO dataset. The Path Aggregation Network (PANet) by Liu et al. [
18] adds a bottom-up path aggregation on top of FPN, achieving bidirectional fusion of multi-scale features. The High-Resolution Network (HRNet) proposed by Wang et al. [
19] maintains high-resolution feature representations throughout the network, effectively avoiding the loss of spatial information for small objects. In boundary optimization, the boundary-aware loss function proposed by Ke et al. [
20] guides the network to focus on boundary pixels by increasing the gradient weight in boundary regions, significantly improving the IoU of segmentation boundaries. OCRNet by Yuan et al. [
21] enhanced feature response in boundary regions through object-contextual representations. Li et al. [
22] proposed the Semantic Flow Network (SFNet), which optimizes boundary localization accuracy through feature alignment. In the application domain of industrial surface defect segmentation, Tian et al. [
23] systematically reviewed key technologies for small-sized industrial defect detection and segmentation, summarizing three core strategies: multi-scale feature fusion, high-resolution representation, and data augmentation. Tabernik et al. [
24] proposed a segmentation-based method for industrial surface defect detection, transforming the detection task into a pixel-set segmentation task, achieving an AP of 97.2% in steel surface defect detection. Zhang et al. [
25] designed a high-resolution representation network for fine-crack segmentation, achieving an F1-score of 86.2% for cracks with a width < 2 mm.
To address the aforementioned challenges, this paper proposes an improved Defect-Mask2Former model for wood small-sized defect segmentation, using Mask2Former as the baseline. Through the synergistic optimization of the AGPE and DBCC modules, precise segmentation of small-sized defects was achieved. To provide reliable location prior information for small-sized defects, we adopt the EFCW-YOLO model [
26], a lightweight wood defect detection model proposed in our previous work, which achieves an mAP50 of 90.26% for small-sized defects and provides high-confidence candidate boxes for subsequent segmentation. The core contributions of this paper are as follows:
(1) We propose an original Attention-Guided Pyramid Enhancement (AGPE) module: This module integrates location prior information from an upstream detector. Through a three-stage mechanism of “feature purification–cross-scale enhancement–small-sized anchoring”, it specifically suppressed wood texture interference and enhanced feature responses for small-sized defects, reducing the small-sized defect miss rate from 20.78% to 5.83%.
(2) We design a Defect Boundary Calibration and Correction (DBCC) module: This module features a dual-branch parallel architecture and a boundary-aware loss function. It optimizes boundary localization specifically to meet the ≤3% size measurement accuracy requirement of the GB/T26899-2022 standard, reducing the diameter measurement error from 8.31% to 2.86% and achieving a deep integration of segmentation results with industrial grading needs.
(3) We construct a dedicated dataset for wood small-sized defects, PlankDefSeg: A custom image acquisition device was designed to build a dataset of 3500 pixel-level annotated images covering five types of small-sized defects across 6 industrial wood species. Three targeted data augmentation strategies are also proposed to address the issues of sparse samples and class imbalance for small-sized defects.
(4) We realize an end-to-end technical closed-loop for GLT grading: The segmentation results of Defect-Mask2Former are integrated into the GLT grading process, forming a standardized “segmentation–measurement–grading” pipeline. The system achieved a grading accuracy of 94.3%, and its efficiency was 20 times higher than manual grading, providing a complete solution for industrial implementation.
2. Results
2.1. Training Process Analysis
Figure 1 showed the training process curves of the Defect-Mask2Former model.
Table 1 presented the loss and performance metrics at key epochs during training. The model achieved optimal performance at the 70th epoch, after which training was stopped. The overall training process was stable, with no obvious overfitting.
From the training process curves in
Figure 1 and
Table 1, the following conclusions could be drawn:
- (1)
Fast Model Convergence: In the first 30 epochs, the training loss rapidly decreased from 2.85 to 1.35, a drop of 52.6%. The small-sized mIoU increased from 42.30% to 68.50%, a gain of 26.2 percentage points, indicating that the model could quickly learn features of wood small-sized defects.
- (2)
Epochs 30–50 as Fine-Tuning Period: During this phase, the loss decreased more slowly. The small-sized mIoU steadily improved from 68.50% to 81.40%, and the boundary IoU increased from 64.70% to 78.60%. The model began to focus on fine-grained learning of defect boundaries.
- (3)
Convergence after Epoch 50: After 50 epochs, the loss remained largely stable. The small-sized mIoU and boundary IoU slowly increased, reaching optimal values at epoch 70 (small-sized mIoU = 85.30%, boundary IoU = 82.10%). Thereafter, slight fluctuations due to mild overfitting occurred, but overall performance remained stable.
- (4)
Strong Model Generalization: The gap between training and validation loss remained consistently small. At optimal performance, the validation loss (1.05) was about 2.19 times the training loss (0.48), indicating no significant overfitting and strong generalization ability, which was attributable to the targeted data augmentation strategies and the lightweight model design.
- (5)
The generalization gap between training and validation losses at optimal performance was 0.57, and the overfitting coefficient was 1.19. The validation loss remained stable throughout training without a subsequent increase, indicating no significant overfitting. This is attributable to the data augmentation strategies, lightweight module design, and early stopping mechanism.
Under the established experimental environment, the training time per epoch for the Defect-Mask2Former model was about 12 min. Training to optimal performance (70 epochs) took approximately 14 h, demonstrating good training efficiency for subsequent model tuning and improvement, which was consistent with findings in efficient training strategies.
2.2. Baseline Model Comparison
To validate the overall performance of the Defect-Mask2Former model, it was compared with four mainstream semantic segmentation models on the PlankDefSeg test set. These models were selected based on their representativeness (covering encoder–decoder, atrous convolution, and Transformer-based architectures), state-of-the-art performance, and applicability in industrial defect segmentation. All models were trained from scratch with the same hyperparameters to ensure fairness.
Table 2 presented the core evaluation metrics for each model. The segmentation results of the baseline models were shown in
Figure 2.
From the comparison results in
Figure 2, it could be seen that the proposed Defect-Mask2Former model achieved the best performance across all evaluation metrics, comprehensively surpassing existing mainstream semantic and instance segmentation models. The detailed analysis was as follows:
Segmentation Accuracy Advantage: Defect-Mask2Former achieved a small-sized mIoU of 85.34%, a 10.09 percentage point improvement over SegFormer (75.25%). The core reason was the location prior guidance and cross-scale feature enhancement of the AGPE module, which solved the problem of weak feature response for small defects. The boundary IoU reached 82.11%, a 17.82 percentage point improvement over Mask2Former, demonstrating the precise calibration effect of the DBCC module on boundary localization.
Quantification Accuracy Met Industrial Needs: The size measurement error of 2.86% strictly met the ≤3% requirement of the GB/T26899-2022 national standard. In contrast, the measurement errors of comparison models all exceeded 6% (U-Net reached 10.38%), making them unable to support the calculation of quantitative indicators for GLT grading. This was one of the core industrial values of the proposed model.
Real-Time Balance: Although the inference speed was slightly lower than Mask2Former (29.8 FPS), 27.6 FPS still met the high-speed detection requirements of the production line, indicating that the lightweight design effectively controlled computational overhead. The inference speed of Defect-Mask2Former was 27.6 FPS, calculated based on an average inference time of 36.2 ms per image (batch size = 1) on an NVIDIA RTX 4090 GPU. The time breakdown is approximately 4.5 ms for preprocessing, 29.8 ms for forward inference, and 1.9 ms for post-processing.
Cost-Effectiveness Advantage: The model achieved a significant improvement in accuracy with only a 2 M increase in parameters, rather than relying on parameter stacking. This validated the rationality and efficiency of the AGPE and DBCC module designs, reducing hardware costs for industrial deployment.
2.3. Ablation Study of AGPE and DBCC Modules
To clarify the individual effects and synergistic impact of the AGPE and DBCC modules, four controlled experiments were designed: (1) M1 (baseline group, without any modules); (2) M2 (AGPE module only); (3) M3 (DBCC module only); (4) M4 (dual-module synergy, our proposed model).
Table 3 showed the ablation study results, and
Figure 3 compared the segmentation results for different configurations.
The ablation study results showed:
When only the AGPE module was introduced, the small-sized mIoU increased from 67.50% to 82.54% (+15.04 percentage points), and the miss rate dropped from 20.78% to 8.28% (−12.5 percentage points). However, the boundary IoU only improved by 9.59 percentage points, and the diameter measurement error decreased to 5.62%. This validated that the AGPE module, through its “feature purification–cross-scale enhancement–small-sized anchoring” mechanism, effectively strengthened feature responses for small-sized defects and suppressed wood texture interference, significantly reducing the miss rate, but its effect on boundary optimization was limited.
When only the DBCC module was introduced, the boundary IoU increased from 64.29% to 79.45% (+15.16 percentage points), and the diameter measurement error dropped from 8.31% to 4.17% (−4.14 percentage points). However, the small-sized mIoU improved by only 8.92 percentage points, and the miss rate decreased by only 3.23 percentage points. This indicated that the DBCC module’s dual-branch architecture and boundary-aware loss function could precisely optimize segmentation boundaries and improved size measurement accuracy, but its improvement on miss problems was limited.
When both AGPE and DBCC modules were introduced simultaneously, all metrics reached their optimum: small-sized mIoU of 85.34%, miss rate of 5.83%, boundary IoU of 82.11%, and diameter measurement error of 2.86%. The two modules formed a functional complement: AGPE provided high-quality initial segmentation masks for DBCC (reducing misses, suppressing texture interference), and DBCC further calibrated the boundaries on this basis, achieving a “no-miss + high-precision” segmentation effect.
It was noteworthy that the inference speed of M4 (27.6 FPS) was slightly higher than that of M3 (28.2 FPS). This phenomenon stemmed from the feature filtering effect of the AGPE module—its Gaussian attention mask focused on defect regions, reducing the effective processing range for the DBCC module and partially offsetting the computational overhead introduced by the dual modules, ensuring the model still met the production line’s real-time requirement (≥25 FPS) while significantly improving accuracy.
2.4. Segmentation Performance for Different Defect Types
To further validate the model’s suitability for various types of small-sized defects, the performance of Defect-Mask2Former and the original Mask2Former on the five defect classes was compared. The results were shown in
Table 4.
An in-depth analysis of the performance improvements for each defect type revealed that Defect-Mask2Former demonstrated significant optimization effects across defects with different visual characteristics, although its mechanism of action varied slightly.
First, for fine mineral streaks, the most challenging defect type (contrast ratio with wood growth rings was only 1.3:1), the original model achieved an mIoU of only 55.81%. Defect-Mask2Former improved it to 81.17% (+25.36 percentage points). This benefited from the AGPE module’s Gaussian attention mask, which effectively suppressed interference from growth-ring texture, separating the feature response of fine mineral streaks from the background. The DBCC module further calibrated the boundaries, controlling the measurement error to 3.76%.
Second, for fine cracks (linear, width < 2 mm), the original model suffered from a high miss rate (24.67%) due to weak feature response. The cross-scale feature enhancement mechanism of the AGPE module had a targeted enhancement effect on linear small-sized features, increasing the mIoU by 21.24 percentage points, achieving a boundary IoU of 80.00% and a measurement error of 3.12%, meeting industrial grading requirements.
For high-contrast small live/dead knots (punctate, grayscale contrast > 3.0:1), the original model already possessed some segmentation ability (mIoU > 74%). Defect-Mask2Former, by quickly locking onto the target area with the AGPE module and precisely calibrating boundaries with the DBCC module, raised the mIoU to over 87%. The measurement error for small live knots was only 2.20%, the lowest among all defect types.
Finally, for small knot voids (round/elliptical holes, clear edges but small size), the local attention anchoring mechanism of the AGPE module could precisely focus on the hole area. The morphological post-processing of the DBCC module filled tiny gaps in the boundary, achieving an mIoU of 87.26% and a boundary IoU of 84.55%, with segmentation results close to those for small live/dead knots.
Overall, Defect-Mask2Former achieved significant performance improvements across all types of small-sized defects, with particularly outstanding optimization effects for complex defects such as low-contrast, linear, and texture-similar ones. This validated the model’s scene adaptability and robustness.
2.5. Visualization of Segmentation Results
To intuitively demonstrate the fine-grained impact of the AGPE and DBCC modules on the segmentation results, this section selected the most representative defect samples from the baseline comparison and performed a local magnification comparison of the segmentation boundaries between the baseline model (Mask2Former) and our proposed model (Defect-Mask2Former), as shown in
Figure 4.
For linear defects like fine cracks, the continuity of the segmentation was crucial. As seen in the first row of
Figure 4, the baseline model’s segmentation result was incomplete along the crack (indicated by the red arrow), which directly affected the accurate measurement of crack length. Defect-Mask2Former, benefiting from the gradient-guided fusion and morphological closing operation of the DBCC module, output a continuous and smooth crack profile, completely preserving the linear trajectory of the crack.
For defects with regular geometric shapes, such as small knots and knot voids, boundary adherence was key to evaluating segmentation quality. Rows two, three, and five of
Figure 4 showed the segmentation results of both models on circular knot voids. The baseline model’s segmentation boundaries showed enlargement (indicated by blue arrows) and jagged edges (indicated by green arrows), deviating from the true circular contour of the void. In contrast, Defect-Mask2Former segmented boundaries that were smooth and closely adhered to the edge of the knot void. This pixel-level boundary calibration capability was a direct visualization of the reduction in size measurement error from 8.31% to 2.86%.
Fine mineral streaks were easily confused with wood texture. The fourth row of
Figure 4 showed that while the baseline model detected the target, its mask range was incomplete (indicated by the white arrow). Defect-Mask2Former, by suppressing texture interference through the AGPE module, achieved a segmentation result with clearer boundaries, proving the effectiveness of the AGPE module in the feature purification stage.
The above visual comparison results indicated that the AGPE module, by suppressing texture interference and enhancing feature responses, provided a “cleaner” initial segmentation region for the DBCC module. The DBCC module, with its dual-branch architecture and boundary-aware loss, then converted this advantage into final pixel-level boundary accuracy, collectively achieving a “no-miss + high-precision” segmentation effect.
2.6. Robustness Verification in Extreme Scenarios
To simulate the complex production environment of a wood processing workshop, three types of extreme scenarios were designed to test model robustness: (1) Low-light scenario (brightness reduced to 50% of standard, intensity < 500 lux); (2) Dense distribution of small defects (single image containing > 5 small-sized defects, density > 8 per 100 cm
2); (3) High texture interference scenario (samples of Pinus radiata with growth ring density > 50 rings/cm
2).
Table 5 presented the robustness test results, and
Figure 5 showed the segmentation results in different scenarios.
Analysis of the experimental results showed:
Low-light scenario: Low light reduced the image signal-to-noise ratio, decreasing defect-to-background contrast. The original model achieved an mIoU of only 65.46%. Defect-Mask2Former’s AGPE module enhanced the feature signal-to-noise ratio through feature purification, and the DBCC module still had good calibration ability for low-contrast boundaries, achieving an mIoU of 82.15%, which met the practical needs of unstable workshop lighting.
Dense distribution of small defects: When a single image contained seven small-sized defects (density 8.75 per 100 cm2), defects were close together, and features easily overlapped. The original model’s segmentation accuracy decreased (mIoU = 68.92%) due to redundant search space. The location prior guidance of the AGPE module restricted the feature extraction range for each defect to within its candidate box, reducing interference between defects, resulting in a miss rate of only 7.24%, with only one tiny defect with extremely high overlap being slightly missed.
High texture interference scenario: For Pinus radiata with growth ring density > 50 rings/cm2, the background texture was dense and highly similar to the features of fine mineral streaks and fine cracks. Texture misclassification accounted for over 40% of errors in the original model (mIoU = 67.23%). The Gaussian attention mask of the AGPE module effectively suppressed the feature response of dense growth rings, allowing defect features to stand out. The DBCC module further distinguished defect boundaries from texture edges, ultimately achieving an mIoU of 80.57%, validating the model’s ability to suppress high texture interference.
The test results in the three extreme scenarios indicated that Defect-Mask2Former could adapt to the detection needs of complex production environments, providing reliability assurance for industrial deployment.
2.7. Validation in GLT Grading Application
The segmentation results of Defect-Mask2Former were integrated into the GLT grading process to build an end-to-end “segmentation–measurement–grading” system. Grading decisions strictly followed the GB/T26899-2022 national standard, with core grading indicators shown in
Table 6.
The test dataset contained 200 industrial-grade GLT timbers (covering grades I/II/III, with 60, 80, and 60 timbers per grade, respectively) (
Table 7). The automatic grading results were compared with manual grading results from three forestry experts to verify system performance.
Efficiency Improvement: Manual grading of a single board averaged 30 s (including defect identification, size measurement, index calculation, and grade determination). The automated system took only 1.5 s per timber (including image input, model inference, size calculation, and grade output), a 20-fold efficiency increase, meeting the high-speed detection demand of 60 timbers per minute.
Error Analysis: Among the 12 incorrectly graded timbers, seven were due to small-sized defect density being close to the grade threshold (e.g., borderline Grade II samples with density 7.8–8.2 per 100 cm2), three were due to the concentrated knot diameter ratio calculation being affected by defects at the board edge, and two were due to slightly exceeding the 3% measurement error for tiny defects under low-light conditions. Future work could improve accuracy by optimizing the adaptive adjustment mechanism of grading thresholds and enhancing edge defect handling logic.
The application validation results demonstrated that the segmentation results of Defect-Mask2Former could accurately support the calculation of GLT grading indicators. The end-to-end system achieved a grading accuracy of 94.3%, and its efficiency met industrial production requirements, providing a feasible technical solution for intelligent quality control of GLT.
3. Discussion
3.1. Analysis of the AGPE Module’s Mechanism
The AGPE module addressed the core problems of weak feature response and severe texture interference for wood small-sized defects through its three-stage synergistic mechanism: feature purification, cross-scale enhancement, and small-sized anchoring. The core of its mechanism lay in the deep integration of upstream detection location priors with multi-scale feature enhancement, achieving precise focusing and feature enhancement in defect regions. In the feature purification stage, the Gaussian soft mask generated based on the EFCW-YOLO candidate boxes did not perform hard thresholding on background regions. Instead, it suppressed texture interference through gradual weight distribution while retaining contextual information around the defect, increasing the signal-to-noise ratio of small-sized defect features by over 35% and effectively avoiding the boundary feature fracture caused by hard masks [
27]. The deformable convolution introduced in the cross-scale enhancement stage adaptively adjusted sampling points by learning the irregular morphology of wood defects, fitting the edge features of defects better than traditional convolution [
28]. Furthermore, the feature concatenation of P2 and P3 layers achieved complementarity between shallow details and deep semantics, compensating the insufficiency of single-scale features in representing small-sized defects.
The small-sized anchoring mechanism of local spatial attention was a key design of the AGPE module. The 3 × 3 local attention kernel precisely matched the pixel scale of wood small-sized defects (around 30 × 30 pixels). It dynamically generated attention weights within the candidate box region, enabling the network to focus on the defect core area rather than extracting features indiscriminately across the entire image, significantly reducing the search space of the segmentation network [
29]. Additionally, the AGPE module operated only on the P2 layer features output by the pixel decoder, without major modifications to the backbone network or Transformer decoder. The parameter increment was only 0.8 M, and the computation increment was 2.1 GFLOPs, achieving lightweight enhancement and laying the foundation for maintaining the model’s real-time performance.
It should be noted that the performance of the AGPE module depends on the localization accuracy of the upstream detection model. The EFCW-YOLO model used in this study achieved an mAP50 of 90.26% and a precision of 89.24% for small-sized defects, providing a reliable location prior for AGPE. If the detection model missed defects or had localization deviations, the AGPE module would not be able to effectively enhance the corresponding defects. Therefore, in practical industrial applications, expanding the candidate box boundary by 10% or introducing multi-scale detection candidate box fusion strategies could improve the AGPE module’s robustness to detection deviations. For extremely small defects (diameter < 1 mm) completely missed by the detection model, pixel saliency detection on high-resolution feature maps could be combined to supplement location prior information, further reducing the miss rate.
3.2. Boundary Optimization Effect of the DBCC Module
The dual-branch parallel architecture of the DBCC module ensured both the integrity of the defect region and the precision of boundary localization. Its core innovation lay in decoupling mask segmentation from boundary calibration and then adaptively fusing them, solving the problem of traditional single-branch models focusing insufficiently on boundaries when optimizing for mIoU [
30]. The mask branch retained the original segmentation logic of Mask2Former, ensuring the overall correct identification of the defect region and providing a high-quality base mask for boundary calibration. The boundary branch used lightweight depthwise separable convolution, reducing computation to 1/8 of traditional convolution while specifically learning defect boundary features. By using Dice loss [
31] to address the class imbalance caused by scarce boundary pixels, it significantly enhanced feature response in boundary regions.
The gradient-guided fusion mechanism was the core link of boundary calibration. Using the gradient of the ground truth boundary extracted by the Sobel operator as a supervisory signal, it achieved pixel-level adaptive weight distribution: in regions with gradient value > 0 (boundaries), a higher weight (0.7–0.9) was assigned to the boundary branch to strengthen boundary pixel calibration; in regions with gradient value = 0 (defect interior), a higher weight (0.8–0.95) was assigned to the mask branch to ensure segmentation stability. Experiments validated that this fusion mechanism reduced jagged edges by over 75% and improved the classification accuracy of boundary pixels from 78.2% to 91.5%.
The morphological post-processing layer, consisting of a 3 × 3 closing operation and area threshold filtering, completed the fine-grained optimization of boundaries. The dilation operation filled tiny gaps (≤1 pixel) in the boundary, preventing area underestimation during size measurement. The erosion operation smoothed boundary burrs, making the segmented contour better fit the true defect morphology. Simultaneously, isolated noise points with area < 10 pixels were filtered out, eliminating false boundaries caused by wood texture noise. It should be noted that the post-processing parameters needed to match the pixel resolution of the dataset. In this study, the area threshold of 10 pixels corresponds to a physical size of 0.01 cm2 under a pixel resolution of 0.1 mm/pixel, suitable for detecting defects with diameter ≥ 1 mm. If applied to a higher-resolution image acquisition system, the threshold parameters should be adjusted proportionally to ensure consistent optimization effectiveness.
The DBCC module reduced the average boundary deviation for wood small-sized defects to ≤0.5 pixels and the size measurement error from 8.31% to 2.86%. The core reason was that it directly linked the optimization objective from “improving boundary IoU” to “meeting industrial size measurement accuracy requirements.” Through the three-stage linkage of boundary-aware loss, gradient fusion, and morphological optimization, it achieved deep adaptation between segmentation boundaries and physical size quantification, which was the core advantage of this module compared to traditional boundary optimization methods.
3.3. Comparative Analysis with Existing Methods
3.3.1. Comparison with Traditional Semantic Segmentation Models
Compared with traditional encoder–decoder architecture models like U-Net, and DeepLabv3+, the mask classification paradigm adopted by Defect-Mask2Former, combined with the improvements from AGPE and DBCC modules, demonstrated two core advantages in wood small-sized defect segmentation. First was the flexibility of the object query mechanism: 100 learnable object queries could adaptively match different numbers and scales of defect instances in an image without presupposing an upper limit on defect count. Compared with the pixel-wise point classification of U-Net and DeepLabv3+, this was more suitable for scenarios where the number of wood defects varied randomly. Second was the scene specificity of feature enhancement and boundary calibration: traditional models used generic multi-scale feature fusion designs that did not account for the feature similarity between wood texture and defects. In contrast, the AGPE module precisely suppressed texture interference through location priors, and the DBCC module optimized boundaries for industrial measurement precision, enabling the model’s segmentation performance in wood scenes to far exceed that of general-purpose models.
From the performance data (
Table 2), Defect-Mask2Former’s small-sized mIoU was 19.27 percentage points higher than U-Net and 16.41 percentage points higher than DeepLabv3+. Its miss rate was 15–27 percentage points lower than those models, while its inference speed was only slightly lower than traditional lightweight models, achieving a balance between accuracy and speed. Although traditional models have advantages in inference speed, their segmentation accuracy and size measurement error were far from meeting the requirements of industrial GLT grading. Defect-Mask2Former, by adding only a small amount of computation, achieved a qualitative leap in performance, making it more suitable for industrial deployment.
3.3.2. Comparison with Mainstream Segmentation Models
While specialized instance segmentation models like SOLOv2 and QueryInst have shown promise in general object segmentation [
32,
33], they lack customized designs for wood defect scenarios, such as utilizing detection priors or industrial measurement-oriented boundary calibration. Consequently, without the targeted improvements of AGPE and DBCC, their performance on small-sized wood defects is expected to be significantly lower than Defect-Mask2Former, as evidenced by the substantial performance gap (over 10 percentage points in mIoU) observed between our model and other general-purpose baselines like Mask2Former and SegFormer in
Table 2.
3.4. Industrial Application Value
The core value of this study lies in constructing an end-to-end standardized technical process from “image acquisition–defect segmentation–size measurement–GLT grading,” achieving a deep integration between wood small-sized defect segmentation and industrial GLT grading. This solved three core problems faced by traditional manual grading and general detection models in industrial applications.
First was the low efficiency and inconsistent standards of manual grading: manual grading took about 30 s per timber and was subject to subjective factors, with grading consistency only around 80% [
34]. In contrast, the proposed automated system took only 1.5 s per timber, achieved a grading accuracy of 94.3%, and fully adhered to the GB/T26899-2022 national standard, ensuring objective and unified grading results.
Second was the high miss rate and large measurement error of general models for small-sized defects: existing models generally had a miss rate >20% and a size measurement error >8%, failing to meet the precision requirements of industrial grading. Defect-Mask2Former reduced the miss rate to 5.83% and controlled the measurement error to 2.86%, fully meeting the ≤3% accuracy requirement of the national standard.
Third was the poor adaptability of detection systems to production lines: the image acquisition device designed in this study could accommodate GLT with widths of 50–600 mm and lengths of 100–4000 mm. The model’s inference speed reached 27.6 FPS, meeting the ≥25 FPS real-time detection requirement of the production line. The overall deployment cost of the device and model was low, making it easy to promote in small and medium-sized wood processing enterprises.
From an industry perspective, the AGPE and DBCC modules proposed in this study provided a reference method for industrial surface small-sized defect segmentation. Their design ideas—detection-segmentation cascade feature enhancement and boundary calibration oriented towards industrial accuracy—could be extended to surface defect detection in other industrial products such as steel, boards, and ceramics, offering a technical reference for the intelligent upgrade of industrial vision inspection. Furthermore, the PlankDefSeg dataset constructed in this study provided a dedicated data foundation for wood defect segmentation research, filling the gap in datasets for wood small-sized defect segmentation and promoting the application and development of computer vision technology in the wood processing industry.
3.5. Limitations and Future Work
Although this study achieved good results in wood small-sized defect segmentation and GLT grading applications, there are still limitations in the following five areas, which are also key directions for future research:
Scale and Diversity of the Dataset Need Expansion: The PlankDefSeg dataset constructed in this study covers six industrial wood species and five types of small-sized defects, totaling 3500 pixel-level annotated images. While sufficient for model training and validation, the coverage of wood species and defect types was limited. It did not include other mainstream industrial woods like oak and walnut, nor common defects such as insect holes, dents, and resin pockets. The model’s generalization ability to unseen species and defects needed validation. Additionally, the data collection scene was primarily in a laboratory environment, which differs from the complex environments of wood processing workshops (e.g., dust, vibration, strong light reflection). Future work will involve collaboration across multiple regions and enterprises to collect defect images from over 20 mainstream industrial wood species worldwide, adding defect types like insect holes, dents, and resin pockets to construct a multi-scene wood defect dataset of ≥10,000 images. Federated learning methods will also be introduced to enable joint training on multi-source data while protecting enterprise data privacy, enhancing the model’s cross-scene and cross-species generalization.
Real-Time Performance and Edge Deployment Need Optimization: Defect-Mask2Former achieved an inference speed of 27.6 FPS on an NVIDIA RTX 4090 server, meeting basic real-time detection requirements. However, on edge computing devices (e.g., NVIDIA Jetson AGX Orin, NVIDIA Corporation, Santa Clara, California, USA; RK3588, Rockchip Electronics Co., Ltd., Shanghai, China), the inference speed dropped to 15–20 FPS, which could not adapt to the detection needs of high-speed production lines (>60 timbers/minute). Also, the model had 46 M parameters, posing some requirements on the memory resources of edge devices. Future work will employ lightweight methods such as model pruning, quantization, and knowledge distillation to design a lightweight version of Defect-Mask2Former (Defect-Mask2Former-Lite), reducing parameters to below 20 M and computation to below 100 GFLOPs, ensuring inference speed ≥30 FPS on edge devices [
35]. Model quantization techniques will also be used to convert the model from FP32 to FP16 or even INT8, further reducing device resource consumption and enabling lightweight edge deployment. For instance, INT8 quantization can theoretically increase inference speed by 2–3 times, and TensorRT optimization can further reduce latency, enabling deployment on edge devices while maintaining ≥30 FPS.
End-to-End Joint Training of Detection and Segmentation Not Yet Realized: In this study, the AGPE module relied on the independent output of the upstream EFCW-YOLO detection model. The detection and segmentation models were trained separately and used in a cascaded inference mode, without achieving end-to-end joint optimization. This means errors from the detection model could propagate to the segmentation model, and the overall training and inference efficiency of the combined system was lower. Future work will design an integrated detection–segmentation network architecture, fusing the detection branch of EFCW-YOLO and the segmentation branch of Defect-Mask2Former into a single end-to-end network. This would allow sharing of feature extraction layers in the backbone, enabling end-to-end learning of location prior information. A joint loss function will also be designed, combining detection localization loss with segmentation mask loss and boundary loss to achieve synergistic optimization of detection and segmentation, reducing error propagation and improving the overall performance and inference efficiency of the model.
2D Segmentation Cannot Detect Internal Wood Defects: The model in this study only processes 2D images of the wood surface and can only detect surface defects. It cannot detect internal defects such as internal cracks, voids, or pith, which are important factors affecting the mechanical properties of GLT and are key criteria in GLT quality grading [
36]. Future work will integrate non-destructive testing techniques like laser ultrasonic testing (LUT) and X-ray detection to fuse 2D image information from the wood surface with 3D structural information from the interior [
37]. This will involve building multimodal wood defect detection and segmentation models, introducing 3D convolutions (3D-CNN) and vision Transformers (ViT) to achieve integrated detection and segmentation of surface and internal defects. This would provide more comprehensive defect information for GLT quality grading, further improving grading accuracy and reliability [
38].
Size Measurement Accuracy Can Be Further Improved: Although the current size measurement error (SME) of 2.86% strictly meets the ≤3% accuracy requirement of the GB/T26899-2022 national standard, there is still room for further improvement to accommodate higher-precision grading scenarios. Future work will focus on two technical directions: (1) high-resolution imaging: upgrading the industrial camera from the current 3072 × 2048 resolution to over 6000 × 4000 pixels, which can provide richer spatial details and refine boundary localization, potentially reducing SME to below 2%; and (2) multi-view fusion and 3D reconstruction: employing multiple cameras to capture defect images from different angles, performing 3D reconstruction to eliminate perspective distortion from single-view imaging, thereby improving the accuracy of defect size quantification, especially for irregularly shaped defects.
Additionally, future research will explore few-shot learning methods, such as prototypical networks [
39] and meta-learning approaches [
40], to address the problem of insufficient samples for rare wood defects. These methods enable rapid adaptation to new defect categories with only a few labeled samples, reducing the dependency on large-scale annotated datasets. Adaptive illumination enhancement algorithms will be introduced to improve model robustness under complex lighting conditions in wood processing workshops. A visual industrial inspection software will be developed to achieve integrated management of image acquisition, defect detection, grade determination, and result traceability, further enhancing the technology’s industrial applicability and practicality.
5. Conclusions
Addressing the industrial need for precise segmentation of small-sized wood defects, this paper proposes an improved Defect-Mask2Former semantic segmentation model, based on the Mask2Former baseline, that integrates an Attention-Guided Pyramid Enhancement (AGPE) module and a Defect Boundary Calibration and Correction (DBCC) module. Through the synergistic optimization of these dual modules, the model solved core problems faced by existing methods in wood small-sized defect segmentation, such as weak feature response, severe texture interference, ambiguous boundary localization, and excessive measurement error. A dedicated image acquisition device was also designed, and the PlankDefSeg dataset for wood small-sized defect segmentation was constructed, realizing an end-to-end technical closed-loop from image acquisition to GLT grading. The main research conclusions were as follows:
The AGPE module effectively strengthened features of small-sized defects, significantly reducing the miss rate. Through its three-stage mechanism of “candidate-box-guided feature purification–cross-scale enhancement via deformable convolution–small-sized anchoring via local spatial attention,” the AGPE module introduced upstream detection location priors into the segmentation network. It precisely suppressed wood texture interference and strengthened feature representation for small-sized defects. This reduced the model’s miss rate for wood small-sized defects from 20.78% to 5.83% and increased the small-sized mIoU to 82.54%. The module added only 0.8 M parameters and 2.1 GFLOPs of computation, achieving lightweight feature enhancement.
The DBCC module achieved fine-grained boundary calibration, meeting industrial measurement accuracy requirements. Employing a dual-branch parallel architecture (mask branch and boundary branch), combined with a gradient-guided fusion mechanism and morphological post-processing, the DBCC module reduced the average boundary deviation to ≤0.5 pixels and the size measurement error from 8.31% to 2.86%. This strictly met the ≤3% accuracy requirement of the GB/T26899-2022 national standard, solving the problem of boundary localization being disconnected from industrial measurement needs in traditional models.
The Defect-Mask2Former model balanced accuracy and real-time performance. Experimental results on the PlankDefSeg dataset showed that Defect-Mask2Former achieved a small-sized defect mIoU of 85.34%, a 17.84 percentage point improvement over the original Mask2Former, and a 10–26 percentage point improvement over mainstream models like U-Net, DeepLabv3, and SegFormer. The model’s inference speed reached 27.6 FPS, meeting the ≥25 FPS real-time detection requirement for GLT production lines, making it a wood small-sized defect segmentation solution that balanced accuracy and speed.
The model exhibited strong robustness in extreme industrial scenarios. In three types of extreme scenarios—low light (50% brightness), dense distribution of small defects (density > 8 per 100 cm2), and high texture interference (growth ring density > 50 rings/cm2)—Defect-Mask2Former achieved small-sized mIoU of 82.15%, 81.82%, and 80.57%, respectively, representing improvements of 16–13 percentage points over the original Mask2Former. This indicated strong adaptability to complex scenes in wood processing workshops.
End-to-end intelligent GLT grading was achieved, with significant industrial application value. By integrating the segmentation results of Defect-Mask2Former into the GLT grading process, a standardized “segmentation-measurement-grading” closed-loop was formed. On a test set of 200 timbers, the automated grading results achieved 94.3% consistency with expert manual grading, with accuracies of 95.8% for grade I, 94.2% for grade II, and 93.1% for grade III. The grading time per timber was reduced from 30 s manually to 1.5 s, a 20-fold increase in efficiency, providing reliable technical support for intelligent quality control on GLT production lines.