3.2. Test Dataset-Based Model Evaluation and Inference Results
Figure 6 presents the test dataset-based evaluation results for YOLOv8n, YOLOv11n, and RT-DETR under the Original, AUG, ChatGPT-4.o, and ChatGPT-5.5 dataset conditions. The results were analyzed using Precision, Recall, mAP@0.5, and mAP@0.5:0.95 for Total, Smoke, and Flame classes. Overall, the generated-image-based training conditions produced competitive or improved results compared with the original and conventional augmentation conditions, although the degree of improvement varied depending on detector architecture, metric, class, and added-image scale. In the Precision results, YOLOv11n showed the most pronounced improvement under the ChatGPT-4.o condition. In particular, the ChatGPT-4.o 2400 condition produced the highest Precision values for YOLOv11n, reaching 0.740 for Total, 0.720 for Smoke, and 0.773 for Flame. This indicates that the generated images contributed to reducing false-positive detections while maintaining class-level detection reliability. RT-DETR also showed competitive Precision performance, with best values of 0.705 for Total, 0.644 for Smoke, and 0.766 for Flame. In the case of YOLOv8n, the Flame Precision reached 0.731 under the AUG condition, whereas the ChatGPT-5.5 3000 condition showed high Precision values for Total and Smoke. These results suggest that the generated-image effect was not identical across detector architectures, but it generally helped maintain or improve prediction reliability.
The Recall results showed a different tendency from Precision. Across all detector architectures, Recall values were lower than Precision values, indicating that missed detections remained a more challenging issue than false-positive suppression. The best Recall values for Flame were 0.683 for YOLOv8n, 0.682 for YOLOv11n, and 0.665 for RT-DETR, whereas the corresponding Smoke Recall values remained lower, at approximately 0.470, 0.478, and 0.467, respectively. This class-wise gap indicates that Smoke detection was more difficult than Flame detection, most likely because smoke regions have ambiguous boundaries, low contrast, and large variations in density and shape. Nevertheless, the ChatGPT-4.o-based conditions improved or maintained Recall in YOLOv8n and YOLOv11n, suggesting that prompt-generated smoke and flame patterns provided useful supplementary training diversity.
For mAP@0.5, the generated-image-based conditions also showed meaningful performance tendencies. YOLOv11n achieved the highest overall mAP@0.5 among the three detector architectures, with best values of 0.632 for Total and 0.719 for Flame. YOLOv8n achieved best values of 0.628 for Total and 0.708 for Flame, while RT-DETR reached 0.580 for Total and 0.692 for Flame. These results indicate that YOLOv11n benefited most from the expanded dataset in terms of overall detection accuracy. In contrast, Smoke mAP@0.5 remained lower than Flame mAP@0.5 across all detector architectures, confirming that smoke detection remained the limiting factor in the overall fire detection performance.
The mAP@0.5:0.95 results further demonstrate the difficulty of precise fire-object localization. Compared with mAP@0.5, the mAP@0.5:0.95 values were substantially lower for all detector architectures and dataset conditions. The best Total mAP@0.5:0.95 values were 0.360 for YOLOv8n, 0.363 for YOLOv11n, and 0.315 for RT-DETR, while the best Flame values were 0.417, 0.413, and 0.384, respectively. This decrease is expected because mAP@0.5:0.95 applies stricter IoU thresholds and therefore requires more accurate bounding-box localization. Flame and smoke objects often have irregular shapes, blurred boundaries, and partial occlusions, which makes high-IoU localization more difficult than simple object presence detection.
Overall, the test dataset-based evaluation indicates three major findings. First, Flame detection consistently outperformed Smoke detection across all detector architectures and metrics. Second, ChatGPT-4.o- and ChatGPT-5.5-generated images improved or maintained model performance compared with conventional augmentation, but the optimal added-image scale differed by detector and metric. Third, the performance improvement was not simply proportional to the number of added images, suggesting that the quality and distributional relevance of the generated images were more important than dataset size alone. These results support the use of validated generated images as supplementary training data for fire detection models, while also showing that their effectiveness should be evaluated separately for each detector architecture, class, and evaluation metric.
Based on the test-dataset evaluation results, representative improved events were selected for statistical verification.
Table 6 summarizes the paired bootstrap analysis conducted to determine whether the observed performance differences between the ChatGPT-4.o-based models and the corresponding conventional augmentation models were statistically reliable.
As shown in
Table 6, the YOLOv11 model trained with the ChatGPT-4.o 2400 dataset improved Precision from 0.635 to 0.687 compared with the AUG 2400 model, corresponding to a difference of +0.052 with a 95% confidence interval (CI) of [0.021, 0.082]. For Recall, the YOLOv11 model trained with the ChatGPT-4.o 3000 dataset improved the value from 0.569 to 0.601 compared with the AUG 2400 model, yielding a difference of +0.031 with a 95% CI of [0.005, 0.054].
Since both confidence intervals were entirely above zero, the selected improvements in Precision and Recall were statistically supported under test-image-level bootstrap resampling. A similar trend was observed for the mAP-based metrics. The YOLOv11 model trained with the ChatGPT-4.o 3000 dataset improved mAP@0.5 from 0.506 to 0.571 compared with the AUG 3000 model, yielding a difference of +0.065 with a 95% CI of [0.037, 0.090]. In addition, mAP@0.5:0.95 increased from 0.244 to 0.282, with a difference of +0.038 and a 95% CI of [0.022, 0.054]. These results indicate that the selected ChatGPT-4.o-based training conditions produced statistically reliable improvements over the corresponding conventional augmentation conditions in both detection accuracy and stricter localization-based evaluation.
Figure 7 presents representative inference results on the test dataset using selected model weights from each data configuration. The Original and AUG models detected major flame regions in several cases, but their responses to smoke regions and partially occluded fire areas were less consistent. In contrast, the ChatGPT-4.o- and ChatGPT-5.5-based models showed more stable detection responses across diverse fire scenes, including building fires, dense smoke plumes, and complex multi-object fire situations. In particular, the generated-image-based models detected flame and smoke objects in scenes with irregular fire shapes and low-contrast smoke regions, which supports the quantitative improvements observed in the test-dataset evaluation.
Overall, the test-dataset evaluation, bootstrap analysis, and inference visualization consistently demonstrate that generated fire images can provide useful supplementary training information for object detection-based fire detection models. However, the performance gains were dependent on the detector architecture, evaluation metric, class type, and added-image scale. Therefore, the results suggest that generated images are effective when properly screened and incorporated into the training dataset, rather than indicating universal superiority over conventional augmentation under all conditions.
To examine the detection behavior for small fire-related objects, AP50 was additionally calculated for ground-truth objects whose normalized bounding-box area was less than 5% of the image area. As summarized in
Table 7, the GPT-4o-3000 training conditions yielded higher small-object mAP@0.5 values than the corresponding AUG3000 conditions for all detector architectures. The mAP@0.5 increased from 0.194 to 0.212 for YOLOv8, from 0.202 to 0.214 for YOLOv11, and from 0.180 to 0.188 for RT-DETR. These results indicate that the GPT-4o-generated images were beneficial for improving small-object detection performance, although the magnitude of improvement varied across detector architectures.
The class-wise AP50 results show that this improvement was mainly associated with smoke detection. Smoke AP50 was consistently higher than Flame AP50 in all experimental conditions, ranging from 0.337 to 0.400, whereas Flame AP50 remained low, ranging from 0.020 to 0.033. This result suggests that small flame regions remain difficult to detect, even when generated images are added to the training dataset. Small flame objects generally occupy only a limited number of pixels and often show unstable boundary characteristics, making them more susceptible to confusion with light sources, reflections, and complex background patterns.
The row-normalized confusion matrices in
Figure 8 provide further insight into the error characteristics of the models. Compared with the AUG3000 models, the GPT-4o-3000 models generally showed higher diagonal values for both Smoke and Flame, indicating improved class-wise correct classification. For YOLOv11, the correct classification rate increased from 65.3% to 66.9% for Smoke and from 37.7% to 45.7% for Flame. For RT-DETR, the corresponding values increased from 72.5% to 75.4% for Smoke and from 51.0% to 57.5% for Flame. In addition, the proportion of Smoke and Flame objects assigned to the Background column decreased in most GPT-4o-3000 conditions, indicating that generated images contributed to reducing missed detections.
However, the confusion matrices also show that two error types remained prominent. First, a considerable proportion of Flame objects was still assigned to Background, particularly for YOLO-based models, which is consistent with the low AP50 values observed for small flame targets. Second, the Background row was frequently classified as Smoke, indicating that smoke-like textures, low-contrast regions, and visually ambiguous background patterns can still produce false-positive smoke detections. Therefore, although the GPT-4o-3000 dataset improved small-object detection and reduced some missed detections, further refinement is required to improve small-flame localization and suppress smoke-related false positives.