Managing Cost–Stability Trade-Offs in Industrial Object Detection: A Unified Decision Support Framework
Abstract
1. Introduction
At which confidence threshold should this detector be operated in my specific environment, given my cost structure and robustness requirements?
- In real systems, detectors output continuous confidence scores that must be converted into binary decisions—e.g., “defect/no defect”, “smoke/no smoke”—to trigger downstream actions. Therefore, selecting a single operating threshold along the confidence axis is unavoidable. However, this choice must be made under three intertwined sources of imprecision and uncertainty:
- (i)
- A threshold-specific evaluation for binary decisions. The perception module must ultimately output binary decisions to be actionable. However, global metrics such as the mAP average the performance over all thresholds, providing limited insight into local behavior [31,32]. Practitioners thus resort to auxiliary single-point criteria such as the F1 score [33] or localization recall precision (LRP) [34], but these lack a built-in notion of robustness around the operating point.
- (ii)
- Asymmetric and partially known error costs. The consequences of false positives (FPs) and false negatives (FNs) are rarely symmetric. In safety-critical domains, FNs can be catastrophic, while in others, FPs may dominate costs. In practice, relative costs are often approximate soft domain preferences rather than exact numbers [35,36,37]. Standard metrics implicitly assume balanced costs, and a systematic framework handling these soft preferences is lacking.
- (iii)
- Robustness to distribution shifts around the operating threshold. Once deployed, detectors inevitably face data distribution shifts [38,39,40]. Even modest calibration changes can move the model along the confidence axis, making a validation-optimal threshold brittle. Existing metrics focus on single points or global aggregates, failing to explicitly quantify the local stability around the chosen threshold.
- From a soft-computing viewpoint, these requirements highlight the need to handle imprecision, uncertainty, and inherent ambiguity in detector outputs while supporting tractable deployment decisions—exactly the type of setting where soft computing is positioned as a flexible alternative to rigid, hard-computing formulations [22]. In our context, this softness manifests in three aspects: (1) cost asymmetry, (2) acceptable localization ranges, and (3) robustness bands. Similar soft-preference ideas have been effectively used to encode flexible preferences in clustering and supply-chain optimization [41,42].
- Phase 1 (knowledge elicitation and parameterization): Qualitative field requirements are translated into three domain parameters (k, , ). These act as soft domain preferences reflecting domain-specific trade-offs.
- Phase 2 (robust and cost-sensitive evaluation): We define a cost-sensitive LRP extended with a robustness band objective. For each candidate threshold s, DSR-LRP aggregates the performance over the interval , jointly capturing the mean error and variability. This converts the qualitative tolerance for uncertainty into a quantitative objective.
- Phase 3 (analysis and interpretable recommendation): We identify the optimal threshold and compile an interpretable package containing the recommended point, a quantitative rationale, and visual evidence.
- Aligned with the principles of soft computing, our framework exploits the tolerance for imprecision and asymmetric costs to achieve robust deployment decisions with minimal overhead. We evaluated DSR-LRP on four diverse datasets. The results confirm that DSR-LRP consistently favors operating thresholds that are stable under distribution shifts and aligned with domain priorities.
2. Practitioner-Centric Evaluation Tools for Object Detection
2.1. Introduction: Evaluation Tools as Decision Support Components
2.2. The Ambiguity of Global Aggregation vs. Binary Decisions
2.3. Rigid Counting vs. Soft Domain Preferences
2.4. Pointwise Optima vs. Robustness to Distribution Shifts
2.5. Summary: The Need for a Unified Framework
3. Methodology: The DSR-LRP Decision Support Framework
3.1. Phase 1: Knowledge Elicitation and Parameterization
- Classification cost ratio (, ): encodes soft preferences over relative error costs, where applies when missed detections are more critical (e.g., wildfire monitoring) and applies when false alarms waste resources (e.g., unnecessary maintenance).
- Localization importance (): controls the tolerance for bounding-box imprecision, where imposes a strict penalty, while applies no additional penalty beyond the minimum IoU threshold .
- Robustness band (): defines the width of the interval over which stability is evaluated to anticipate performance changes under post-deployment distribution shifts; for example, setting means that, for a given candidate threshold s, both the performance and stability are evaluated across the entire interval centered on s.

- Implementation Memo
- Practical Implications

3.2. Phase 2: Integration of Cost Sensitivity and a Robustness Band
3.2.1. Foundation: Cost-Sensitive LRP for Pointwise Evaluation
Re-Parameterization Using Elicited Knowledge
Why the New Normalization Is Necessary
Implementation Memo
Practical Implications
3.2.2. Band Extension: Robustness Band Objective Function
Implementation Memo
Practical Implications
3.3. Phase 3: Analysis and Interpretable Recommendation
- Optimal point: This component presents the candidate’s optimal operating threshold and its corresponding score, the recommended pair . This point is identified by finding the minimum of the DSR-LRP score curve ().
- Automated rationale report: This provides an automatically generated, descriptive summary under the specified Phase 1 parameters. It reports the recommended operating threshold , its score , and the shift relative to the LRP optimum , defined as . For example, “With , , and , the recommended operating threshold is (with a DSR-LRP score of 0.249); relative to the LRP optimum , the shift is .” This report is descriptive by design; causal interpretation is left to practitioners.
- Visual evidence figure: Beyond a single optimum, this figure provides a multi-faceted analysis of the performance landscape around the selected threshold. It visualizes the performance curves over the entire threshold range, with the confidence score on the x-axis and the LRP and DSR-LRP scores on the y-axis. This figure allows practitioners to intuitively see where is situated in the overall performance landscape and confirms that it lies within a flat region, indicating low performance degradation even under minor distribution shifts. This serves as direct visual evidence of why the selected point is assessed as robust.
- This multi-faceted delivery of evidence maximizes the transparency of the evaluation results and empowers practitioners to make final decisions with evidence-driven confidence. An example of the three-part recommendation package is shown in Figure 3.
- Implementation Memo
- Practical Implications
4. Experiments
4.1. Dataset Information
4.1.1. Blood Cell Dataset
4.1.2. Wildfire Smoke Dataset
4.1.3. Pothole Dataset
4.1.4. Semiconductor Manufacturing Equipment Sensor Dataset
4.2. Confidence Score Calibration Methods
4.2.1. Post Hoc Calibration
- Histogram binning: This method divides the confidence scores into several bins and uses the average accuracy within each bin as the new calibrated confidence. The calibration function is:where is the set of predictions whose confidence score falls into the ith interval, and is the outcome of the jth prediction (1 for a correct detection, 0 otherwise).
- Logistic calibration: This method applies a linear logit model for calibration, defined as , within the standard sigmoid calibration map . Therefore, the final correction map is given by:
- Beta calibration: This method employs a more flexible non-linear model for the log-odds term , which is then transformed by the same sigmoid map . The final correction map is given by:
4.2.2. Auxiliary Loss Calibration with Monte Carlo (MC) Dropout
4.3. Experimental Results and Analysis
4.3.1. Blood Cell Dataset
4.3.2. Wildfire Smoke Dataset
4.3.3. Pothole Dataset
4.3.4. Semiconductor Manufacturing Equipment Sensor Dataset
4.3.5. Cross-Scenario Interpretation of Parameter Roles
| Calibration | mAP | mAP50 | D-ECE (%) | LRP () | DSR-LRP |
|---|---|---|---|---|---|
| Original | 0.347 | 0.609 | 0.73 | 0.497 (0.23–0.24) | 0.501 (0.24) |
| Histogram binning | 0.300 | 0.560 | 0.37 | 0.536 (0.31–0.33) | 0.541 (0.31) |
| Logistic | 0.347 | 0.609 | 0.54 | 0.529 (0.08) | 0.531 (0.05) |
| Beta | 0.347 | 0.609 | 0.37 | 0.522 (0.20) | 0.529 (0.23) |
| MC dropout | 0.345 | 0.630 | 2.00 | 0.505 (0.31) | 0.510 (0.31) |
5. Conclusions
5.1. Summary of Contributions
- Mitigating threshold brittleness. By evaluating the performance within a robustness band rather than at a single point, DSR-LRP tends to steer the operating point away from unstable regions prone to sharp degradation. This capability was highlighted in our pothole case study, where the framework recommended a more robust operating point that reduced the sensitivity to simulated shifts from 15.6% (standard optimum) to 1.8%.
- Operationalizing soft domain preferences. Through its configurable parameters (k and ), DSR-LRP translates high-level, often qualitative domain priorities into the final deployment decision (as illustrated by the contrasting cost structures in the wildfire and pothole datasets).
- Adaptive consistency. In our experiments, when the model was already stable, DSR-LRP tended to avoid unnecessary adjustments and converge to the pointwise optimum, acting as a reliable guardrail that intervenes mainly when uncertainty poses a risk (as demonstrated in the semiconductor dataset).
- In conclusion, DSR-LRP is more than a simple evaluation metric; it is a soft-computing-oriented tool that helps bridge the gap between rigid academic benchmarks and the flexible, robust needs of industry. By explicitly managing uncertainty and maximizing transparency, it empowers practitioners to make reliable, evidence-driven deployment decisions.
5.2. Limitations and Future Work
- Computational overhead. DSR-LRP introduces additional computation because the LRP must be evaluated at multiple points within the robustness band. However, this evaluation is performed post hoc on stored detection outputs and does not require rerunning detector inference. Therefore, the practical overhead is modest; for example, with and , only 11 band points are evaluated for each candidate threshold. The cost increases linearly with a wider or finer , but remains a one-time offline cost during evaluation.
- From post hoc evaluation to in-training optimization. Our framework currently applies its core logic (Phases 2 and 3) only after model training. A natural next step is to formulate a differentiable loss function inspired by DSR-LRP’s principles. This would enable models to internalize stability and cost-awareness directly during the learning process, evolving from systems that are merely evaluated for robustness to those that are trained for it.
- From expert tuning to systematic preference elicitation. Phase 1 currently relies on practitioner knowledge to set parameters (k, , and ), which may introduce subjectivity. Future work could develop systematic methodologies for this process, for instance, by adapting multi-criteria decision-making (MCDM) frameworks [54] to automatically elicit and optimize these soft preferences from data or interactive user feedback. Although this paper provides default settings in Section 3.1 and a cross-scenario interpretation of parameter roles in Section 4.3.5, the recommended threshold can still depend on these practitioner-defined choices. A controlled within-dataset factorial sweep over k, , and , together with multi-seed statistical significance testing, is therefore left as future work. More broadly, future research could also adapt large-scale group decision-making approaches, such as Trillo et al. [55], when multiple stakeholders are involved in defining deployment preferences.
- Integration with candidate-level deployment audits. A natural extension is to integrate DSR-LRP with candidate-level deployment audit frameworks such as CSEF [50]. In such a two-tier pipeline, DSR-LRP could first identify a cost-aware and locally robust operating threshold for each candidate, after which CSEF-style auditing could evaluate whether the selected candidate satisfies broader deployment-level stability and operational constraints. This integration would preserve the distinction between threshold-level operating-point selection and candidate-level deployment auditing, while allowing both perspectives to support a unified deployment decision process.
- Generalization to foundation and multimodal detectors. DSR-LRP is architecture-agnostic in principle because it only requires predicted boxes, class labels, confidence scores, and ground-truth annotations. However, open-vocabulary, multimodal, or foundation-model-based detectors may introduce additional uncertainty sources, such as prompt sensitivity, semantic ambiguity among classes, and class-dependent confidence calibration. Extending DSR-LRP to these emerging architectures is a promising future direction.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| Symbol | Description |
| s | A candidate confidence threshold. |
| The final optimal operating threshold recommended by the | |
| DSR-LRP framework. | |
| The threshold that minimizes the pointwise LRP score without | |
| robustness consideration. | |
| k | The classification cost ratio (). A value greater than 1 implies |
| a higher cost for FNs, and less than 1 for FPs. | |
| Cost multipliers for FPs and FNs, derived from the cost ratio k. | |
| The localization importance, a value between 0 and 1 that controls the | |
| tolerance for imprecise bounding boxes. | |
| The robustness band, defining the interval width for evaluating the performance | |
| stability against potential distribution shifts. | |
| The step size for sweeping thresholds (e.g., 0.01). | |
| N | The number of discrete steps within the robustness band, calculated as |
| . | |
| The cost-sensitive localization recall precision (LRP) error score at a single | |
| threshold s. | |
| The final score for a candidate threshold s with a robustness band , | |
| considering both the average performance and variability within the band. | |
| The number of true positives. | |
| The number of false positives. | |
| The number of false negatives. | |
| X | The set of ground-truth bounding boxes. |
| The set of detections whose confidence scores are higher than the threshold s. | |
| Z | The normalization constant to keep the LRP score within the range [0, 1]. |
| The minimum intersection over the union (IoU) threshold for a detection to be | |
| considered a TP. |
Appendix A. Algorithm for DSR-LRP Framework
| Algorithm A1 DSR-LRP framework (part 1 of 2). | |
| 1: | Phase 1: Knowledge elicitation and parameterization |
| 2: | Based on practitioners’ soft domain preferences, define the core domain-specific parameters: |
| 3: | - k: Classification cost ratio () |
| 4: | - : Localization importance () |
| 5: | - : Robustness band (expected magnitude of confidence score shifts) |
| 6: | Input: |
| 7: | - : All detection results from the candidate |
| 8: | - : All ground-truth data |
| 9: | - : Minimum IoU threshold for a TP |
| 10: | Preparation: |
| 11: | Initialize empty dictionaries: , |
| 12: | Calculate cost multipliers: ; |
| 13: | Return: Parameters and initialized dictionaries. |
| 14: | Phase 2: Robust and cost-sensitive evaluation |
| 15: | Preliminary: Result Classification |
| 16: | For each confidence threshold s, classify detection results in against using the IoU threshold to obtain counts of . |
| 17: | 2.1. Calculate cost-sensitive LRP |
| 18: | For each point s, calculate the pointwise LRP score using the formula: |
| 19: | Return: A dictionary mapping each threshold to its LRP score, . |
| 20: | 2.2. Calculate DSR-LRP |
| 21: | Input: |
| 22: | - : Step size for sweeping thresholds within the robustness band (e.g., 0.01). |
| 23: | For a given candidate threshold s, evaluate stability over the band using the formula: |
| where is the number of discrete steps within the band. | |
| 24: | Return: A dictionary mapping each threshold to its DSR-LRP score, . |
| Algorithm A2 DSR-LRP framework (part 2 of 2). | |
| 1: | Input: |
| 2: | - : Dictionary of LRP scores from Part 1 |
| 3: | - : Dictionary of DSR-LRP scores from Part 1 |
| 4: | - Original parameters |
| 5: | Phase 3: Analysis and interpretable recommendation |
| 6: | Identify optimal thresholds by finding the minimum values in the returned dictionaries: |
| 7: | |
| 8: | |
| 9: | Compile the final recommendation package, , with the following components: |
| 10: | - Optimal point: The recommended pair . |
| 11: | - Rationale: A quantitative summary of the cost-stability trade-off, including the shift from to under the given parameters. |
| 12: | - Visual evidence: A figure plotting the computed and curves to visualize the performance landscape. |
| 13: | Return: The complete recommendation package, . |
References
- Zou, Z.; Chen, K.; Shi, Z.; Guo, Y.; Ye, J. Object Detection in 20 Years: A Survey. Proc. IEEE 2023, 111, 257–276. [Google Scholar] [CrossRef]
- Sun, Y.; Sun, Z.; Chen, W. The evolution of object detection methods. Eng. Appl. Artif. Intell. 2024, 133, 108458. [Google Scholar] [CrossRef]
- Li, Z.; Dong, Y.; Shen, L.; Liu, Y.; Pei, Y.; Yang, H.; Zheng, L.; Ma, J. Development and challenges of object detection: A survey. Neurocomputing 2024, 598, 128102. [Google Scholar] [CrossRef]
- Liu, L.; Ouyang, W.; Wang, X.; Fieguth, P.; Chen, J.; Liu, X.; Pietikäinen, M. Deep Learning for Generic Object Detection: A Survey. Int. J. Comput. Vis. 2020, 128, 261–318. [Google Scholar] [CrossRef]
- Bhatt, P.; Malhan, R.; Rajendran, P.; Shah, B.; Thakar, S.; Yoon, Y.; Gupta, S. Object Detection in 20 Years: A Survey. J. Comput. Inf. Sci. Eng. 2021, 21, 040801. [Google Scholar] [CrossRef]
- Yang, J.; Liu, Z. A novel real-time steel surface defect detection method with enhanced feature extraction and adaptive fusion. Eng. Appl. Artif. Intell. 2024, 138, 109289. [Google Scholar] [CrossRef]
- Chen, S.; Jiang, S.; Wang, X.; Sun, P.; Hua, C.; Sun, J. An efficient detector for detecting surface defects on cold-rolled steel strips. Eng. Appl. Artif. Intell. 2024, 138, 109325. [Google Scholar] [CrossRef]
- Sun, P.; Hua, C.; Ding, W.; Hua, C.; Liu, P.; Lei, Z. Ceramic tableware surface defect detection based on deep learning. Eng. Appl. Artif. Intell. 2025, 141, 109723. [Google Scholar] [CrossRef]
- Grigorescu, S.; Trasnea, B.; Cocias, T.; Macesanu, G. A survey of deep learning techniques for autonomous driving. J. Field Robot. 2020, 37, 362–386. [Google Scholar] [CrossRef]
- Muhammad, K.; Ullah, A.; Lloret, J.; Ser, J.D.; de Albuquerque, V.H.C. Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions. IEEE Trans. Intell. Transp. Syst. 2021, 22, 4316–4336. [Google Scholar] [CrossRef]
- Kuutti, S.; Bowden, R.; Jin, Y.; Barber, P.; Fallah, S. A Survey of Deep Learning Applications to Autonomous Vehicle Control. IEEE Trans. Intell. Transp. Syst. 2021, 22, 712–733. [Google Scholar] [CrossRef]
- Thottempudi, P.; Jambek, A.B.B.; Kumar, V.; Acharya, B.; Moreira, F. Resilient object detection for autonomous vehicles: Integrating deep learning and sensor fusion in adverse conditions. Eng. Appl. Artif. Intell. 2025, 151, 110563. [Google Scholar] [CrossRef]
- Li, X.; Lin, K.; Meng, M.; Li, X.; Li, L.; Hong, Y.; Chen, J. A Survey of ADAS Perceptions with Development in China. IEEE Trans. Intell. Transp. Syst. 2022, 23, 14188–14203. [Google Scholar] [CrossRef]
- Litjens, G.; Kooi, T.; Bejnordi, B.; Setio, A.; Ciompi, F.; Ghafoorian, M.; Laak, J.; Ginneken, B.; Sánchez, C. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef]
- Suganyadevi, S.; Seethalakshmi, V.; Balasamy, K. A review on deep learning in medical image analysis. Int. J. Multimed. Inf. Retr. 2022, 11, 19–38. [Google Scholar] [CrossRef] [PubMed]
- Razzak, M.; Naz, S.; Zaib, A. Deep Learning for Medical Image Processing: Overview, Challenges and the Future. In Classification in BioApps; Springer International Publishing: Cham, Switzerland, 2018; Volume 26, pp. 323–350. [Google Scholar]
- Liu, X.; Gao, K.; Liu, B.; Pan, C.; Liang, K.; Yan, L.; Ma, J.; He, F.; Zhang, S.; Pan, S.; et al. Advances in Deep Learning-Based Medical Image Analysis. Health Data Sci. 2021, 2021, 8786793. [Google Scholar] [CrossRef]
- Chen, X.; Wang, X.; Zhang, K.; Fung, K.; Thai, T.; Moore, K.; Mannel, R.; Liu, H.; Zheng, B.; Qiu, Y. Recent advances and clinical applications of deep learning in medical image analysis. Med. Image Anal. 2022, 79, 102444. [Google Scholar] [CrossRef] [PubMed]
- Cano-Ortiz, S.; Iglesias, L.; Árbol, P.; Castro-Fresno, D. Improving detection of asphalt distresses with deep learning-based diffusion model for intelligent road maintenance. Dev. Built Environ. 2024, 17, 100315. [Google Scholar] [CrossRef]
- Dhiman, A.; Klette, R. Pothole Detection Using Computer Vision and Learning. IEEE Trans. Intell. Transp. Syst. 2020, 21, 3536–3550. [Google Scholar] [CrossRef]
- Ma, N.; Fan, J.; Wang, W.; Wu, J.; Jiang, Y.; Xie, L.; Fan, R. Computer vision for road imaging and pothole detection: A state-of-the-art review of systems and algorithms. Transp. Saf. Environ. 2022, 4, tdac026. [Google Scholar] [CrossRef]
- Zadeh, L.A. Fuzzy logic, neural networks, and soft computing. Commun. ACM 1994, 37, 77–84. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2016; pp. 779–788. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the Computer Vision—ECCV 2020, Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2020; Volume 12346, pp. 213–229. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Everingham, M.; Gool, L.; Williams, C.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef]
- Lin, T.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C. Microsoft COCO: Common Objects in Context. In Proceedings of the Computer Vision—ECCV 2014. Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2014; Volume 8693, pp. 740–755. [Google Scholar]
- Liu, Y.; Sun, P.; Wergeles, N.; Shang, Y. A survey and performance evaluation of deep learning methods for small object detection. Expert Syst. Appl. 2021, 172, 114602. [Google Scholar] [CrossRef]
- Goswami, P.; Aggarwal, L.; Kumar, A.; Kanwar, R.; Vasisht, U. Real-time evaluation of object detection models across open world scenarios. Appl. Soft Comput. 2024, 163, 111921. [Google Scholar] [CrossRef]
- Wenkel, S.; Alhazmi, K.; Liiv, T.; Alrshoud, S.; Simon, M. Confidence Score: The Forgotten Dimension of Object Detection Performance Evaluation. Sensors 2021, 21, 4350. [Google Scholar] [CrossRef] [PubMed]
- Jena, R.; Zhornyak, L.; Doiphode, N.; Chaudhari, P.; Buch, V.; Gee, J.; Shi, J. Beyond mAP: Towards Better Evaluation of Instance Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 11309–11318. [Google Scholar]
- Powers, D.W. Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. Int. J. Mach. Learn. Technol. 2011, 2, 37–63. [Google Scholar]
- Oksuz, K.; Cam, B.; Akbas, E.; Kalkan, S. Localization Recall Precision (LRP): A New Performance Metric for Object Detection. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 504–519. [Google Scholar]
- Frumosu, F.; Khan, A.; Schiøler, H.; Kulahci, M.; Zaki, M.; Westermann-Rasmussen, P. Cost-sensitive learning classification strategy for predicting product failures. Expert Syst. Appl. 2020, 161, 113653. [Google Scholar] [CrossRef]
- Sbeyti, M.; Karg, M.; Wirth, C.; Klein, N.; Albayrak, S. Cost-Sensitive Uncertainty-Based Failure Recognition for Object Detection. In Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence, Barcelona, Spain, 15–19 July 2024; Volume 244, pp. 1890–1900. [Google Scholar]
- Viaene, S.; Dedene, G. Cost-sensitive learning and decision making revisited. Eur. J. Oper. Res. 2005, 166, 212–220. [Google Scholar] [CrossRef]
- Al-Emadi, S.; Yang, Y.; Ofli, F. Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2025; pp. 8299–8309. [Google Scholar]
- Khan, T.; Singh, K.; Bhati, B.; Ahmad, K.; Al-Rasheed, A.; Getahun, M.; Soufiene, B. Trust-driven approach to enhance early forest fire detection using machine learning. Sci. Rep. 2025, 15, 14480. [Google Scholar] [CrossRef]
- Zhang, Y.; Carballo, A.; Yang, H.; Takeda, K. Perception and sensing for autonomous vehicles under adverse weather conditions: A survey. ISPRS J. Photogramm. Remote Sens. 2023, 196, 146–177. [Google Scholar] [CrossRef]
- Zhang, Z.; Yu, X.; Tao, R.; Zhang, X.; Li, H.; Lu, J.; Zhou, J. Knowledge augmentation-based soft constraints for semi-supervised clustering. Appl. Soft Comput. 2023, 144, 110484. [Google Scholar]
- Dehshiri, S.J.H.; Amiri, M.; Olfat, L.; Pishvaee, M.S. A robust fuzzy stochastic multi-objective model for stone paper closed-loop supply chain design considering the flexibility of soft constraints based on Me measure. Appl. Soft Comput. 2023, 134, 109944. [Google Scholar]
- Kaushal, M.; Khehra, B.S.; Sharma, A. Soft computing based object detection and tracking approaches: State-of-the-art survey. Appl. Soft Comput. 2018, 70, 423–464. [Google Scholar] [CrossRef]
- Liu, Y.; Zhou, C.; Guo, D.; Wang, K.; Pang, W.; Zhai, Y. A decision support system using soft computing for modern international container transportation services. Appl. Soft Comput. 2010, 10, 1087–1095. [Google Scholar] [CrossRef]
- Khelifa, B.; Laouar, M.R. A holonic intelligent decision support system for urban project planning by ant colony optimization algorithm. Appl. Soft Comput. 2020, 96, 106621. [Google Scholar] [CrossRef]
- Küppers, F.; Kronenberger, J.; Shantia, A.; Haselhoff, A. Multivariate Confidence Calibration for Object Detection. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); IEEE: New York, NY, USA, 2020; pp. 1322–1330. [Google Scholar]
- Padilla, R.; Netto, S.L.; da Silva, E.A.B. A Survey on Performance Metrics for Object-Detection Algorithms. In Proceedings of the 2020 International Conference on Systems, Signals and Image Processing (IWSSIP); IEEE: New York, NY, USA, 2020; pp. 237–242. [Google Scholar]
- Hall, D.; Dayoub, F.; Skinner, J.; Zhang, H.; Miller, D.; Corke, P.; Carneiro, G.; Angelova, A.; Suenderhauf, N. Probabilistic Object Detection: Definition and Evaluation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2020; pp. 1031–1040. [Google Scholar]
- van Rijsbergen, C.J. Information Retrieval, 2nd ed.; Butterworths: London, UK, 1979. [Google Scholar]
- Lee, K.; Hong, J.; Kim, B.; Song, Y.; Lee, D. Cost-Stability Evaluation Framework for Reliable Deployment of Object Detectors in Industrial Applications. IEEE Access 2025, 13, 207566–207580. [Google Scholar] [CrossRef]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K. On calibration of modern neural networks. In Proceedings of the ICML’17: Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 6–11 August 2017; Volume 70, pp. 1321–1330. [Google Scholar]
- Zhang, J.; Cho, J.; Zhou, X.; Krähenbühl, P. NMS Strikes Back. arXiv 2022, arXiv:2212.06137. [Google Scholar] [CrossRef]
- Pathiraja, B.; Gunawardhana, M.; Khan, M. Multiclass Confidence and Localization Calibration for Object Detection. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 19734–19743. [Google Scholar]
- Lee, D.; Kim, K. Interactive weighting of bias and variance in dual response surface optimization. Expert Syst. Appl. ESWA 2012, 39, 5900–5906. [Google Scholar]
- Trillo, J.R.; Herrera-Viedma, E.; Morente-Molinera, J.A.; Cabrerizo, F.J. A large scale group decision making system based on sentiment analysis cluster. Inf. Fusion 2023, 91, 633–643. [Google Scholar]










| Approach | Brief Description | Practitioner-Centric Deployment Requirements | ||
|---|---|---|---|---|
| Local Threshold | Cost Weighting | Local Robustness Around Threshold | ||
| mAP [47] | Global PR-curve summary over confidence thresholds. | × | × | × |
| D-ECE [46] | Calibration error over confidence levels. | × | × | × |
| F1 score [33] | Single-threshold harmonic mean of precision and recall. | ✓ | × | × |
| score [49] | Single-threshold metric with recall-oriented weighting. | ✓ | ▵ | × |
| LRP [34] | Single-threshold localization recall precision error. | ✓ | ▵ | × |
| CSEF [50] | Candidate-level deployment audit using cost-sensitive operating performance and stability checks. | ✓ | ✓ | ▵ |
| DSR-LRP | Threshold-level objective with cost-sensitive LRP and local robustness band. | ✓ | ✓ | ✓ |
| Blood Cell | Image | Class | ||
|---|---|---|---|---|
| Platelets | RBCs | WBCs | ||
| Train | 255 | 249 | 2936 | 263 |
| Validation | 73 | 76 | 819 | 72 |
| Test | 36 | 36 | 398 | 37 |
| Wildfire smoke | Image | Class | ||
| Smoke | ||||
| Train | 516 | 516 | ||
| Validation | 147 | 147 | ||
| Test | 74 | 74 | ||
| Pothole | Image | Class | ||
| Pothole | ||||
| Train | 465 | 1256 | ||
| Validation | 133 | 330 | ||
| Test | 67 | 154 | ||
| SE sensor | Image | Class | ||
| Average | Deviation | Drift | ||
| Train | 1205 | 582 | 498 | 275 |
| Validation | 347 | 167 | 176 | 78 |
| Test | 173 | 83 | 58 | 40 |
| Calibration | mAP | mAP50 | D-ECE (%) | LRP () | DSR-LRP |
|---|---|---|---|---|---|
| Original | 0.577 | 0.870 | 3.49 | 0.491 (0.40) | 0.494 (0.38) |
| Histogram binning | 0.546 | 0.837 | 2.37 | 0.510 (0.31) | 0.513 (0.34–0.35) |
| Logistic | 0.577 | 0.870 | 3.64 | 0.493 (0.22) | 0.493 (0.24) |
| Beta | 0.577 | 0.870 | 2.38 | 0.493 (0.43) | 0.494 (0.40–0.41) |
| MC dropout | 0.574 | 0.865 | 6.85 | 0.490 (0.44) | 0.495 (0.40–0.41) |
| Calibration | mAP | mAP50 | D-ECE (%) | LRP () | DSR-LRP |
|---|---|---|---|---|---|
| Original | 0.498 | 0.875 | 0.97 | 0.216 (0.15) | 0.260 (0.21) |
| Histogram binning | 0.400 | 0.761 | 0.25 | 0.337 (0.15–0.18) | 0.352 (0.14) |
| Logistic | 0.498 | 0.875 | 0.48 | 0.227 (0.02) | 0.273 (0.11) |
| Beta | 0.498 | 0.875 | 0.32 | 0.216 (0.09) | 0.249 (0.14) |
| MC dropout | 0.533 | 0.913 | 2.14 | 0.188 (0.36–0.37) | 0.219 (0.34–0.35) |
| Calibration | mAP | mAP50 | D-ECE (%) | LRP () | DSR-LRP |
|---|---|---|---|---|---|
| Original | 0.502 | 0.770 | 1.46 | 0.338 (0.38) | 0.348 (0.52–0.53) |
| Histogram binning | 0.419 | 0.681 | 0.47 | 0.410 (0.41–0.42) | 0.412 (0.55) |
| Logistic | 0.502 | 0.770 | 1.07 | 0.338 (0.46) | 0.344 (0.50) |
| Beta | 0.502 | 0.770 | 0.66 | 0.338 (0.58) | 0.347 (0.62) |
| MC dropout | 0.486 | 0.757 | 3.93 | 0.320 (0.42) | 0.336 (0.44) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lee, K.; Hong, J.; Kim, B.-S.; Song, Y.; Lee, D.-H. Managing Cost–Stability Trade-Offs in Industrial Object Detection: A Unified Decision Support Framework. Algorithms 2026, 19, 409. https://doi.org/10.3390/a19050409
Lee K, Hong J, Kim B-S, Song Y, Lee D-H. Managing Cost–Stability Trade-Offs in Industrial Object Detection: A Unified Decision Support Framework. Algorithms. 2026; 19(5):409. https://doi.org/10.3390/a19050409
Chicago/Turabian StyleLee, Kuhyun, Jihoon Hong, Beom-Seok Kim, Yuna Song, and Dong-Hee Lee. 2026. "Managing Cost–Stability Trade-Offs in Industrial Object Detection: A Unified Decision Support Framework" Algorithms 19, no. 5: 409. https://doi.org/10.3390/a19050409
APA StyleLee, K., Hong, J., Kim, B.-S., Song, Y., & Lee, D.-H. (2026). Managing Cost–Stability Trade-Offs in Industrial Object Detection: A Unified Decision Support Framework. Algorithms, 19(5), 409. https://doi.org/10.3390/a19050409

