Next Article in Journal
Analysis of Interlayer Stress Transfer Mechanism of Bimetallic Composite Pipes
Previous Article in Journal
Toxic Efficacy of Nerium oleander Extracts as a Precise Strategy for Managing Pomacea canaliculata (Gastropoda) Eggs
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Research Progress on Intelligent Color-Sorting Equipment for Post-Harvest Chili Peppers: Machine Vision, Pneumatic Actuation, and System Integration

School of Agricultural Engineering, Jiangsu University, Zhenjiang 212013, China
*
Author to whom correspondence should be addressed.
Processes 2026, 14(18), 2991; https://doi.org/10.3390/pr14182991 (registering DOI)
Submission received: 13 August 2026 / Revised: 12 September 2026 / Accepted: 16 September 2026 / Published: 19 September 2026
(This article belongs to the Section AI-Enabled Process Engineering)

Abstract

Post-harvest sorting is essential for converting the biological variability of chili peppers into consistent commercial grades, efficient processing, and higher market value. Rapid advances in machine vision, multimodal sensing, deep learning, and intelligent actuation are transforming sorting equipment from rule-based classifiers into integrated perception-to-execution systems. However, existing studies often evaluate isolated algorithms or components, leaving limited evidence that recognition accuracy translates into reliable sorting during continuous operation. We conducted a structured narrative review of English-language studies published from January 2008 to July 2026 using Web of Science, Scopus, AESC, and PubMed. Evidence was synthesized along a perception-to-execution chain, distinguishing pepper-specific sorting systems and component studies from cross-crop engineering evidence. Direct system-level evidence was concentrated in a small number of chili and bell pepper studies, whereas much of the technical discussion drew on apple, tomato, potato, sweet potato, and onion research. Visible, spectral, fluorescence, and multimodal features can distinguish ripeness, color grades, and surface defects, although performance remains sensitive to cultivar, illumination, pose, and dataset design. A small number of pepper sorting studies support the feasibility of integrating vision, conveying, and physical separation under specific operating conditions. Cross-crop studies inform the discussion of localization, pneumatic actuation, and system integration, but do not establish the performance of these approaches in chili pepper sorting. Reported performance is difficult to compare because studies rarely standardize latency, target association, false and missed rejections, product damage, energy use, and long-term reliability. Robust deployment therefore requires cross-batch datasets, synchronized target-level traceability, and evaluation protocols linking perception outputs to final physical destinations. This review identifies the boundaries of current pepper-specific evidence and proposes a system-level evaluation framework and research priorities for testing the transferability of cross-crop engineering approaches to pepper sorting.

1. Introduction

1.1. Research Background and Industrial Needs

Chili peppers are consumed fresh or processed into foods and seasonings. Variety, color, shape, ripeness, and visible defects determine grading, raw-material allocation, and product consistency [1,2]. Post-harvest sorting must translate this biological variability into consistent grade boundaries.
Chili-sorting research spans two dimensions: perception technology and equipment configuration. Visible-light imaging captures surface color, shape, and defects, whereas multi-band or multimodal sensing can distinguish quality attributes that overlap in conventional color spaces. Equipment ranges from systems integrating image analysis, conveying, and actuation to simpler devices that use color cues to trigger diversion [3,4,5,6]. Integrated systems classify products from multiple features; color-triggered systems require less computation and hardware. Direct comparison remains difficult because material presentation, grade definitions, conveying modes, and evaluation endpoints differ. Classification accuracy alone therefore does not represent final sorting performance. Reviews of machine vision identify image quality, dataset representativeness, and environmental adaptability as fundamental requirements [7,8]. Online sorting adds conveying and actuation constraints: each decision must reach the correct actuator within the rejection window. Evaluation should therefore include throughput, system false-rejection and missed-rejection rates, response latency, material damage, and inter-batch stability. Higher throughput reduces the time available for acquisition, computation, and actuation, while decision thresholds trade false rejection against missed rejection. Deep learning has reduced reliance on manually designed features for fruit detection, classification, and defect identification. RGB imaging and spectral feature selection have both been applied to chili ripeness classification [9,10,11,12]. RGB describes visible surface characteristics, whereas selected spectral bands can distinguish stages that overlap in conventional color spaces. Their practical value depends on whether these gains justify added sensing and computation under illumination variation, reflection, occlusion, pose variation, sample heterogeneity, and limited edge-computing resources.
Chili-sorting studies identify persistent challenges due to varietal diversity, continuous color gradients, irregular geometries, and interactions between color and other quality attributes [2]. Fixed thresholds and a single offline accuracy value are therefore insufficient for evaluating sorting equipment. This review examines how crop attributes, visual perception, target characterization, conveying, positioning, rejection, and system evaluation jointly determine whether correct identification produces correct physical separation.

1.2. Definitions and Scope of Post-Harvest Color Sorting and Intelligent Color-Sorting Equipment for Peppers

Post-harvest color sorting of peppers refers to the online grading or removal of individual, spatially distinguishable peppers according to predefined visual criteria. These criteria include color grade, visual indicators of maturity, and visible color abnormalities. The process should follow nondestructive-testing principles while limiting handling damage to an acceptable level. Color, shape, and external defects have been used for pepper classification and sorting [9,11,13,14]. Within the scope of this review, color information must directly contribute to an online sorting decision. General post-harvest quality prediction without a corresponding sorting action is excluded.
Intelligent color-sorting equipment comprises modules for image acquisition, information processing, quality classification, target information generation, conveyor control, and physical separation [3,4,5,13,14]. For discrete objects, detection, localization, or tracking must be completed before actuation. Each classification result must also be associated with the correct physical target [15]. Continuous material streams may instead use region segmentation, line-scan localization, or channel-based triggering. An offline classifier represents only the perception component and cannot characterize whole-system sorting performance.
This review defines its scope according to the endpoint of each decision. Visible-light imaging is used to evaluate surface color, shape, and visible abnormalities. multispectral, fluorescence, and visible–near-infrared sensing can provide complementary information about maturity, tissue condition, and subsurface characteristics [10,16,17,18]. However, the additional sensing and computational complexity should be justified by a net improvement in whole-system output performance. Studies should clearly define the grading categories or abnormality types and link each decision to a corresponding sorting action. Predictions of sugar content, moisture, nutritional composition, or flavor are treated only as background evidence when they do not alter online sorting decisions.
Accordingly, this review distinguishes three functional levels. The core sorting tasks comprise color grading, visual maturity assessment, and the detection of visible color abnormalities. Equipment functions comprise target information generation, conveyor-based positioning, and physical separation. Internal or subsurface attributes are considered extended capabilities only when they affect online sorting decisions.

1.3. Research Progress, Limitations of Existing Reviews, and the Perspective Adopted in This Review

Post-harvest quality assessment now combines visible appearance inspection, machine vision, spectral sensing, and artificial intelligence [19,20,21]. These methods support different decision functions and engineering endpoints, so reviews should not be organized by sensing modality alone. Existing reviews emphasize different endpoints across crops. Chili studies focus on handling, preservation, and crop-specific requirements, whereas tomato and citrus studies emphasize nondestructive testing and predictive performance [20,21,22,23]. Berry research spans the processing chain [20,21,22,23]. Cross-crop sensors, features, and algorithms may inform chili sorting, but differences in morphology and conveying conditions limit direct generalization. Differences in samples, imaging conditions, and evaluation metrics also prevent direct comparison among offline results. Perception-level evidence on maturity, color classification, and feature extraction [9,10,11,24] does not establish reliable physical diversion. A method may classify images accurately without generating the information required for correct and timely diversion. Online sorting additionally requires target identity, spatial position, and classification output to remain associated with the same pepper until actuation. This review therefore organizes evidence by perception-, target-, and system-level endpoints, focusing on the interface between information measurement and executable sorting.
Odor-based sorting provides a non-visual alternative [16], but its engineering value depends on whether it improves physical sorting outcomes. System-level studies have reported online bell pepper sorting systems and implementations based on deep convolution networks [3,4]. When the relationship between the dataset and physical platform is unclear, these results should be treated as indirect rather than complete system-level evidence. The conveyor-based red chili grading system in [5] and the color-triggered bell pepper sorting prototype in [6] are not directly interchangeable. They differ in material presentation, actuation logic, and evaluation endpoints. More broadly, inconsistent evaluation conditions limit comparisons among alternative technical approaches. Accordingly, this review uses pepper sorting as the application focus, synthesizes the limited direct equipment evidence, and evaluates cross-crop studies as sources of candidate methods and engineering hypotheses requiring pepper-specific validation. Table 1 summarizes the available evidence and its applicability boundaries.

1.4. Research Questions, Technical Framework, and Organization of the Review

Post-harvest color sorting of peppers requires color information to be converted into real-time sorting decisions. Classification accuracy describes only the perception component of this process. A machine vision system for food processing comprises data acquisition, representation, discrimination, and decision output. Sorting equipment must additionally integrate material feeding, conveying, illumination, imaging, and actuation mechanisms [25,26]. This review organizes the literature around a technical chain extending from objects and operating conditions to perception, target information, positioning, execution, and system evaluation. Errors introduced at upstream stages can propagate through this chain and cause false rejections, missed rejections, or product damage.
This review addresses three research questions. Q1: How can perception and recognition methods generate reliable and actionable target information across different sorting tasks and operating conditions [9,10,11,27,28]? Q2: How should conveying, positioning, and pneumatic rejection be coordinated under different operating speeds, object spacing, and target densities? This coordination must convert correct identification into the correct physical separation of each target. Q3: Which metrics and criteria are required to evaluate models, grading schemes, and complete sorting systems? Relevant measures include model accuracy, system missed-rejection rate, system false-rejection rate, throughput, response latency, material damage, and inter-batch stability.
Section 3 defines the sorting objects, operating conditions, and equipment requirements. Section 4 and Section 5 address Q1 through system architectures and sensing methods. Section 6 addresses Q2 through conveying, positioning, and rejection, while Section 7 addresses Q3 through evaluation, interfaces, and closed-loop reliability. Section 8 identifies measurable development priorities, and Section 9 summarizes the conclusions and future directions. Figure 1 links the sorting requirements, technical chain, evidence gaps, and development pathways.
The horizontal comparisons consider object and task characteristics, sensing and algorithmic methods, operating conditions, online capabilities, evaluation endpoints, and reported limitations. Cross-crop evidence can inform method transfer, but its applicability must be assessed against the morphological and operational characteristics of peppers. Review articles help characterize domain boundaries, whereas original studies provide evidence for performance claims and engineering conclusions. Comparisons should therefore consider algorithms, operating conditions, and the capacity for stable physical sorting together.

2. Literature Search, Screening, and Evidence Synthesis Methods

To improve transparency and traceability, we used a structured narrative search process with predefined evidence sources, search scope, eligibility criteria, screening, and evidence classification. Effect sizes were not pooled, and no meta-analysis was performed.

2.1. Literature Search Strategy

The primary databases were the Web of Science Core Collection and Scopus. The Agricultural & Environmental Science Collection (AESC) and PubMed were searched as supplementary databases. Searches were conducted between 22 and 25 July 2026 and covered publications from 1 January 2008 to 25 July 2026. The search focused on English-language publications.
Search terms were grouped by target object, task, technology, and system. Terms within each group were combined using “OR”, and the groups were combined using “AND”. Supplementary searches covered visual perception, material handling, pneumatic actuation, and system integration to capture studies using broad machine-related terminology.
To assess transferability, supplementary searches combined each object term with visual perception, material handling, pneumatic actuation, and system-integration terms. Non-pepper fruit and vegetable studies were labelled separately and used only as potentially transferable evidence.
The principal comparisons and thematic synthesis were based on English-language publications. This restriction ensured consistency in bibliographic information, search boundaries, and evidence-classification criteria. Chinese-language studies were used only as contextual evidence concerning industry conditions, technological development, or application scenarios. They were not combined with the core English-language evidence in quantitative performance comparisons.
The database-specific search strategies are reported in Table 2.

2.2. Inclusion and Exclusion Criteria for Literature

These criteria defined a traceable evidence base for a structured narrative review rather than an exhaustive systematic review.

2.2.1. Inclusion Criteria

Studies meeting the following requirements were included.
Priority was given to studies on color grading, ripeness assessment, appearance anomaly detection, online inspection and rejection, and multi-subsystem integration for chili or sweet pepper. Other fruit and vegetable studies were included only as “transferable technological evidence” if they provided visual perception, single-particle conveying, target localization, pneumatic actuation, or system evaluation methods directly relevant to chili color sorting. Such evidence was used for comparing technical approaches, identifying engineering constraints, and formulating hypotheses for further validation, but not for directly proving the performance of chili sorting equipment.
Included studies were required to involve at least one technical module: image acquisition and color recognition (color space, color difference and threshold methods, machine learning and deep learning-based ripeness classification, detection of obvious color defects); material conveying, pose control, and motion perception; pneumatic ejection actuation and rejection control; visual-execution timing coordination and system integration.
Supporting research providing chili variety-specific color characteristics, ripening dynamics, physical properties of materials, damage mechanisms, or quality specifications that are directly relevant to color sorting equipment design was also included. However, such literature was treated separately from core color sorting technology studies during data extraction and synthesis.
Peer-reviewed journal articles and reviews with sufficient technical or experimental detail were prioritized. Conference papers and doctoral dissertations were supplementary evidence only when they provided unique, verifiable information. Journal versions were retained unless conference papers reported distinct conditions or independent results.

2.2.2. Exclusion Criteria

Core evidence excluded non-academic or promotional sources; non-pepper studies without transferable technical modules, experimental conditions, or system-evaluation information; studies unrelated to color quality, ripeness, material properties, post-harvest handling, or color-sorting conditions; and general processing or packaging studies unrelated to sorting, conveying, visual inspection, actuation, or closed-loop management. Patents were used only to identify industrial pathways or equipment configurations, not to support performance conclusions.
Duplicate records were removed. For multiple reports of the same platform, dataset, or experiment, the most complete version was retained; others were kept only for substantial improvements, new conditions, or independent findings. Publications without full text, adequate methodological detail, interpretable key results, or clear relevance were excluded from core comparative evidence.

2.3. Literature Screening and Quality Assessment

Retrieval results were imported into EndNote, where duplicates were first automatically removed by DOI, followed by manual verification based on title, author, year, and journal. Subsequently, initial screening of titles and abstracts and full-text eligibility assessment were conducted. The initial screening excluded records outside the scope of post-harvest color sorting, nondestructive sensing, conveying and positioning, diversion execution, or whole-system evaluation. During full-text assessment, the research subject, experimental context, sample units, data partitioning, technical modules, evaluation endpoints, and availability of full text were verified. Each full-text article was assigned only one primary reason for exclusion to avoid double counting. Quality assessment was based on four predefined dimensions: study design and operational context, sample size and independence of experimental units, duration of replication or validation trials, and completeness of methods and results reporting. Studies meeting fewer than three of these criteria were excluded from core synthesis; however, those providing unique and clearly defined engineering insights were retained only as background evidence. The number of screened studies and reasons for exclusion are shown in Figure 2. Studies retained after screening and quality assessment were further classified according to the directness of the study population and the certainty of the outcome measures, to define the scope of conclusions supported by different studies. The operational criteria and scoring rules used to assess study quality are presented in Table 3.

2.4. Evidence Classification and Synthesis Rules

E1 evidence supported performance assessments, with E2 providing supplementary pepper-specific support. E3 and E4 were used only for transferable methods, engineering constraints, and hypotheses requiring validation, not as independent evidence of overall chili pepper equipment performance (Table 4).
To avoid treating the number of publications as equivalent to the amount of independent evidence, related publications were grouped into study families. Publications were considered related when they used the same dataset, experimental platform, or successive iterations of the same technology.
Multiple publications within a family could report algorithm-, interface-, or system-level results. However, each independent experimental platform was counted only once in the evidence synthesis. Citations to individual publications were retained to preserve the traceability of the reported findings. When platform independence could not be established, the publications were labeled as “possibly related” and assigned less weight in the synthesis.
Included studies differ in crop, task definition, operating conditions, measurement units, denominators, and system boundaries. Reported values are therefore retained under their source definitions and used only for qualitative evidence synthesis. Cross-study numerical ranking is not performed, and non-Capsicum performance is not interpreted as measured or expected performance of Capsicum sorting systems without a harmonized benchmark. Evidence hierarchy used for synthesis and interpretation in Table 4.

3. Post-Harvest Online Intelligent Color-Sorting Tasks and System Requirements

3.1. Post-Harvest Processing Flow and Online Color-Sorting Tasks for Peppers

Online pepper color sorting integrates conveying, visual perception, grade determination, actuation, and result recording. Overall performance depends on synchronized conveying, decision timing, and actuation accuracy [5,29]; therefore, offline classification accuracy indicates perception performance but cannot represent system-level sorting performance.
Nondestructive inspection methods include appearance imaging, which characterizes color and visual grade, and volatile-signal sensing, which captures odor-related quality attributes [5,16,29]. Although these methods may be complementary, few studies have compared them using consistent samples, operating conditions, and evaluation criteria. Sensor selection for online color sorting should therefore be based on the target attributes and required data-acquisition frequency.
Research on other crops primarily provides transferable engineering architectures rather than pepper-specific grading thresholds. Tomato machine-vision sorting systems illustrate approaches to organizing online interfaces, whereas YOLO-RGDD offers strategies for detecting complex surface defects [30,31]. These systems differ from pepper color sorting in task definition, annotation granularity, and the cost of misclassification. Nevertheless, they provide useful reference architectures for continuous conveying, online communication, and defect modeling. Joint validation of these components is required before they can be directly adapted to pepper sorting.
Existing studies on peppers have not fully validated the integration of color assessment, maturity estimation, anomaly detection, and physical grading within a single system [5,16,29]. This limitation makes it difficult to compare algorithms and prototypes. Future studies should report model classification metrics and final-bin miss and false-rejection rates, throughput, and end-to-end latency. They should also clearly distinguish among offline model evaluation, online prototype testing, and experiments with physical sorting systems. The following sections discuss technologies for perception, conveying, and actuation.

3.2. Pepper Variety, Color, Maturity, and Color-Grade Characteristics

The overall color of peppers is determined by pigment accumulation, genetic background, and the optical properties of the fruit surface. As red peppers ripen, their total carotenoid content increases. Color intensity is also associated with capsanthin and its esters, the transcription levels of DXS and PSY-1, and pigment-storage structures [32]. The cuticle and epicuticular wax directly affect surface gloss, whereas carotenoids are more closely associated with L*, a*, and b* values [33]. Redness, color intensity, and gloss are therefore related but not interchangeable. A single color threshold may conflate samples with similar color but different gloss, or similar gloss but different color.
Visible-image studies have generally used three approaches: classification based on handcrafted features, representations obtained through dimensional reduction, and deep feature learning. K-nearest neighbors (KNN), principal component analysis (PCA), and convolutional neural networks (CNNs) have shown the feasibility of classifying peppers under controlled conditions [34,35,36]. However, the term “variety” is not used consistently across studies. It may refer to a genetic type, fruit-color type, or maturity stage, and the resulting model outputs therefore have different biological interpretations. Studies of sweet peppers have also directly estimated maturity stages and selected optimal spectral bands [9,10]. Variety classification and maturity assessment use different labeling schemes. Consequently, variety classification cannot replace maturity assessment or provide universally applicable commercial grading thresholds across varieties.
The performance gains achieved by deep-learning models vary among pepper categories. Following data augmentation, YOLOv8m achieved higher overall accuracy, recall, and mean average precision (mAP). The largest improvement was observed for Cambium, whereas accuracy decreased for Fidalga and Habanero, possibly because of synthetic noise and visual similarities between yellow, round peppers [37]. Aggregate metrics may obscure confusion between adjacent classes. In practical systems, performance depends on the reliable separation of neighboring grades in mixed batches under operating conditions.
Physiological, optical, and classification evidence can be integrated into a testable framework for color-grade stratification. Samples may be classified according to variety or type, maturity-related color stage, overall color intensity, and surface gloss. Because harvest maturity and post-harvest storage affect nutritional composition and phenotype, labels should also record post-harvest duration and storage conditions [38]. Studies using KNN, PCA, CNNs, and YOLOv8m support the feasibility of automated classification but do not yet justify uniform grading across mixed batches [34,35,36,37]. Cross-variety, cross-batch, and cross-storage-stage experiments are needed to assess feature stability and align model categories with commercial grades.
Figure 3 distinguishes biological and optical color attributes from machine-vision classification and commercial grading, clarifying why category recognition does not by itself constitute product grading. The next section addresses abnormal coloration associated with damage, processing, or physiological change.

3.3. Visible Color Abnormalities and Their Relationship to Overall Color Assessment

Visible color abnormalities must be defined relative to the normal characteristics of each variety, maturity stage, and surface-gloss condition. They can be classified as overall color deviations, localized damage, or changes associated with underlying physiological processes. Pigment composition and harvest timing affect the overall color values of red peppers. In sweet peppers, ripeness and freshness can be characterized using digital imaging, chlorophyll fluorescence, and visible/near-infrared (Vis/NIR) signals [39,40,41]. Color abnormalities therefore cannot be defined by fixed color ranges alone. Before selecting thresholds or models, engineering applications should specify the detection scale, sample condition, and reference baseline.
Visible imaging is suitable for characterizing overall color patterns, although the sensitivity of individual color indices varies with material and hue. Among 12 ornamental pepper materials, VARI and NGRDI distinguished green fruits more effectively, whereas BGI was more suitable for reddish-brown materials [42]. Harvest-stage studies found that ASTA color values increased from 114 to 178 for Hanbando and from 115 to 140 for Dabotop. However, Cap/Y and R/Y produced inconsistent rankings of varieties in terms of redness [41]. Color indices are therefore most appropriate for relative comparisons within a restricted range of materials. They should not be treated as unified metrics across varieties or harvest stages.
Localized abnormalities at an early stage illustrate the limitations of visible-light information. In experiments involving mechanically cut green peppers, the calibration and prediction accuracy of white-light color and texture models were below 0.4. Adding chlorophyll fluorescence increased calibration accuracy to above 0.86 and prediction accuracy to above 0.96 for several classifiers [43]. Studies using fluorescence to characterize sweet peppers and hyperspectral imaging to detect preharvest and post-harvest defects support a similar interpretation [44,45]. Complementary spectral information can reveal tissue differences that are not fully represented by visible color. However, this advantage primarily concerns early damage or latent changes. Current evidence does not support combining overall color, ripeness, and localized anomalies into a single assessment task.
Multimodal fusion supplements visible-light information rather than replacing it. For prediction of the overall color values of granulated red pepper, a convolutional neural network combined visible L*a*b* and fluorescence RGB inputs. High-level fusion outperformed low- and mid-level fusion, with a validation coefficient of determination of 0.828 and a root mean square error of 0.351 [46]. However, the study focused on a well-defined prediction task and did not classify localized damage in whole fruits. The fusion level should therefore be selected for the intended task, and these benefits should not be generalized to all visible abnormalities.
Sensor selection should match the task, acquisition timing, and deployment cost. A unified dataset should synchronize annotations of color, ripeness, and damage while controlling variety, harvest stage, and fruit morphology, thereby isolating each modality’s contribution [39,40,43,44,47].

3.4. Perception Reliability Under Pose Variation, Occlusion, and Complex Imaging Conditions

Pose variation and occlusion primarily cause the loss of information in individual views, rather than merely reducing classifier accuracy. Online detection of non-segregated potatoes shows that contact, mutual occlusion, and soil contamination impair contour-based target separation [48]. In sweet pepper studies, camera position affects fruit visibility under leaf occlusion, whereas observations from multiple positions improve capture rates [49]. Although these findings arise from conveyor and plant-based settings, respectively, both highlight the importance of viewing geometry and material configuration. High-speed color sorting also requires cross-view target association to prevent duplicate counting and identity mismatches.
Multi-view observations require robust registration and target association to produce consistent sorting decisions. NDT-6D uses color and geometric cues for six-degree-of-freedom registration in agricultural robotics, thereby improving the stability of structural matching [50]. Together with evidence that camera position affects pepper visibility [49], these findings suggest a three-stage framework: observation layout, cross-view registration, and target fusion. However, the cited studies examined agricultural-robot point clouds and greenhouse observations rather than continuous conveyor systems. Neither study validated conveyor timing, inter-target spacing, identity consistency, or synchronization with actuation. They should therefore be regarded as methodological references rather than direct evidence of online sorting performance.
Lighting variation, reflections, dust, and humidity affect color and texture through different mechanisms. These factors should not be grouped indiscriminately under the term “complex background”. An RGB–near-infrared framework for tomatoes enabled detection, segmentation, counting, and size measurement under variable illumination [51]. The results suggest that complementary spectral bands can reduce the sensitivity of RGB imaging to lighting variation. However, the study did not independently examine interference from water stains, dust, or strong reflections. Engineering experiments should isolate these factors to identify the sources of performance degradation. Such tests can guide improvements in illumination, optical design, and cleaning mechanisms.
When diurnal variation weakens color cues, thermal and depth imaging can provide complementary temperature and three-dimensional structural information. This approach has been validated for pepper detection under both daytime and nighttime conditions [52]. By contrast, RGB–near-infrared fusion aims to stabilize appearance representations, and thus addresses a different information gap [51,52]. Differences in sensor configuration and operating environment currently preclude identification of an optimal combination for reflective, dusty, or humid conditions. Improvements in detection must also be integrated with target tracking and rejection actions before they can benefit online sorting.
Current evidence supports multi-view observation, cross-view registration, and complementary sensing under pose, occlusion, and illumination variation [48,49,50,51,52] (Table 5). Benchmarks should control orientation, illuminance, and surface contamination at fixed material flow and conveyor speed, and report detection, association, and physical rejection rates.

3.5. Functional Requirements and System-Level Performance Metrics for Intelligent Color-Sorting Equipment

Intelligent color-sorting equipment integrates feeding, detection, color and ripeness assessment, grading, actuation, and traceability. Existing studies support external-appearance sorting but provide limited evidence on internal-quality sensing, actuation errors, and long-term stability [3,4,6].
Evaluation at the classification level should emphasize the separation of adjacent grades and the production of actionable decisions. Simplified sensing and logic-control systems may be suitable for a small number of categories. However, circuit- and Proteus-based approaches have not undergone multi-class online evaluation under real operating conditions [6]. Two online systems used a multi-layer perceptron and an improved ResNet50, respectively. The former achieved 93.2% accuracy, 86.4% precision, 84.0% sensitivity, and 95.7% specificity [3]. The latter reported in-line accuracy of 98.7%, precision of 97.0%, sensitivity of 96.9%, specificity of 99.0%, F1 of 96.9%, and a separately labelled overall accuracy of 96.9% for five-grade sorting [4]. The 98.7% statistic is retained under its source label because its class aggregation has not been verified; it is not interchangeable with the 96.9% overall result. Differences in datasets, class definitions, and operating conditions preclude direct comparisons of model performance. Equipment studies should consistently report grade-specific metrics, explicitly averaged F1 scores, and separate model and final-bin confusion matrices and should distinguish classification errors from actuation errors.
Real-time evaluation must distinguish algorithmic latency from system throughput. The two online systems processed individual samples in approximately 0.2 s and 4 ms, respectively. Nevertheless, both achieved an overall sorting rate of approximately 3000 samples per hour [3,4]. Differences in computational scope and hardware configuration prevent these values from providing a direct ranking of model speed. System capacity also depends on conveying, triggering, object spacing, communication delays, and actuator reset time. Simulation-based schemes that integrate detection, color assessment, and rejection mechanisms have not reported measured online throughput [6]. Algorithm-level reporting should include inference latency and computing-hardware specifications. System-level reporting should include end-to-end latency, throughput, actuation response time, and stability during continuous operation. Conveyor speed, frame rate, and rejection timing should also be specified.
Conveying and actuation mechanisms should be evaluated separately from the vision model. A small-scale mechanical potato-grading prototype achieved 94.88% accuracy and a processing capacity of 13.95 t/h after optimization [53]. The adjusted variables included slide-rail height and angle, chain speed, and belt speed. Relative to the first-generation prototype, accuracy and processing capacity improved by 3.84% and 12.94%, respectively [53]. Its grading mechanism differs from vision-based pepper rejection, so the reported values are not directly comparable. Nevertheless, the study illustrates how conveying trajectories, material orientation, and mechanical timing can affect final performance. Pepper color-sorting experiments should report the physical throughput, final-bin MRsys and FRRsys, actuation success per issued command, induced damage rate, and correctly sorted output. They should also include integrated tests of vision and actuation.
Deployment-level evaluation requires a balance among accuracy, computational demand, and adaptability to operating conditions. Fruit-grading studies have proposed multi-view and spectral imaging, lightweight models, databases, and transfer learning to improve efficiency and robustness [54]. A multi-domain, multi-scale fusion model for Sichuan pepper (*Zanthoxylum*) had a storage footprint of 5.84 MB. It achieved 98.34% accuracy for ripeness classification and a segmentation mAP50 of 88.8% [55]. Because the study concerned *Zanthoxylum* rather than *Capsicum*, these results provide only engineering reference points for lightweight deployment. They should not be used as acceptance thresholds for pepper-sorting equipment. Deployment studies should also report frame rate, memory use, computational footprint, and consistency across varieties and batches. Performance degradation under seasonal and lighting variations should be quantified.
Table 6 links functional requirements to metrics and evidence boundaries, showing why classification accuracy alone cannot represent system performance. Pepper color-sorting studies still lack unified, physical validation of perception, decision-making, and actuation. Future work should report end-to-end latency, throughput, continuous-operation stability, and failure modes under defined varieties, batches, classes, and operating environments [3,4,6,53,54,55].

4. Technological Evolution and System Architecture of Color-Sorting Equipment

Post-harvest pepper sorting depends on the complete sensing-to-actuation chain, including cultivar, appearance, contamination, orientation, occlusion, and conveyor dynamics. This section compares human-integrated judgement, rule-based photoelectric or 2D vision, and data-driven intelligent vision through their decision logic, information flow, and closed-loop capability. Capsicum evidence is prioritized, while related-crop results are interpreted under the transferability criteria in Section 2.4.

4.1. Human-Integrated Color Perception and Experience-Based Decision-Making

Manual color sorting requires operators to integrate information about color, shape, and surface condition when assigning grades. This approach allows operators to handle some atypical samples. However, differences among evaluators, operator fatigue, and batch variation can cause grading scales to drift [56]. The central limitation is that human-integrated color perception is difficult to quantify, verify, and transfer consistently. Translating such judgments into equipment rules therefore requires integrated visual perception to be decomposed into measurable variables with reproducible grading boundaries.
Early prototypes formalized experiential judgments using HSV color, grayscale thresholds, projected area, and related size proxies in tomato, grape, and mango systems [56,57]. These systems demonstrated rule-based actuation but remained sensitive to illumination, crop geometry, and feeding conditions. They therefore show how grading experience can be formalized, not how manual-sorting performance should be quantified. Because human judgments also provide training labels and acceptance benchmarks, pepper studies should report evaluator number, inter-evaluator agreement, repeatability error, and grading-boundary uncertainty across cultivars and batches [56,57,58,59]. Unstable reference labels can cause high fitting accuracy to reproduce an evaluator- or batch-specific scale. Table 7 summarizes the functional evidence.

4.2. Rule-Based Photoelectric and 2D Machine Vision

Rule-based photoelectric and 2D machine-vision systems convert controlled acquisition, handcrafted color or size features, and predefined rules into mechanical sorting [62,63]. Camera systems use RGB, HSV, CIE XYZ, or CIE L*a*b* representations, whereas dedicated RGB sensors shorten the control chain. One apple device demonstrated laboratory feasibility but was not tested under complex production conditions [63]. Another required 260 ms per fruit in software, whereas the full line processed 96 fruits per minute [62]. Thus, throughput reflects acquisition, conveying, and actuation as well as inference. Systems using learned grade mappings mark the transition to data-driven vision [64,65]. Color measurements alone cannot establish comprehensive quality, and evaluation should link sensing to final-bin errors, throughput, and perception-to-actuation latency.

4.3. Data-Driven Intelligent Machine Vision

Data-driven intelligent vision extends color classification to localization, multi-attribute grading, instance separation, and action planning, linking perception with mechanical handling [57,66,67,68,69,70]. Equipment-level performance nevertheless remains constrained by illumination, occlusion, target geometry, and actuator coupling.
Capsicum studies demonstrate real-time recognition at the perception layer [67], while related-crop studies address adhered-object separation and illumination generalization [69,70]. However, these results do not verify continuous feeding, identity-preserving tracking, timed diversion, or final-bin outcomes on chili pepper lines. Section 5 compares recognition models. Here, the evidence supports perception feasibility but not complete-system validation.

4.4. Comparison of Sorting Principles, Processes, and Performance Across Technical Paradigms

The paradigms differ in decision formation and coupling to physical sorting. Human-integrated judgement is flexible but difficult to standardize, while rule-based 2D vision is repeatable under controlled acquisition. Data-driven vision expands information coverage and object localization but increases data and deployment demands [60,61,71,72], as summarized with the other functional trade-offs in Table 7.
System-level and model-level values in this evidence base describe different endpoints rather than a common performance scale. The integrated apple system linked three-camera acquisition to physical grading [72]. It reported 95.76% to 99.45% for single features, a mean multi-feature accuracy of 95.49%, and a field grading accuracy of 94.12%. Its lower field result and 1.2 s cycle coincided with full-line execution and a bottom-surface blind area.
Multispectral apple responses overlapped across most cultivars, showing that additional channels do not automatically resolve cultivar variation [61]. SB3D-NET reported 95.54% training and 90.74% validation accuracy across five soybean cultivars, but not system-level throughput [71]. The apparent differences across these studies therefore reflect crop morphology, cultivar separability, sensing coverage, validation design, and evaluation endpoint rather than technical superiority.
Additional sensing or model complexity is justified only when it improves system endpoints under stated operating conditions, including surface coverage, final-bin accuracy, MRsys, FRRsys, latency, throughput, damage, and sustained stability (Figure 4 and Figure 5).

4.5. Overview of Intelligent Color-Sorting Equipment: Components, Information Flow, and Closed-Loop Architecture

Intelligent color-sorting equipment integrates material presentation, perception, grade determination, timed execution, and outcome verification within one electromechanical information system [73,74,75,76,77].
Synchronized material and information flows convert each target’s grade and location into an actuator command and verified destination. Reported platforms place acquisition, fusion, inference, communication, and actuation locally, at the edge, or centrally. Their engineering value depends on coupling sensing, delay estimation, and mechanical motion. Figure 6 illustrates representative pneumatic execution mechanisms, while Section 7 examines interfaces and feedback.
Local platforms minimize sensing and communication overhead for stable, single-attribute tasks but offer limited tolerance to appearance variation and online verification [73,74,75,76]. Edge processing shortens network paths while increasing local resource, thermal, and maintenance demands [76]. Centralized multimodal systems improve cross-sensor synchronization but add calibration, transfer, and cycle-time costs [77].
Execution timing remains a system-level constraint because command delay and material misalignment can separate vision grading from final sorting. Diagnosis therefore requires synchronized command, actuator, and target-motion records [73].
Most reported systems remain quasi-closed-loop because item-level destination outcomes are evaluated offline [73,74,75]. A complete closed loop requires end-of-line sensing, traceable target identities, and synchronized outcome records for delay correction, threshold adjustment, model maintenance, and fault diagnosis (Figure 7).
Architecture selection should follow operating constraints, with local platforms favoring short response paths under limited sensing requirements. Centralized multimodal platforms favor information coverage when calibration and cycle time permit. These trade-offs must be validated under the intended illumination, cultivar, throughput, and actuator conditions [73,74,75,76,77].
Across these architectures, performance variation can be traced to two coupled bottlenecks. Variable appearance affects perception, while communication, timing, alignment, and actuator response affect physical execution. Pepper validation must therefore connect inter-batch recognition stability, target association, command timing, and final-bin verification within the same operating condition.
A practical design should specify the target attribute, required information, available cycle time, and acceptable error consequence before selecting sensors or actuators. This sequence prevents equipment complexity from increasing without a measurable improvement in sorting performance.

5. Machine Vision Perception and Chili Color Recognition Technologies

5.1. Image Acquisition, Illumination Control, and the Optical Imaging Environment of Color Sorting Equipment

Front-end image quality constrains recognition because cameras, lenses, illumination, sample pose, optical geometry, exposure, and calibration jointly determine color fidelity, reflections, motion blur, and timing. Information lost through poor exposure, spectral gaps, or motion blur is difficult to recover downstream. Visible-light imaging rapidly captures surface color, shape, and position for conveyor-based inspection [24], whereas active infrared imaging detects subsurface tissue responses but requires more complex optics, calibration, and acquisition [78]. The two routes are complementary in function rather than interchangeable in system design.
Fluorescence imaging is an active-excitation technique. Ultraviolet-induced fluorescence varies among sweet pepper varieties, and the corresponding macroscopic and microscopic features are not always consistent. When combined with machine learning, these tissue responses can be used to identify damage in green peppers [43,45]. However, the signals are affected by variety, pigmentation, and tissue condition. Direct application of RGB color-difference thresholds is therefore inappropriate. In practical systems, the excitation source, optical filters, exposure settings, and calibration procedure must be carefully controlled. Ambient light must also be effectively isolated.
Tissue absorption and scattering determine the penetration depth of active optical signals. Previous studies have defined the “effective penetration depth” as the depth at which incident intensity decreases to approximately 37% or 1% of its initial value. However, the measured depth also depends on light intensity, incident angle, source–detector distance, and tissue structure [79]. Increasing the power of the light source therefore does not necessarily provide reliable information from deeper tissue layers. Compared with visible-light surface imaging, subsurface detection requires more complex optical systems, stricter calibration, and longer acquisition times.
The usefulness of an additional imaging channel depends on the specific task. Digital images and laser-scattering signals have been used to predict moisture content and color during sweet pepper drying. In that application, yellow peppers showed a stronger correlation with moisture content than red and green peppers [80]. This finding indicates that signals associated with a particular color or processing condition may not transfer directly to the online sorting of fresh peppers. Front-end channels should therefore be selected according to predefined grades and timing requirements, rather than by simply increasing the number of signal types.
Visible-light imaging enables rapid conveyor-based surface inspection but has limited subsurface sensitivity [24]. Fluorescence and active infrared imaging reveal tissue condition but require controlled excitation, optical isolation, and calibration [43,45,78]. Channel selection should balance color fidelity, anomaly visibility, pose tolerance, exposure, signal-to-noise ratio, and cycle time. Shorter exposure reduces motion blur but may lower signal quality, whereas longer acquisition restricts the sorting cycle. Reported gains must therefore be interpreted under the stated illumination, calibration, material pose, and acquisition speed [24,43,45,47,78,79].

5.2. Traditional Methods Based on Color Spaces, Color Differences, Color Ratios, and Thresholding

Traditional color-based classification generally comprises four sequential stages: color representation, target or defect segmentation, explicit feature extraction, and rule-based classification. Common features include channel statistics, color ratios, texture, shape, and defect-region descriptors. Decisions are then made using thresholding, statistical classifiers, or nearest-neighbor rules [34,81]. These methods are interpretable and computationally efficient. However, errors can accumulate throughout the processing pipeline, and performance depends strongly on imaging consistency and segmentation rules.
Under controlled background conditions, RGB and HSV features supported a five-category chili classifier trained on 210 images and tested on 90 images [34]. The reported precision, recall, and accuracy were 1.0, corresponding to image-level overall accuracy of 90/90, although the averaging convention was not reported. The available description does not establish fruit-level independence, class balance, or external-batch testing. If multiple views of one fruit cross dataset subsets, image-level partitioning can overestimate generalization. The result therefore demonstrates separability within the reported closed setting, not transfer across cultivars, batches, or production conditions.
CIELAB coordinates, chromaticity measures, and color-difference parameters provide a more standardized description of color. They characterize lightness, red–green and yellow–blue components, saturation, and differences between samples [82]. Under different LED lighting conditions, chili peppers were evaluated using L*, a*, b*, C*, H°, and ΔE*ab. Combined red and blue lighting produced higher ASTA color values and chromaticity, whereas blue light promoted capsaicin accumulation. These results indicate that production conditions can influence both measured color and composition. Standardized color spaces can help define interpretable grade boundaries. Nevertheless, recalibration remains necessary for different light sources, camera responses, varieties, and ripeness stages.
Color ratios are calculated from normalized channel values or ratios between channels and may partially compensate for variations in overall brightness. The studies discussed here evaluated RGB statistics, multispectral features, CIELAB coordinates, and color-difference parameters. However, they did not directly validate the transferability of color ratios across imaging conditions. Color ratios should therefore be treated as candidate features rather than inherently robust indicators. Their performance should be compared with that of raw channels and standardized color-difference measures under identical conditions.
Thresholding is commonly used for target or defect segmentation, but it cannot independently determine product quality. In chili classification, RGB images were first converted to HSV. Otsu thresholding and morphological opening were then applied to the saturation channel before color and shape features were classified using KNN [34]. Similarly, dual-color apple grading used a cascaded pipeline comprising initial segmentation and subsequent screening based on statistical, texture, and geometric features [81]. Both approaches relied on a limited set of explicit features. Consequently, segmentation errors propagated to subsequent classification stages. More complex downstream classifiers may not fully compensate for changes caused by illumination or background drift in the front-end segmentation rules.
The five-category chili result and the 93.5% apple defect-classification result were obtained for different crops, class definitions, sample structures, and decision tasks [34,81]. Their numerical difference cannot isolate an algorithmic effect. It instead indicates that traditional methods must be interpreted within the imaging, labeling, and partitioning conditions of each study.
Traditional methods are inexpensive, interpretable baselines for color calibration and lightweight classification under stable imaging and well-defined grades. Their thresholds and handcrafted features require recalibration when illumination, background, pose, or data distributions shift. Comparisons should therefore use matched samples and fixed partitioning, illumination, conveyor, latency, and final-bin criteria.

5.3. Machine Learning- and Deep Learning-Based Recognition of Ripeness and Color Grades

Performance across maturity and color-grade studies should be compared by class definition, sample diversity, partitioning, input information, output granularity, and model complexity [9,10,11,39,40,83,84,85]. Interpretable indices suit stable grades and imaging, whereas sensor fusion and deep models can address ambiguous color, localization, and complex backgrounds. These gains require additional data, calibration, computation, and deployment resources (Table 8). As cross-crop evidence, CAM-YOLO integrates a convolutional block attention module into YOLOv5 for tomato detection and ripeness classification, including overlapping and small targets [86]. Its transfer to pepper sorting requires validation under pepper-specific imaging and conveying conditions.
Validation should preserve fruit identity across partitions and use independent batches. Reports should specify image-, fruit-, batch-, origin-, cultivar-, and device-level splits because continuous ripening and cultivar or batch variation can blur adjacent commercial grades. Class-specific confusion is therefore more informative than overall accuracy alone.

5.4. Methods for Detecting and Classifying Visible Color Abnormalities

Visible color abnormalities can be assessed through image classification, object detection, semantic segmentation, or instance-level processing [91,92,93,94,95,99]. The appropriate output depends on whether the actuator needs only a fruit-level decision or also the abnormality location and area.
Image classification supports low-cost screening but cannot localize defects. Detection provides actionable positions, while segmentation quantifies irregular regions and severity. These gains require progressively greater annotation and computational effort [91,92,94,95,99]. A related methodological example is YOLOV8-CMS for citrus leaves, which combines disease classification with segmentation-based grading using the lesion-to-leaf area ratio [100]. Transfer to post-harvest pepper inspection would require fruit-specific defect labels and validation.
Representative benchmarks show that the model version or the presence of additional spectral bands alone does not determine practical value [92,94]. Active learning may reduce annotation effort, but the deployment benefit still depends on output granularity and batch-level generalization [93].
Mango image classification reported about 98.5% cross-validation accuracy and an AUC of 0.98 [91]. A pear benchmark reported mAP@0.5 values of 63.6% to 71.2% across 27 detectors [92]. Apple segmentation produced F-scores of 0.7938 for RGB and 0.7888 for selected RGB-NIR bands [94]. These values do not form a common ranking because classification accuracy, detection mAP, and segmentation F-score measure different units and error structures. Their variation reflects task granularity, class separability, annotation form, and metric definition in addition to model performance (Figure 8).
Method selection should therefore follow the downstream sorting decision and required output granularity. Classification supports fruit-level screening, detection provides actionable positions, and segmentation quantifies abnormality extent [91,92,94,95,99]. These outputs should remain linked to target identity and confidence. Pepper datasets should distinguish natural color variation, transitional ripening, and actionable defects while recording variety, batch, origin, illumination, viewpoint, and minimum detectable area [43,91,92,93,94,95,96,99,101].
Active learning can reduce annotation effort, but this benefit should be reported separately from recognition and line-level outcomes [93]. Integrated trials should report abnormality recall, minimum detectable area, per-object latency, MRsys, and FRRsys, linking perception to final-bin costs.

5.5. Multispectral and Multimodal Color Perception Methods

Multispectral and multimodal sensing is valuable when additional channels improve task-specific contrast without excessive acquisition, registration, or computational cost [44,47,102,103]. Pepper studies show spectral separability for defects, maturity, and impurities, while selected bands can reduce hardware and data loads [10,44,47,97].
Visible, near-infrared, and fluorescence signals can complement RGB for weakly visible tissue states [39,40,46]. Benefits may be offset by cultivar overlap, channel redundancy, registration error, noise, limited samples, calibration, latency, and maintenance [61,94,98]. RGB should remain the baseline, with added channels justified by measured operational gain. In jujubes, hyperspectral defect maps were combined with curvature-error correction, stem-end identification, and market-specific rules to determine retention, rejection, and appearance grades [104]. This cross-crop study illustrates the additional processing needed to convert pixel-level classifications into fruit-level sorting decisions.
No current framework covers all three tasks, so a common-batch line evaluation should compare RGB, selected-band, full-spectrum, and fusion routes and report grade, position, confidence, and localized abnormalities.

5.6. Robustness, Generalization, and the Generation of Target Information for Color Sorting in Complex Scenarios

Robustness, cross-domain generalization, and uncertainty identification address distinct levels of model failure [105,106]. Intra-domain robustness concerns variations in occlusion, reflection, object scale, and background. Cross-domain generalization measures performance degradation across varieties, origins, production batches, equipment, and illumination conditions. Uncertainty identification addresses low-confidence predictions and out-of-distribution inputs. High accuracy on a single dataset indicates performance only within the evaluated domain and cannot replace cross-condition validation.
Architectural optimization can improve detection performance within a given data domain. In a four-category citrus disease detection task, deformable convolutions, feature pyramids, attention mechanisms, and loss-function optimization improved multiple performance metrics [107]. Precision, recall, and mAP@0.5 increased from 94.5%, 92.1%, and 93.2% to 98.7%, 95.9%, and 97.7%, respectively. These changes corresponded to gains of 4.2, 3.8, and 4.5 percentage points. The results indicate that architectural modifications can improve detection in complex backgrounds. However, validation was confined to the same task and data domain. Comparable improvements therefore cannot be assumed for chili peppers from different varieties, origins, or production batches.
When errors arise primarily from domain shifts associated with crop species or imaging scenes, data-level domain adaptation may provide a more direct solution. One study transferred an orange detector to apples and tomatoes using image translation and pseudo-label-based self-supervised learning [108]. The method increased mAP from 65.3% to 87.5% for apples and from 71.1% to 76.9% for tomatoes. These changes represented improvements of 22.2 and 5.8 percentage points, respectively. The gains depend on similarities between the source and target domains in shape, scale, and foreground–background characteristics. Moreover, pseudo-label thresholds still require manual selection, and uncertainty under cross-domain conditions remains unresolved.
Architectural optimization and domain adaptation address different sources of error [107,108], but neither establishes cross-cultivar, cross-batch, or cross-device robustness for chili sorting. Validation should distinguish controlled within-domain variation, independent external domains, and continuous conveyor operation (Figure 9).
Perception outputs should provide only the fields required downstream: target identifier, position, color grade or abnormality class, and confidence or uncertainty [3,4,5,11]. Timestamps, trajectories, validity status, and system traceability belong to the coordination and integration interfaces discussed in Section 6 and Section 7. Current evidence does not establish generalization across pepper varieties, origins, batches, devices, and illumination conditions, so these domains require external validation.

6. Temporal and Spatial Coordination of Material Conveyance, Target Localization, and Pneumatic Rejection

Vision outputs support sorting only when material presentation, localization, timing, actuation, and outcome verification are coordinated along the feeding-to-rejection sequence.

6.1. Feeding, Spreading, Flattening, and Posture Regulation

Front-end handling should produce separable, observable targets because image count, visible surface area, contour quality, and pose stability constrain downstream evidence [109,110,111,112]. In one multichannel prototype, increasing flow from one to three targets per second reduced images per fruit from 24 to 9 [109]. This illustrates the trade-off between throughput, surface coverage, and separation (Figure 10).
Material pose governs visible surface coverage. Single-fruit placement maximizes separation but provides one view; controlled single-layer feeding limits overlap; and rolling increases coverage at a cycle-time cost. Random piling instead increases occlusion, reflection, and identity ambiguity [109,110,111,112]. These effects precede model selection.
The 120 controlled apple images and chili shape model address different tasks, so their accuracies are not directly comparable [110,112]. The former depends on surface visibility, whereas the latter does not establish singulation or multi-view coverage.
Front-end optimization should therefore maximize relevant surface visibility while recognition models accommodate residual variation (Table 9).

6.2. Surface Cleaning, Material Conditioning, and Imaging Standardization

Material conditioning should establish a state suitable for reliable imaging. Spectral, multimodal, and thermal methods can screen moisture, ripeness, freshness, and cold-related changes [40,113,114,115,116], but direct evidence for online mechanical cleaning remains limited.
A controlled test should vary dust load, cleaning airflow, surface moisture, and conveyor speed on the same line. It should measure residual contamination, post-cleaning image contrast, classification stability, final-bin errors, airflow demand, and conditioning-induced damage.

6.3. Conveyor Speed, Material Spacing, and Multi-Channel Processing Capacity

Conveyor design is governed by the joint time budget for acquisition, computation, communication, localization, and actuation [3,4,5,6,24,117,118,119,120].

6.3.1. Conveyor Speed Has an Operating-Condition-Dependent Effective Range

Conveyor speed affects performance through four coupled paths: fewer views per item, greater motion blur, shorter inter-arrival intervals, and less actuator recovery time [109,119,120]. Material size, contact mode, inclination, and grading geometry shift the feasible range [120]. A reported speed is therefore one operating point on a speed-accuracy-throughput curve.
Parallel processing and edge inference can shorten computation, but they cannot restore lost surface coverage or actuator recovery time [118,119]. Multichannel designs increase capacity only until shared cameras, controllers, communication links, conveyors, or air supplies become bottlenecks [117,118,119,120].

6.3.2. Material Spacing Translates Spatial Configuration into a System-Level Time Budget

When materials pass sequentially through a single channel, the inter-arrival interval can be expressed as Δt = s/v. Here, s denotes target spacing, and v denotes conveyor speed. At a spacing of 100 mm and a speed of 0.6 m/s, the ideal inter-arrival interval is approximately 0.167 s. This interval corresponds to an ideal arrival rate of approximately six items per second [119].
This value is a time budget derived from the spatial configuration, rather than a measured end-to-end throughput. The spatial configuration becomes an effective processing capacity only when imaging, transmission, classification, communication, and actuation are all completed within this interval.
Multiple cameras improve coverage, multiple processors reduce local queues, and multiple material channels increase physical capacity [117,118,119,120]. Shared conveyors, controllers, communication links, or air supplies can nevertheless become system bottlenecks.
End-stage re-inspection can audit contaminants or misrouted grades when reference labels are available [117]. It complements, but does not replace, synchronized records of belt speed, spacing, target identity, and channel-level output.

6.3.3. Multi-Channel Processing Capacity Depends on the Technical Route and Shared Bottlenecks

Effective multichannel capacity is a stable output after shared-resource occupancy and recovery time; material lanes, imaging resources, computing nodes, and actuator channels are reported separately.

6.3.4. System-Level Discrimination Framework for Intelligent Chili Color Sorting Equipment

A minimum report should state conveyor speed (v), target spacing (s), Δt = s/v, channel count (Nch), processing time (Tproc), and end-to-end latency (Tsys). The complete chain must fit the arrival interval.

6.4. Target Localization, Identity Maintenance, and Arrival Time Prediction

Online localization combines target detection, identity maintenance, coordinate transformation, and arrival-time prediction [121,122,123,124]. Its purpose is to convert a visual object into an executable state for the controller.

6.4.1. Target Localization: From Detection to Executable Position Generation

Localization technologies provide different state variables and should not be ranked by one accuracy value. Tracking preserves identity continuity, detection supplies two-dimensional position, RGB-D sensing adds three-dimensional action points, and motion estimation predicts arrival time [121,122,123,124].
A representative tracking pipeline reported MOTA of 93.6%, MOTP of 85.5%, and 23 ms per frame for 8 to 20 targets [121]. These metrics quantify identity and position estimation within that pipeline. They are not equivalent to RGB-D localization error or final-bin rejection success.

6.4.2. Trajectory Prediction: Divergence Between Explicit State Estimation and Implicit Motion Assumptions

Motion handling ranges from explicit velocity estimation to continuous detection and data association [121,123]. Fast single-frame recognition is sufficient only when material motion is tightly constrained and execution remains observable [122].
Infrared tracking addressed weak-light identity continuity and reported MOTA of 0.85 while counting 67 of 70 targets [123]. RGB-D and grasping studies instead provide geometric action points [122,124]. The apparent performance differences therefore reflect distinct outputs and operating conditions rather than interchangeable localization accuracy.

6.4.3. Multimodal Perception: Trade-Offs Among Environmental Adaptability, Calibration, and Latency Cost

Infrared tracking is applicable when visible contrast is weak, whereas RGB-D sensing is valuable when three-dimensional geometry is required [123,124]. Both add calibration, synchronization, reconstruction, hardware, and latency costs that must fit the conveyor time budget.
Integrated evaluation should report localization error, identity switches, arrival-window error, command latency, and final rejection success on the same conveyor and material set.

6.5. Pneumatic Jet Mechanisms, Nozzle Arrangement, and Valve-Control Response

Pneumatic rejection uses zonal arrays, pulsed jets, or continuous suction-based diversion [125,126,127,128,129], which prioritize spatial coverage, transient impulse, and flow-field stability, respectively. Existing prototypes characterize array throughput, pulse-nozzle force, and suction-field routing under specific configurations [125,126,127]. Individualized pulses favor well-separated targets requiring positional precision; zonal arrays broaden coverage with coarser control; continuous diversion suits dense or posture-coupled flows. Route selection remains conditional on target mass, pose, density, spacing, speed, pressure, and acceptable damage. These factors govern impulse, nozzle interference, trigger-window width, displacement, and damage (Figure 11).
End-to-end delay should be partitioned into command-to-valve, valve-to-effective-jet, and jet-to-target-displacement intervals, including adjacent-nozzle interference. Switching specifications should therefore be linked to measured jet rise, decay, and target displacement before entering the conveyor timing budget [127].
Actuation can be selected according to target density, required precision, and acceptable damage. Pepper-specific tests should quantify the final displacement and induced damage [128,129,130,131,132], while contact-based studies provide transferable metrics for success, force, damage, and operation time [128,129]. Compliant mechanisms also offer force-control principles that should guide non-contact rejection tests.

6.6. Perception–Execution Timing Alignment, Diversion Accuracy, and Material-Damage Control

System errors propagate from perception and timing to execution and final material state; predictive tracking, actuator compensation, closed-loop robotic perception, and compliant mechanics address different links [133,134,135,136,137]. Evaluation should combine future-state prediction with measured actuator delay and distinguish recognition errors, timing mismatches, unreachable poses, unstable contact, and release-induced damage [135,136,137]. A unified test should report final placement, localization uncertainty, end-to-end latency, task-completion rate, effective throughput, and material-damage rate. Pepper evidence identifies compliance, rotational-to-travel-speed matching, and mechanical impact as variables coupling speed, impact, and damage [137].

6.7. Summary: From Component Performance to System-Level Evaluation

Together, observable presentation, persistent identity, bounded timing error, controlled actuation, and verified outcomes link component performance to the integration analysis in Section 7.

7. Vision–Execution Closed Loop and Intelligent Color-Sorting System Integration

7.1. Interface Design for Vision, Conveying, Control, and Pneumatic Units

Closed-loop integration requires a shared object-event record that preserves target identity across sensing and actuation. The record should contain the target ID, timestamp, execution status, and final destination. These fields allow recognition, timing, and actuation errors to be traced [3,4,48,50,125,130,131,133,138] (Table 10).

7.2. Real-Time Coordination of Color Recognition, Communication, Localization, and Rejection

Real-time coordination requires all modules to share one clock and spatial reference. Total latency must remain within conveyor travel time.
RGB and spectral routes trade acquisition complexity against separability [9,39,40,47,97,143]. Matched tests should assign one target identifier and reconstruct the coordinate error, trigger delay, and final-bin outcome [50,133,144,145].

7.3. Edge Computing, Model Deployment, and Mechatronic–Optical System Integration

Deployment ranges from mobile-server assistance to edge-cloud processing and fully local control [4,140,141,142]. Moving computation closer to the actuator reduces network exposure but increases local resource, thermal, and maintenance demands.
Deployment should be selected by the complete timing and resource budget rather than isolated inference time. Comparable latency reporting requires the processor or GPU, input resolution, batch size, network transfer, preprocessing, inference, postprocessing, and the actuator-communication boundary. Unreported fields are indicated by NR in Table 11b.
Reported values include 0.859 s for a mobile-server response and 40 ms per 1000 rows for an edge model [140,141]. Local routes reported 12.3 ms per image segment and about 4 ms per sample in an online sweet-pepper system [4,142]. These values use different workload units and architectures. They identify local timing constraints but do not establish a hardware speed ranking.

7.4. Evaluation of Sorting Accuracy, System Missed-Rejection Rate, System False-Rejection Rate, and Throughput

For binary retain/reject sorting, the positive class comprises fruits requiring rejection under the reference quality criteria; qualifying fruits form the negative class. System counts are determined by final physical destinations. TPsys and FNsys are rejection-required fruits reaching the reject and retain bins, respectively; FPsys and TNsys are qualifying fruits reaching those bins. The following binary metrics assume that every evaluated fruit is assigned to one of these two destinations.
A c c u r a c y s y s = T P s y s + T N s y s T P s y s + T N s y s + F P s y s + F N s y s
M R s y s = F N s y s T P s y s + F N s y s
F R R s y s = F P s y s T N s y s + F P s y s
Q s y s = N o u t t
For throughput, Nout is the number of individual fruits crossing the combined final-output boundary during elapsed time t, regardless of sorting correctness. Mass throughput uses the corresponding output mass divided by t. Report the observation window, downtime treatment, channel count and units. Correctly sorted throughput Qcorrect counts correctly assigned fruits per unit time, whereas Qreject,correct counts only correctly rejected fruits and depends on the incoming defect prevalence.
Input and output counts must be reconciled. Lost, uncollected, unresolved and manual-review items must be reported separately with their true classes where available; they must not be silently excluded from the input total. If the two-bin assumption is unmet, report outcome coverage and label the binary rates as conditional on resolved destinations. Ratios with a zero denominator are undefined, not zero.
Model counts use predicted labels rather than final-bin destinations. With the same positive class, model recall divides TPmodel by all actual positive samples (TPmodel + FNmodel), whereas model precision divides TPmodel by all predicted positive samples (TPmodel + FPmodel). Thus, 1 − Recallmodel equals the model false-negative rate; 1 − Precisionmodel is the false-discovery proportion among predicted positives. The model false-positive rate divides FPmodel by all actual negative samples (TNmodel + FPmodel). The system missed-rejection rate can additionally reflect tracking, timing and actuation failures. It equals 1 − Recallsys only when recall uses the same final-bin counts. The model false-negative rate and system missed-rejection rate differ in their evaluation endpoints; the false-discovery proportion and false-rejection rate differ in their denominators. Detection studies must also state the target class, counting unit, matching rule and confidence threshold.
For K mutually exclusive grades, overall accuracy should be reported separately from class-specific and averaged metrics. Let Cij denote the number of items of true grade i assigned to grade j, with the assignment endpoint identified as either a model prediction or a final output bin. For a complete K-class assignment, overall accuracy is defined below.
O A = k = 1 K C k k i = 1 K j = 1 K C i j
For one-versus-rest evaluation, each named grade k is treated as positive and all other grades as negative. Per-grade recall divides correctly assigned members of grade k by all true members of that grade; per-grade precision divides them by all assignments to that grade. Macro-averaging gives each grade equal weight, whereas support-weighted averaging weights grades by their true sample counts. Macro-F1 averages the grade-specific F1 scores and need not equal the harmonic mean of macro-precision and macro-recall. Balanced accuracy denotes macro-recall here. Averages across features, folds or runs must be identified separately. Multi-grade assignments can be collapsed into binary MRsys and FRRsys only after the rejection-required grades are explicitly mapped to the positive class.
The complements of reported recall and precision are 5.91% and 5.54% for the three-grade apple model [147]. The corresponding arithmetic complements are 5.1% and 3.9% for the four-grade fingered-citron detector [148]. Because target classes, counting units, averaging conventions, and evaluation endpoints differ, these values are not comparable binary miss or false-discovery rates. They are also not estimates of final-bin MRsys or FRRsys. The same restriction applies to cross-study accuracy comparisons throughout Section 4, Section 5, Section 6 and Section 7.
Available evidence separates model evaluations from continuous-conveying and final-bin system tests [4,146,147,148,151]. Only the latter directly connects recognition, material handling, and physical assignment.
Table 11a preserves source-defined system endpoints, while Table 11b records the conditions that limit their comparability. Model frame rate, physical throughput, and final-bin accuracy should be reported separately because workload, channel count, spacing, hardware, and actuation timing differ across studies [4,146,147,148,151]. Closed-loop tests should reconcile all input and output items, retain unresolved outcomes, and report separate model-level and final-bin confusion matrices under stratified sample conditions.

7.5. Evaluation of System Latency, Energy Consumption, Material Damage, and Reliability

7.5.1. Evaluation Boundaries and Indicator Hierarchy

System latency, throughput, energy, damage, and reliability must be evaluated at the same operating point.

7.5.2. System Latency, Speed Effects, and Throughput Capacity

Higher conveyor speed can raise capacity but narrow sensing and actuation margins [147,149]. Repeated trials indicate short-term stability only; sustained performance requires continuous-operation and failure-recovery evidence [150].

7.5.3. Energy Consumption and Material Damage: Limited Engineering Quantification

Published energy data are incomplete because lamp ratings describe optical load but exclude conveying, computing, pneumatics, cooling, and auxiliary systems [150]. Consequently, a low lamp rating does not demonstrate low system energy consumption, particularly when throughput and operating time differ.
Energy and damage must be interpreted at the same operating point. Energy per unit mass depends on total system power and sustained throughput, while conveying, collision, ejection, and drop impact can create damage [4,43,73,125,131,149,150]. Reporting only one dimension can therefore favor a configuration that transfers cost to another part of the system.
Existing pepper damage studies show that minor surface defects can be detected [43], but they do not measure damage created by the sorting line. Production tests should inspect skin, stem, and internal tissue before and after conveying and rejection.

7.5.4. Reliability: The Gap Between Recognition Stability and Equipment Availability

Short trials of recognition and actuator repeatability cannot establish equipment availability. Reliability evidence should instead report continuous operating time, uptime, failure frequency, recovery time, throughput, and environmental range [4,149,150,152].

7.5.5. Evidence-Based Integrated Evaluation and the Absence of Cross-System Rankings

No study jointly reports final-bin accuracy, throughput, end-to-end latency, energy per unit mass, induced damage, and operating time. Missing dimensions preclude defensible cross-system ranking. Comparable evaluation requires these outcomes, operating faults, and recovery to be reported under aligned sample, illumination, conveyor, hardware, and system-boundary conditions.

7.6. System Observability, Data Traceability, and Production-Line Validation

7.6.1. From Output-Based Assessment to System Observability

Observability, traceability, and line validation form one chain. An observable system should retain acquisition conditions, model outputs, actuator status, and final outcomes to distinguish sample variation from system failures [153,154,155,156].

7.6.2. From Individual Predictions to Traceable Data Chains

Traceability should preserve a replayable chain from raw data and model version to command and final destination, including cross-view correspondence [153,154,155].

7.6.3. Hierarchical Deployment Models for Production Lines

Batch screening, online control, anomaly diagnosis, and contaminant monitoring each require a defined target scale, decision threshold, and error consequence [153,154,155,156].

7.6.4. Comparison of Evidence from Representative Studies

Table 11b consolidates the principal sources of cross-study performance variation identified across Section 4, Section 5, Section 6 and Section 7. It preserves source-defined endpoints and marks unreported information as NR, thereby separating technical effects from differences in experimental design.
Sample and label heterogeneity is the first source of variation. Restricted cultivars, sharply separated classes, small datasets, and image-level random splits can produce high within-domain performance. Independent fruit, batch, cultivar, origin, or device partitions impose broader distribution shifts. Classification, detection, segmentation, tracking, and final-bin sorting must also remain separate because their labels, counting units, and error consequences differ [34,61,71,72,91,92,94,107,108].
Imaging and scene heterogeneity forms the second source. Controlled illumination and fixed poses reduce color drift, reflection, occlusion, and motion blur [24,47,72]. Active excitation, multi-view acquisition, and continuous conveying add calibration and timing constraints [43,45,109,119,123]. A sensing gain is therefore transferable only when the target contrast persists across cultivars, poses, and operating speeds.
System and deployment heterogeneity forms the third source. Local inference can shorten communication paths, but hardware resources, queues, thermal load, and actuator recovery still constrain sustained throughput. Centralized or multimodal platforms expand information coverage but add synchronization and calibration costs [4,72,119,121,140,141,142,146,147,148,149,150,151]. Frame rate, inference latency, material throughput, and final-bin accuracy are comparable only when workload, channel count, spacing, hardware, and timing boundaries are aligned.
These boundaries define the applicable conditions of the principal technical routes. Rule-based RGB methods are appropriate for stable illumination and reproducible grade thresholds. Deep models better accommodate complex appearance and localization but require fruit-level and independent-batch validation. Multispectral or multimodal systems are justified when added channels improve a predefined operational endpoint enough to offset calibration, latency, hardware, and maintenance costs. Tracking, RGB-D sensing, and pneumatic routes must additionally satisfy the target-density, pose, timing, and damage constraints of the intended production line.

7.6.5. Research Gaps and a Verifiable Pathway for Production-Line Validation

The literature does not yet establish a unified diagnostic protocol for chili color sorting. The proposed framework therefore classifies failures as acquisition, recognition, communication or timing, and actuation events, each linked to observable variables.
Validation should progress from acquisition stability to model performance, spatiotemporal coordination, and final-bin outcomes. The final stage should jointly report MRsys, FRRsys, throughput, energy, damage, faults, and sustained availability under stated production conditions. This factor-based evidence map provides the basis for the system-level challenges examined in Section 8.

8. Key Challenges and Development Trends in Intelligent Color-Sorting Equipment for Chili Peppers

Section 5, Section 6 and Section 7 discussed color perception, material conveying, and the vision–execution closed loop in chili pepper sorting. This section examines the remaining evidence gaps at the overall system level. The practical value of such equipment depends on whether its components can accommodate variations in illumination, cultivar, batch, and processing load. The discussion addresses six key challenges: color robustness, cross-domain generalization, vision–execution synchronization, multi-objective optimization, edge deployment, and the industrial application of multispectral imaging. For each challenge, evidence from the literature, cross-study inferences, and engineering recommendations are organized according to four elements: available evidence, applicability boundaries, issues requiring further validation, and engineering criteria.

8.1. Robustness of Color Recognition Under Complex Illumination and Operating Conditions

Stable illumination is essential for the machine-vision inspection of fruits and vegetables. Comprehensive studies of bell peppers have shown that surface curvature and viewing angle can produce substantial reflection artifacts. Local reflectance values may even exceed those measured from a flat white reference panel [47,157]. Consequently, continuous spectral measurements alone cannot eliminate geometric effects. When several sources of interference coexist, intrinsic color must still be distinguished from apparent color changes caused by the imaging conditions.
Full-spectrum hyperspectral imaging is suitable for identifying spectral bands associated with pigments or quality attributes. By contrast, reduced-band systems can decrease both acquisition costs and computational demands. In strawberry bruise detection, a random forest model using the full spectrum achieved an accuracy of 99.21%. A model based on three selected bands produced comparable performance [158]. This finding indicates only that a small number of bands may retain essential information for a specific task. Chili peppers and strawberries differ in their surface curvature, specular reflection, defect mechanisms, spectral characteristics, and acquisition conditions. The direct transfer of such methods to chili pepper sorting is therefore not justified. Moreover, these findings cannot replace independent evaluations under high-speed sorting, motion blur, and continuous-conveyor conditions.
Existing methods target different sources of interference. Spectral–spatial fusion associates spectral bands with localized anomalies, whereas multi-scale features accommodate variations in defect size. Data augmentation introduces predefined perturbations, while attention mechanisms emphasize salient image regions [158,159,160]. However, differences in training perturbations, object categories, and data partitions preclude direct model ranking based solely on accuracy, F1 score, or mean average precision. Artificial augmentation also cannot fully reproduce production-line disturbances, including dust, surface water films, occlusion, continuous rotation, and inter-batch illumination drift. Model comparisons should therefore report the actual ranges of these disturbances, results from independent batches, and performance degradation under domain shifts. Such information is needed to determine whether a model relies on intrinsic color or on confounding cues related to the background, illumination, or object pose.
Based on the available heterogeneous evidence, this review proposes a hierarchical error-control framework that requires further validation. First, calibrated illumination and multiview imaging should be used to reduce input variability. Second, full-spectrum data can support the selection of stable and informative spectral bands. Third, lightweight models should be trained using disturbance ranges representative of actual production lines. Cross-condition evaluations should report performance-retention rates, processing speed, confidence estimates, and procedures for handling low-confidence samples. This framework integrates physical control, information compression, and model robustness. Nevertheless, its benefits must be verified using identical batches and operating conditions.

8.2. Model Generalization Across Cultivars, Origins, and Batches

Machine vision, near-infrared and hyperspectral imaging, and multimodal models can effectively discriminate samples drawn from controlled distributions. However, within-domain performance does not demonstrate generalization across cultivars, geographical origins, or production batches. Studies of fruits and vegetables have identified several mechanisms underlying domain shifts [11,108,161,162]. For chili peppers, research must integrate evidence from cultivar identification [37], multi-sensor fusion for whole bell peppers [39,40], origin-independent models [151], and cross-year spectral validation of paprika [113]. Random data partitions estimate only within-domain performance, whereas tests using independent domains provide stronger evidence of generalizability. Results from these validation levels should not be combined into a single measure of “generalization accuracy.”
RGB features characterize visually observable differences, whereas near-infrared methods establish class boundaries from spectral responses. Previous studies have reported strong classification performance across datasets containing eleven, two, and four apple cultivars [163,164,165]. The cited cultivar-classification accuracies retain their source labels; overall, macro-averaged and one-versus-rest aggregation have not been harmonized. These findings indicate that both visual and spectral information can be discriminative. However, their applicability remains limited to the cultivars, environments, and batches represented in the respective studies. High within-domain accuracy does not ensure stable performance in previously unseen domains.
The available technical approaches also yield inconsistent conclusions regarding the value of increasing model complexity. Fusion of visible and hyperspectral information can characterize complementary types of variation. However, in apple cultivar identification, a particular multi-head attention architecture did not outperform the best configuration incorporating gray-level co-occurrence matrix texture features [162,163]. Furthermore, complementarity between sensing modalities does not necessarily translate into improved cross-domain generalization. If several sensors capture the same batch-specific biases, a fusion model may simply learn those domain-specific features more comprehensively.
As summarized in Table 12, the evaluated methods provide evidence of within-domain class separability. However, they differ in study objects, class numbers, sampling units, data-partitioning strategies, and evaluation metrics. The table therefore compares the boundaries of the available evidence rather than ranking the algorithms. Current comparisons remain insufficient to establish the superiority of any specific approach for chili pepper sorting.
Generalization should be evaluated using a tiered protocol. Within-batch random partitioning can estimate the upper performance bound for a given data distribution. Cross-batch holdout testing can assess sensitivity to changes in harvest time and surface conditions. Cross-origin holdout testing can evaluate differences in growing environments and cultivation practices. Finally, cross-cultivar holdout testing can determine whether a model captures cultivar-independent characteristics.
At each tier, studies should report the sampling and partitioning units, macro-averaged F1 score (equal-weight mean of grade-specific F1 scores), balanced accuracy (macro-recall), cross-domain performance degradation, low-confidence abstention rate (abstained predictions divided by all evaluated items), and confidence calibration. Abstention is distinct from physical rejection. Three deployment strategies should also be compared: no adaptation, lightweight calibration, and full retraining. This protocol distinguishes within-domain classification performance from cross-domain transferability. However, its implementation requires long-term datasets with a tiered structure and full sample traceability.

8.3. Vision–Execution Timing Errors and Compensation in High-Speed Color Sorting

The total latency of high-speed color sorting comprises image acquisition, inference, tracking, communication, valve control, and actuator motion. Predictive tracking can preserve target identities and estimate arrival times [133]. Research on vision-based sorting controllers also indicates that recognition outputs must be synchronized with control logic before actuation [75]. Studies of airflow and material dynamics show that identical triggering times can produce different separation trajectories, loss rates, and impact risks [131]. Timely triggering is therefore necessary but insufficient. System-level evaluation should ultimately focus on successful removal and a reduction in mechanical damage.
Δttotal = Δtcapture + Δtinference + Δttracking + Δtcommunication + Δtvalve + Δtmotion
ΔxvΔttotal + Δxlocalization + Δxslip
Equations (1) and (2) define the total system latency and the corresponding displacement error, respectively.
A symmetry-axis method for bell pepper pose estimation achieved mean angular errors of approximately 6.5°, 7.4°, and 6.9° across three datasets [167]. These findings indicate that directional information can be quantified under the evaluated conditions. However, static pose accuracy does not represent timing accuracy during continuous conveying, target occlusion, or multi-object tracking. Vision-derived errors must therefore be propagated through the motion model to the nozzle or end-effector plane.
Actuator studies provide two complementary types of evidence. Improved Prandtl–Ishlinskii models characterize asymmetric hysteresis in pneumatic artificial muscles. Bench testing of flexible pneumatic ejectors reported a cycle time of 108 ms and an equivalent impact force of 400 g. The corresponding root-mean-square trajectory errors were 22.8 and 20.5 mm along two planar directions [130,132]. The former evidence explains actuator hysteresis, whereas the latter demonstrates the feasibility of rapid, compliant contact. However, these results were obtained under different operating conditions. They cannot be directly ranked or used to establish the suitability of a particular mechanism for continuous chili pepper sorting.
A high-speed Delta parallel sorting system integrates object detection, time-optimal trajectory planning, and mechanical-vibration analysis [168]. It therefore provides a systematic reference for the high-speed handling of soft fruits. When such a system is adapted to chili pepper sorting, a unified clock should record timestamps at every processing stage. Evaluation should extend beyond mean latency to include the median, 95th percentile (p95), jitter, and long-term clock drift.
Pose variation, velocity fluctuations, and actuator hysteresis should be incorporated into a unified predictive compensator, with uncompensated control used as the baseline. Evaluation metrics should include positional error, timing error, final-bin MRsys, FRRsys, damage rate, throughput, and energy consumption per sorted unit. These measurements would help quantify the relative contributions of errors from vision, valve control, and airflow. The main engineering challenges are multi-source timestamp synchronization, online identification of pressure-response models, and stable tracking under dense-target conditions.

8.4. Multi-Objective Co-Optimization of Accuracy, Efficiency, Energy Consumption, and Material Damage

Existing studies often optimize either the overall machine cycle time or the detection of latent defects. Few studies jointly evaluate these objectives using identical equipment, batches, and operating speeds. A low-complexity color–size sorting system requires approximately 11.76–12.92 s per sorting cycle. By contrast, near-infrared early-damage detection achieved a test accuracy of up to 98.60% using wavelet processing, feature selection, and classifier ensembles [57,169]. The former provides evidence of system-level operation, whereas the latter describes detection performance. These results are therefore not directly comparable. More broadly, standardized system boundaries for these metrics remain unavailable.
High-dimensional sensing can expand the range of measurable quality attributes. A four-dimensional hyperspectral framework integrating visible and near-infrared (VNIR), short-wave infrared (SWIR), and three-dimensional shape information has been investigated [170,171]. Its spectral bands, imaging modes, and models differ substantially in performance, cost, calibration requirements, and suitability for online operation. Greater information richness does not necessarily produce greater system-level benefits. A new modality has practical value only when it delivers stable gains in independent tests without compromising cycle-time stability or maintainability.
Material damage is also affected by maturity and defects that exist before sorting. Pre-existing damage must therefore be distinguished from damage introduced by the sorting process. Samples should also be stratified according to maturity, firmness, or softness [169,172]. Otherwise, equipment-induced damage may be overestimated because of pre-existing defects. Conversely, it may be underestimated when assessment is limited to visible surface damage.
A multi-objective system optimization framework requiring further validation is given in Equation (3).
F ( x ) = [ M R s y s ( x ) , F R R s y s ( x ) , Q s y s ( x ) , E ( x ) , D 0 ( x ) , D s ( x ) ]
Here, MRsys and FRRsys are the system missed-rejection and false-rejection rates, respectively, defined from final-bin outcomes in Section 7.4, with rejection-required fruit as the positive class; Qsys is total physical output throughput, irrespective of sorting correctness. Model precision and recall are reported separately as diagnostic metrics. E is system energy intensity, including sensing, computation, conveying and compressed-air consumption. D0 and Ds denote newly induced damage immediately after sorting and after storage, respectively. Decision variables include sensor configuration, model architecture, conveyor settings and actuation parameters.
Evaluations should report model macro-F1 and rejection-class recall separately from final-bin MRsys, FRRsys, overall grading accuracy, physical throughput, end-to-end latency, energy intensity, and newly induced damage rate. Throughput is expressed in items s−1 or kg h−1; energy intensity uses J per processed fruit or kWh t−1 with the same material-accounting boundary. Measurements should use matched batches, illumination and operating speeds.
As an illustrative worked example, a candidate setting x with MRsys = 0.05, FRRsys = 0.03, Qsys = 8 items s−1, E = 45 J per processed fruit, D0 = 0.012 and Ds = 0.025 gives F(x) = [0.05, 0.03, −8, 45, 0.012, 0.025]. Illustrative screening ranges are MRsys = 0.01–0.15, FRRsys = 0.01–0.10, Qsys = 2–12 items s−1, E = 20–100 J per processed fruit, D0 = 0.005–0.050 and Ds = 0.010–0.080. These values are hypothetical system-level examples, not conversions of measured model recall, and must be replaced by experimental ranges for the target sorting line.
A color–size sorting system can serve as the baseline for incremental evaluation. Near-infrared, VNIR/SWIR, and three-dimensional information can then be introduced sequentially. For each modality, independent-test gains, latency, energy consumption, and additional material damage should be recorded. Pareto frontiers can present the resulting trade-offs among competing objectives. This design would reveal the system-level cost of improving detection accuracy. Its implementation requires standardized energy-accounting boundaries and paired inspections of the same materials before and after sorting.

8.5. Lightweight Models, Edge Deployment, and Online Adaptation

Lightweight models, edge deployment, and online adaptation encompass input compression, model inference, and the maintenance of runtime performance. Band reduction, portable detection, optimized object detectors, and multi-sensor fusion have demonstrated feasibility in specific settings. Research on edge-based and real-time sorting also emphasizes hardware, communication, and the complete processing pipeline [66,138,141,173,174,175]. Deployment performance should therefore not be evaluated solely through parameter counts or single-inference speed.
A review of deep learning in soybean production likewise discusses lightweight models, transfer learning, sensor fusion, and edge computing, while identifying data-quality limitations, deployment difficulties, and a lack of standardized evaluation benchmarks [176]. These review-level findings provide cross-crop context for the deployment challenges considered here.
Input compression and network enhancement represent distinct design strategies. Hyperspectral dimensionality reduction decreases the number of input variables. However, relevant studies often omit test hardware, batch size, and end-to-end acquisition overhead [174]. An enhanced YOLOv5s model increased mean average precision from 71.4% to 85.1% [177]. However, its size increased from 13.7 to 23.3 MB, while the frame rate decreased from 31.9 to 30.7 frames per second. Neither model size nor frame rate directly represents peak memory consumption or production-line throughput. Because the evaluated tasks also differ, these values should not be used for direct model ranking. Algorithmic improvements must be reassessed based on the target hardware and across the complete processing pipeline.
Evidence from multi-sensor fusion further illustrates the difference between validation performance and external generalization. Fusion of Vis–NIR, SWIR, hyperspectral, ultrasonic, relaxation, and color measurements increased the cross-validation R2 from 0.82 to 0.90 [39]. The corresponding root-mean-square error of cross-validation decreased from 0.24 to 0.18. However, the independent-prediction R2 increased only from 0.81 to 0.82, while the root-mean-square error of prediction decreased from 0.26 to 0.24. New sensing modalities should therefore be evaluated according to their gains on independent test sets. Latency, energy consumption, equipment cost, and calibration burden should also be considered.
Online adaptation is primarily motivated by environmental variation, instrument heterogeneity, cultivar differences, and field-background complexity. However, its effectiveness remains insufficiently validated for chili pepper color sorting [39,177,178]. This review proposes a “drift detection, controlled updating, and regression testing” framework for future evaluation. Relevant metrics should include detection delay, annotation and updating costs, performance recovery, and performance loss in the original domain. Continuous updating may amplify erroneous labels and transient noise. Without evidence of sustained learning on agricultural production lines, online adaptation should not yet be considered a mature capability.
Edge-computing benchmarks should cover data acquisition, preprocessing, inference, communication, and command generation. They should report latency percentiles, peak memory use, unit-level energy efficiency, temperature increases, and throughput after thermal throttling. These measurements should be conducted on the target central processing unit, graphics processing unit, or neural processing unit. Hardware-aware compression, drift detection, and modality selection should be optimized jointly. The central challenge is to maintain predictable resource use and recognition performance during continuous domain shifts and batch transitions.

8.6. Multispectral Sensing, Data Standards, and Industrial Validation

Multispectral sensing aims to translate hyperspectral band discovery into deployable online systems that can be calibrated, transferred, and independently validated. Reviews and commercial applications provide evidence concerning band reduction, optical geometry, and calibration [179,180,181]. Chili-pepper-related studies have investigated hyperspectral detection of paprika fruit defects [44], machine-vision-based maturity estimation of bell peppers [9], and impurity detection in green peppers [97,156]. Other applications include optimal-band selection for sweet pepper ripening [10] and visible–fluorescence fusion for red peppers [46]. Image-based red pepper sorting [29] and hyperspectral identification of foreign materials in chili peppers have also been investigated [182]. General reviews describe potential engineering pathways, whereas pepper-specific studies define task-level feasibility boundaries. Neither type of evidence can replace the other.
Hyperspectral imaging can identify spectral regions associated with ripeness, pigmentation, and defects. Reviews of horticultural applications indicate that frequently used bands are concentrated within approximately 601–950 nm [179,180,183]. Partial least-squares regression and support vector machines also remain widely used. However, full-spectrum systems are constrained by limited throughput, spectral collinearity, and acquisition speed. They may therefore be more suitable for band discovery and failure analysis than for direct online deployment.
Reduced-band systems have received task-specific validation. Studies of external defects across approximately 550–991 nm combined Savitzky–Golay preprocessing, standard normal variate transformation, the successive projections algorithm, and k-nearest-neighbor classification [184]. Reported class-specific detection accuracies were approximately 93% for healthy potatoes, 93% for black/green-skin potatoes, and 83% for scab/mechanical-damage/broken-skin potatoes [184]. Their denominators and one-versus-rest convention have not been verified, so these values are not presented as an overall accuracy or macro-average. These findings indicate that band reduction can preserve task-relevant discriminatory information. They do not establish that the same bands or preprocessing procedures can be transferred directly to chili peppers.
Commercial point-based visible–near-infrared systems support parallel multichannel operation at conveyor speeds of approximately 1 m s−1. Such systems can process up to approximately 10 fruits s−1 [181]. At this speed, an integration time of 20 ms corresponds to a sample displacement of approximately 20 mm. Their advantages reflect long-term optimization of optical geometry, sampling volume, and calibration maintenance. Nevertheless, their limited spatial localization may cause local defects to be missed when the sampling geometry does not align with the defect location.
As summarized in Table 13, full-spectrum imaging, multispectral imaging, and point-based spectroscopy prioritize information richness, spatial resolution, and industrial maturity to different degrees. These technologies have not been compared under common constraints for batch composition, operating speed, and cost. Table 13 should therefore be interpreted as a qualitative comparison rather than a performance ranking. The proposed benefits of hierarchical sensing architectures still require within-batch ablation experiments and external-domain validation.
Current evidence does not support a universally optimal preprocessing strategy across tasks. Traditional pipelines may benefit from Savitzky–Golay and standard normal variate (SG–SNV) preprocessing or feature selection. By contrast, some deep-learning models may operate directly on raw spectra [181,184,185]. These three strategies should be compared using the same independent test set. Evaluations should report predictive performance, cross-batch stability, processing latency, and the cost of transferring calibration between instruments or domains.
Data standards should make performance differences interpretable and reproducible rather than prescribing a single technical pathway. Datasets should document equipment specifications, optical metadata, domain labels, annotation protocols, raw and corrected data, preprocessing parameters, and data-partitioning procedures. Inter-instrument evaluations should additionally record device identity, calibration version, and operating status. Random partitioning may allow nearly identical samples to appear in both training and test sets. Cross-domain holdout testing is therefore more appropriate for evaluating deployment risks.
Industrial validation should cover the complete closed-loop system, including sensing, illumination, conveying, computation, and actuation. Static classification accuracy cannot substitute for measurements of throughput, p95 decision latency, final-bin MRsys, FRRsys, continuous-operation stability, drift, and resource consumption. Based on the available heterogeneous evidence, this review proposes a candidate development pathway. Full-spectrum systems can first identify critical bands and characterize failure modes. multispectral imaging can then localize color anomalies, while point- or area-based spectral modules can be selected according to the required sampling geometry. However, direct comparisons have not established this pathway as optimal. It should therefore be evaluated according to its performance consistency across devices, batches, and grading speeds.
Commercial documentation describes optical sorting solutions for fresh, frozen, whole, and cut peppers that remove unwanted colors, visible defects, and foreign materials. However, commercial availability alone does not establish independently verified performance across chili cultivars, product forms, and operating conditions. Industrial validation should therefore assess whether these systems consistently meet defined throughput and quality requirements under representative production conditions.
Economic viability depends on whether labor savings, improved saleable yield, and increased product value justify capital and operating costs relative to existing sorting practices. Relevant costs include equipment purchase, production-line integration, electricity, compressed air, maintenance, and calibration. False rejection, handling damage, and unplanned downtime can reduce net returns, while seasonal production and low annual utilization may prolong the payback period. More sophisticated sensing is economically justified only when its incremental economic benefits exceed its additional lifecycle costs. Industrial validation should therefore compare candidate systems with existing sorting practices under matched product specifications and quality requirements. Sustained production trials should quantify saleable output and total cost per tonne of sellable product, alongside payback estimates across realistic utilization scenarios. These assessments would identify the conditions under which technical performance translates into commercial value.

8.7. Research Priorities and a Phased Validation Pathway

Research should follow short-, medium-, and long-term phases covering input standardization, external-domain validation, perception–actuation synchronization, system-cost evaluation, and continuous-operation verification. This pathway reflects evidence dependencies rather than an established industry standard.
Short-term work should standardize data partitioning by fruit, batch, origin, and harvest year. Reporting should cover final-bin MRsys, FRRsys, cross-domain degradation, p95/p99 end-to-end latency, and system power consumption.
Medium-term validation should span origins, harvest years, and devices under no adaptation, lightweight recalibration, and full retraining. It should assess compensation for pose, velocity, and actuator hysteresis.
Long-term studies should assess continuous-shift operation, calibration transfer, fault-induced degradation, and maintenance costs through reproducible pilot-scale protocols.

9. Conclusions and Outlook

Post-harvest sorting directly affects the consistency of chili pepper grades, processing efficiency, and market value. This review adopts the framework of “objects and operating conditions → perception and target information → localization and execution → system evaluation.” It integrates direct evidence from chili pepper studies with relevant engineering knowledge from other crops. Based on this framework, three hierarchical evaluation domains are identified: offline perception validation, component-level performance verification, and online whole-machine sorting assessment. Algorithmic metrics must be linked to target association, temporal sequence matching, and execution validation before they can be translated into practical indicators of equipment performance.

9.1. Evidence-Based Key Conclusions

The reviewed literature provides a limited body of direct evidence for pepper sorting equipment and a broader body of cross-crop methodological and engineering evidence. These sources support different levels of inference. Existing prototypes demonstrate physical sorting, but comprehensive validation linking target identity, actuation, and final destination during sustained operation remains limited.
First, pepper-specific perception and component studies support the feasibility of selected sensing and recognition methods under their reported conditions, but do not independently demonstrate complete-system sorting performance.
Second, direct system-level evidence is concentrated in refs. [3,4,5], covering chili and bell pepper sorting machines or prototypes. These studies support the feasibility of linking visual decisions to physical separation under specific operating conditions, but do not establish performance across pepper types, batches, throughput levels, or extended continuous operation.
Third, studies on apple, tomato, potato, sweet potato, onion, and other crops provide engineering approaches and evaluation examples. Their applicability to pepper sorting remains to be tested under pepper-specific material and operating conditions.
Technical configurations should be selected according to material value, grading objectives, allowable damage thresholds, required throughput, and cost constraints. High-speed appearance grading may favor lightweight vision systems operated under controlled illumination. Narrow-band, spectral, or fluorescence sensing should be introduced only when RGB imaging provides insufficient information. Any resulting performance improvement should be verified through ablation experiments conducted under identical conditions.

9.2. Boundaries of the Existing Evidence

A major limitation of the current literature is the scarcity of independently developed, full-scale chili pepper sorting platforms. Even within related series of studies, algorithm-focused and equipment-focused papers often differ substantially in their object definitions, grading criteria, sample units, data-partitioning strategies, and operating conditions.
Most studies report only offline accuracy, mean average precision (mAP), frames per second (FPS), or sensor-response metrics. They rarely provide unified records linking input samples, target identities, classification results, execution commands, and final sorting destinations.
Lighting fluctuations, occlusion, vibration, and conveyor disturbances may also reduce the stability of multi-source sensing. Few studies simultaneously record localization errors, cumulative latency, and execution deviations along the trajectory of each material item. Technical and economic feasibility, maintenance costs, and resilience to component degradation or failure are also rarely evaluated.
Consequently, the proposed evidence hierarchy, interface specifications, and composite metrics should be regarded as testable evaluation frameworks rather than formal industry standards. Industrial maturity, optimal sensing strategies, and the benefits of large-scale deployment require further validation using whole-system operational data.

9.3. Priorities for Industrial Validation

Future research should first establish standardized testing protocols and cross-batch data-isolation rules. Datasets should be partitioned by individual fruit, batch, origin, and cultivar. Grading criteria, equipment specifications, and optical metadata should also be disclosed. These measures would enable meaningful comparisons among RGB, narrow-band, multispectral, and sensor-fusion solutions under consistent operating conditions.
Data acquisition, inference, localization, communication, valve control, and sorting outcomes should then be synchronized using a unified clock. Studies should report comprehensive metrics across different processing volumes. These metrics should include separate model and final-bin confusion matrices, final-bin MRsys and FRRsys with rejection-required fruit as positive, end-to-end latency with p95 and p99 percentiles, energy consumption per unit mass, and sorting-induced product damage.
For ripeness or internal-quality assessment, visual, spectral, and pigment-related features should be calibrated against traceable physicochemical indices. Classification results, spatial coordinates, arrival times, execution actions, and final destinations should be linked along continuous material trajectories. Modular interfaces should also be adopted for perception, control, and execution. Such interfaces would facilitate component replacement, fault isolation, and maintenance-cost evaluation.
A chili pepper color-sorting system can be considered repeatable, maintainable, and scalable only when these performance metrics remain stable across batches, equipment platforms, and extended periods of continuous operation.

Author Contributions

Conceptualization, J.C.; Investigation, J.C.; Writing—original draft, J.C.; Visualization, J.C.; Writing—review and editing, Z.T., Y.W., L.Z. and Y.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Jiangsu University College Student Innovation Training Program (project number: X202610299801) and the 25th batch of college student scientific research project funding project of Jiangsu University (project number: 25B058) and Modern Agricultural Machinery Equipment and Technology Demonstration and Promotion Project in Jiangsu Province (Specialized Agricultural Machinery Pilot and Maturation Project).

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The authors express their sincere gratitude for the valuable technical support and resources that contributed to this research.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Amna; Akram, M.W.; Li, G.Q.; Akram, M.Z.; Faheem, M.; Omar, M.M.; Hassan, M.G. Machine vision-based automatic fruit quality detection and grading. Front. Agric. Sci. Eng. 2025, 12, 274–287. [Google Scholar] [CrossRef] [Scilit]
  2. Tai, S.; Tang, Z.; Li, B.; Wang, S.; Guo, X. Intelligent Recognition and Automated Production of Chili Peppers: A Review Addressing Varietal Diversity and Technological Requirements. Agriculture 2025, 15, 1200. [Google Scholar] [CrossRef] [Scilit]
  3. Mohi-Alden, K.; Omid, M.; Soltani Firouz, M.; Nasiri, A. A machine vision-intelligent modelling based technique for in-line bell pepper sorting. Inf. Process. Agric. 2023, 10, 491–503. [Google Scholar] [CrossRef] [Scilit]
  4. Mohi-Alden, K.; Omid, M.; Soltani Firouz, M.; Nasiri, A. Design and evaluation of an intelligent sorting system for bell pepper using deep convolutional neural networks. J. Food Sci. 2022, 87, 289–301. [Google Scholar] [CrossRef] [Scilit]
  5. Lestari, H.A.; Kurniawan, A.; Wahab, L. Automated Conveyor System of Sorting and Grading for Red Chili Pepper (Capsicum annum L.) using Image Processing and Artificial Neural Network. J. Tek. Pertan. Lampung (J. Agric. Eng.) 2024, 13, 1320–1333. [Google Scholar] [CrossRef] [Scilit]
  6. Seiphepi, G.; Zungeru, A.M.; Gaboitaolelwe, J.; Lebekwe, C.; Mtengi, B. Automatic Bell Pepper Colour Detector and Sorting Machine. Int. J. Eng. Res. Technol. 2020, 13, 3156–3166. [Google Scholar] [CrossRef] [Scilit]
  7. Zhu, Y.; Zhang, S.; Tang, S.; Gao, Q. Research Progress and Applications of Artificial Intelligence in Agricultural Equipment. Agriculture 2025, 15, 1703. [Google Scholar] [CrossRef] [Scilit]
  8. Jiang, C.; Miao, K.; Hu, Z.; Gu, F.; Yi, K. Image Recognition Technology in Smart Agriculture: A Review of Current Applications Challenges and Future Prospects. Processes 2025, 13, 1402. [Google Scholar] [CrossRef] [Scilit]
  9. Villaseñor-Aguilar, M.J.; Bravo-Sánchez, M.G.; Padilla-Medina, J.A.; Vázquez-Vera, J.L.; Guevara-González, R.G.; García-Rodríguez, F.J.; Barranco-Gutiérrez, A.I. A Maturity Estimation of Bell Pepper (Capsicum annuum L.) by Artificial Vision System for Quality Control. Appl. Sci. 2020, 10, 5097. [Google Scholar] [CrossRef] [Scilit]
  10. Muñoz-Postigo, J.; Valero, E.M.; Martínez-Domingo, M.A.; Lara, F.J.; Nieves, J.L.; Romero, J.; Hernández-Andrés, J. Band selection pipeline for maturity stage classification in bell peppers: From full spectrum to simulated camera data. J. Food Eng. 2024, 365, 111824. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, Y.; Ouyang, C.; Peng, H.; Deng, J.; Yang, L.; Chen, H.; Luo, Y.; Jiang, P. YOLO-ALW: An Enhanced High-Precision Model for Chili Maturity Detection. Sensors 2025, 25, 1405. [Google Scholar] [CrossRef] [Scilit]
  12. Wang, C.; Liu, S.; Wang, Y.; Xiong, J.; Zhang, Z.; Zhao, B.; Luo, L.; Lin, G.; He, P. Application of Convolutional Neural Network-Based Detection Methods in Fresh Fruit Production: A Comprehensive Review. Front. Plant Sci. 2022, 13, 868745. [Google Scholar] [CrossRef] [Scilit]
  13. Blasco, J.; Munera, S.; Aleixos, N.; Cubero, S.; Molto, E. Machine Vision-Based Measurement Systems for Fruit and Vegetable Quality Control in post-harvest. In Advances in Biochemical Engineering/Biotechnology; Springer: Berlin/Heidelberg, Germany, 2017; pp. 71–91. [Google Scholar] [CrossRef] [Scilit]
  14. Patel, K.K.; Kar, A.; Jha, S.N.; Khan, M.A. Machine vision system: A tool for quality inspection of food and agricultural products. J. Food Sci. Technol. 2012, 49, 123–141. [Google Scholar] [CrossRef] [Scilit]
  15. Khan, Z.; Shen, Y.; Liu, H. ObjectDetection in Agriculture: A Comprehensive Review of Methods, Applications, Challenges, and Future Directions. Agriculture 2025, 15, 1351. [Google Scholar] [CrossRef] [Scilit]
  16. Rasekh, M.; Karami, H.; Fuentes, S.; Kaveh, M.; Rusinek, R.; Gancarz, M. Preliminary study nondestructive sorting techniques for pepper (Capsicum annuum L.) using odor parameter. LWT 2022, 164, 113667. [Google Scholar] [CrossRef] [Scilit]
  17. Lu, J.; Zhang, M.; Hu, Y.; Ma, W.; Tian, Z.; Liao, H.; Chen, J.; Yang, Y. From Outside to Inside: The Subtle Probing of Globular Fruits and Solanaceous Vegetables Using Machine Vision and Near-Infrared Methods. Agronomy 2024, 14, 2395. [Google Scholar] [CrossRef] [Scilit]
  18. Wang, C.; Li, X.; Zhang, Z.; Luo, X.; Cai, J.; Wang, A. Nondestructive Quality Detection of Characteristic Fruits Based on Vis/NIR Spectroscopy: Principles, Systems, and Applications. Agriculture 2025, 15, 2167. [Google Scholar] [CrossRef] [Scilit]
  19. Saldaña, E.; Siche, R.; Luján, M.; Quevedo, R. Review: Computer vision applied to the inspection and quality control of fruits and vegetables. Braz. J. Food Technol. 2013, 16, 254–272. [Google Scholar] [CrossRef] [Scilit]
  20. Huang, Y.; Li, Z.; Bian, Z.; Jin, H.; Zheng, G.; Hu, D.; Sun, Y.; Fan, C.; Xie, W.; Fang, H. Overview of Deep Learning and Nondestructive Detection Technology for Quality Assessment of Tomatoes. Foods 2025, 14, 286. [Google Scholar] [CrossRef] [Scilit]
  21. Yu, K.; Zhong, M.; Zhu, W.; Rashid, A.; Han, R.; Virk, M.S.; Duan, K.; Zhao, Y.; Ren, X. Advances in Computer Vision and Spectroscopy Techniques for nondestructive Quality Assessment of Citrus Fruits: A Comprehensive Review. Foods 2025, 14, 386. [Google Scholar] [CrossRef] [Scilit]
  22. Tiamiyu, Q.O.; Adebayo, S.E.; Ibrahim, N. Recent advances on post-harvest technologies of bell pepper: A review. Heliyon 2023, 9, e15302. [Google Scholar] [CrossRef] [Scilit]
  23. Wang, D.; Zhang, M.; Mujumdar, A.S.; Yu, D. Advanced Detection Techniques Using Artificial Intelligence in Processing of Berries. Food Eng. Rev. 2022, 14, 176–199. [Google Scholar] [CrossRef] [Scilit]
  24. Jun, Q.; Sasao, A.; Shibusawa, S.; Kondo, N. Extracting External Features of Sweet Peppers Using Machine Vision System on Mobile Fruits Grading Robot. Int. J. Food Eng. 2012, 8, 1. [Google Scholar] [CrossRef] [Scilit]
  25. Zhu, L.; Spachos, P.; Pensini, E.; Plataniotis, K.N. Deep learning and machine vision for food processing: A survey. Curr. Res. Food Sci. 2021, 4, 233–249. [Google Scholar] [CrossRef] [Scilit]
  26. Mohammadi Baneh, N.; Navid, H.; Kafashan, J. Mechatronic components in apple sorting machines with computer vision. J. Food Meas. Charact. 2018, 12, 1135–1155. [Google Scholar] [CrossRef] [Scilit]
  27. Han, B.; Zhang, J.; Almodfer, R.; Wang, Y.; Sun, W.; Bai, T.; Dong, L.; Hou, W. Research on Innovative Apple Grading Technology Driven by Intelligent Vision and Machine Learning. Foods 2025, 14, 258. [Google Scholar] [CrossRef] [Scilit]
  28. Huang, X.; Zhang, D.; He, Y. An intelligent surface quality detection method for defect identification and classification based on YOLOv5 model. R. Soc. Open Sci. 2025, 12, 1–20. [Google Scholar] [CrossRef] [Scilit]
  29. Khuriyati, N.; Pamungkas, A.P.; Pambudi, A.A. The Sorting and Grading of Red Chilli Peppers (Capsicum annuum L.) Using Digital Image Processing. Int. J. Agric. Environ. Sci. 2019, 6, 17–23. [Google Scholar] [CrossRef] [Scilit]
  30. Hou, L.; Liu, Z.; You, J.; Liu, Y.; Xiang, J.; Zhou, J.; Pan, Y. Tomato Sorting System Based on Machine Vision. Electronics 2024, 13, 2114. [Google Scholar] [CrossRef] [Scilit]
  31. Liang, Z.; Zhu, T.; Teng, G.; Zhang, Y.; Gu, Z. YOLO-RGDD: A Novel Method for the Online Detection of Tomato Surface Defects. Foods 2025, 14, 2513. [Google Scholar] [CrossRef] [Scilit]
  32. Berry, H.M.; Rickett, D.V.; Baxter, C.J.; Enfissi, E.M.A.; Fraser, P.D. Carotenoid biosynthesis and sequestration in red chilli pepper fruit and its impact on colour intensity traits. J. Exp. Bot. 2019, 70, 2637–2650. [Google Scholar] [CrossRef] [Scilit]
  33. Li, Z.; Zhao, H.; Jing, Z.; Zhao, Z.; Wang, M.; Gong, M.; Wu, X.; He, Z.; Liao, J.; Liu, M.; et al. Recent Advances in Pepper Fruit Glossiness. Genes 2025, 16, 1319. [Google Scholar] [CrossRef] [Scilit]
  34. Sihombing, Y.F.; Septiarini, A.; Kridalaksana, A.H.; Puspitasari, N. Chili Classification Using Shape and Color Features Based on Image Processing. Sci. J. Inform. 2022, 9, 42–50. [Google Scholar] [CrossRef] [Scilit]
  35. Azzahra, T.; Rahmadan, R.; Abi Maulana, F.; Asmita, I.; Rahayu, E.; Erwis, F. Classification of Capsicum Varieties Using Color Analysis with Convolutional Neural Network. J. ICT Apl. Syst. 2024, 3, 66–74. [Google Scholar] [CrossRef] [Scilit]
  36. Arinal, V.; Sarimole, F.M.; Sugeng, S.; Julianda, R. Chili Pepper Variety Detection System Using the Principal Component Analysis Method. Int. J. Mech. Electr. Civ. Eng. 2024, 1, 72–87. [Google Scholar] [CrossRef] [Scilit]
  37. Barbosa, M.D.O.; Aguiar, F.P.L.; Sousa, S.D.S.; Cordeiro, L.D.S.; Nääs, I.D.A.; Okano, M.T. YOLOv8m for Automated Pepper Variety Identification: Improving Accuracy with Data Augmentation. Appl. Sci. 2025, 15, 7024. [Google Scholar] [CrossRef] [Scilit]
  38. Kasampalis, D.S.; Tsouvaltzis, P.; Ntouros, K.; Gertsis, A.; Gitas, I.; Moshou, D.; Siomos, A.S. Nutritional composition changes in bell pepper as affected by the ripening stage of fruits at harvest or post-harvest storage and assessed non-destructively. J. Sci. Food Agric. 2022, 102, 445–454. [Google Scholar] [CrossRef] [Scilit]
  39. Ignat, T.; Alchanatis, V.; Schmilovitch, Z.E. Maturity prediction of intact bell peppers by sensor fusion. Comput. Electron. Agric. 2014, 104, 9–17. [Google Scholar] [CrossRef] [Scilit]
  40. Kasampalis, D.S.; Tsouvaltzis, P.; Ntouros, K.; Gertsis, A.; Gitas, I.; Siomos, A.S. The use of digital imaging, chlorophyll fluorescence and Vis/NIR spectroscopy in assessing the ripening stage and freshness status of bell pepper fruit. Comput. Electron. Agric. 2021, 187, 106265. [Google Scholar] [CrossRef] [Scilit]
  41. Kim, S.; Youl Ha, T.; Park, J. Characteristics of pigment composition and colour value by the difference of harvesting times in Korean red pepper varieties (Capsicum annuum L.). Int. J. Food Sci. Technol. 2008, 43, 915–920. [Google Scholar] [CrossRef] [Scilit]
  42. Santos, M.B.D.C.; Melo, R.D.S.; Sousa, V.E.D.C.; Medeiros, A.M.; Pessoa, A.M.D.S.; Silva, S.D.C.; do Nascimento, A.M.M.; Barroso, P.A. A Computer Vision-Based Methodology to Estimate Fruit Colour Diversity in Ornamental Pepper (Capsicum spp.). Plant Breed. 2024, 143, pbr.13237. [Google Scholar] [CrossRef] [Scilit]
  43. Fatchurrahman, D.; Castillejo, N.; Hilaili, M.; Russo, L.; Fathi-Najafabadi, A.; Rahman, A. A Novel Damage Inspection Method Using Fluorescence Imaging Combined with Machine Learning Algorithms Applied to Green Bell Pepper. Horticulturae 2024, 10, 1336. [Google Scholar] [CrossRef] [Scilit]
  44. Faqeerzada, M.A.; Kim, Y.N.; Kim, H.; Akter, T.; Kim, H.; Park, M.S.; Kim, M.S.; Baek, I.; Cho, B.K. Hyperspectral imaging system for pre- and post-harvest defect detection in paprika fruit. Postharvest Biol. Technol. 2024, 218, 113151. [Google Scholar] [CrossRef] [Scilit]
  45. Huang, Z.; Takemoto, T.; Omwange, K.A.; Saito, Y.; Kuramoto, M.; Kondo, N. Macroscopic and microscopic characterization of fluorescence properties of multiple sweet pepper cultivars (Capsicum annuum L.) using excitation-emission matrix and UV induced fluorescence imaging. Spectrochim. Acta Part A Mol. Biomol. Spectrosc. 2023, 288, 122094. [Google Scholar] [CrossRef] [Scilit]
  46. Ma, S.; Li, Y.; Peng, Y.; Nie, S.; Wang, W.; Zhang, Y. Fusion of visible and fluorescence imaging through deep neural network for color value prediction of pelletized red peppers. J. Food Sci. 2024, 89, 7410–7421. [Google Scholar] [CrossRef] [Scilit]
  47. Schmilovitch, Z.E.; Ignat, T.; Alchanatis, V.; Gatker, J.; Ostrovsky, V.; Felföldi, J. Hyperspectral imaging of intact bell peppers. Biosyst. Eng. 2014, 117, 83–93. [Google Scholar] [CrossRef] [Scilit]
  48. Al-Mallahi, A.; Kataoka, T.; Okamoto, H.; Shibata, Y. An image processing algorithm for detecting in-line potato tubers without singulation. Comput. Electron. Agric. 2010, 70, 239–244. [Google Scholar] [CrossRef] [Scilit]
  49. Hemming, J.; Ruizendaal, J.; Hofstee, J.; Van Henten, E. Fruit Detectability Analysis for Different Camera Positions in Sweet-Pepper. Sensors 2014, 14, 6032–6044. [Google Scholar] [CrossRef] [Scilit]
  50. Gupta, H.; Lilienthal, A.J.; Andreasson, H.; Kurtser, P. NDT-6D for color registration in agri-robotic applications. J. Field Robot. 2023, 40, 1603–1619. [Google Scholar] [CrossRef] [Scilit]
  51. Li, A.; Wang, C.; Wang, A.; Sun, J.; Gu, F.; Zhang, T. YOLO-MSRF: A Multimodal Segmentation and Refinement Framework for Tomato Fruit Detection and Segmentation with Count and Size Estimation Under Complex Illumination. Agriculture 2026, 16, 277. [Google Scholar] [CrossRef] [Scilit]
  52. Tasneem, Z.; Oka, K.; Tada, N. Sweet pepper detection in day and night greenhouse environments using thermal and depth imaging. ROBOMECH J. 2025, 12, 29. [Google Scholar] [CrossRef] [Scilit]
  53. Zhao, H.; Deng, W.; Xie, S.; Zhao, Z. Performance Optimization and Experimental Study of Small-Scale Potato-Grading Device. Agriculture 2024, 14, 822. [Google Scholar] [CrossRef] [Scilit]
  54. Feng, Z.; Zhu, L.; Zhou, R.; Ntakirutimana, T.; Wei, B.; Wang, D. Transforming Apple Grading: Standards Survey and Deep Learning Insights. Food Bioprocess Technol. 2025, 18, 5954–5969. [Google Scholar] [CrossRef] [Scilit]
  55. Xiang, P.; Pan, F.; Duan, X.; Yang, D.; Hu, M.; He, D.; Zhao, X.; Huang, F. A Method for Sorting High-Quality Fresh Sichuan Pepper Based on a Multi-Domain Multi-Scale Feature Fusion Algorithm. Foods 2024, 13, 2776. [Google Scholar] [CrossRef] [Scilit]
  56. Masum, A.A.; Himel, M.M.H.; Salehin, M.M.; Rahman, K.S.; Ahamed, S.; Kabir, M.; Bhuiyan, M.G.K.; Rahman, A. Development of automated real-time mango grader using machine vision technique. Discov. Agric. 2025, 3, 104. [Google Scholar] [CrossRef] [Scilit]
  57. Dewi, T.; Risma, P.; Oktarina, Y. Fruit sorting robot based on color and size for an agricultural product packaging system. Bull. Electr. Eng. Inform. 2020, 9, 1438–1445. [Google Scholar] [CrossRef] [Scilit]
  58. Al Ohali, Y. Computer vision based date fruit grading system: Design and implementation. J. King Saud Univ. —Comput. Inf. Sci. 2011, 23, 29–36. [Google Scholar] [CrossRef] [Scilit]
  59. Dhakshina Kumar, S.; Esakkirajan, S.; Bama, S.; Keerthiveena, B. A microcontroller based machine vision approach for tomato grading and sorting using SVM classifier. Microprocess. Microsyst. 2020, 76, 103090. [Google Scholar] [CrossRef] [Scilit]
  60. Cubero, S.; Aleixos, N.; Moltó, E.; Gómez-Sanchis, J.; Blasco, J. Advances in Machine Vision Applications for Automatic Inspection and Quality Evaluation of Fruits and Vegetables. Food Bioprocess Technol. 2011, 4, 487–504, Erratum in Food Bioprocess Technol. 2011, 4, 829–830. https://doi.org/10.1007/s11947-011-0585-8. [Google Scholar] [CrossRef] [Scilit]
  61. Bratu, A.M.; Popa, C.; Bojan, M.; Logofatu, P.C.; Petrus, M. Nondestructive methods for fruit quality evaluation. Sci. Rep. 2021, 11, 7782. [Google Scholar] [CrossRef] [Scilit]
  62. Toylan, H.; Kuscu, H. A Real-Time Apple Grading System Using Multicolor Space. Sci. World J. 2014, 2014, 292681. [Google Scholar] [CrossRef] [Scilit]
  63. Elwakeel, A.E.; Mazrou, Y.S.A.; Tantawy, A.A.; Okasha, A.M.; Elmetwalli, A.H.; Elsayed, S.; Makhlouf, A.H. Designing, Optimizing, and Validating a Low-Cost, Multi-Purpose, Automatic System-Based RGB Color Sensor for Sorting Fruits. Agriculture 2023, 13, 1824. [Google Scholar] [CrossRef] [Scilit]
  64. Hoshino, H.; Shindo, T.; Hiraguri, T.; Itoh, N. RGB Color Space-Enhanced Training Data Generation for Cucumber Classification. J. Imaging 2025, 11, 120. [Google Scholar] [CrossRef] [Scilit]
  65. Wang, J.; Xia, D.; Wan, J.; Hou, X.; Shen, G.; Li, S.; Chen, H.; Cui, Q.; Zhou, M.; Wang, J.; et al. Color grading of green Sichuan pepper (Zanthoxylum armatum DC.) dried fruit based on image processing and BP neural network algorithm. Sci. Hortic. 2024, 331, 113171. [Google Scholar] [CrossRef] [Scilit]
  66. Chakraborty, S.K.; Subeesh, A.; Dubey, K.; Jat, D.; Chandel, N.S.; Potdar, R.; Rao, N.R.N.V.G.; Kumar, D. Development of an optimally designed real-time automatic citrus fruit grading–sorting machine leveraging computer vision-based adaptive deep learning model. Eng. Appl. Artif. Intell. 2023, 120, 105826. [Google Scholar] [CrossRef] [Scilit]
  67. Abubeker, K.M.; Abhijit; Akhil, S.; Akshat Kumar, V.K.; Jose, B.K. Computer Vision Assisted Real-Time Bird Eye Chili Classification Using YOLO V5 Framework. J. Artif. Intell. Technol. 2024, 4, 265–271. [Google Scholar] [CrossRef] [Scilit]
  68. Xu, B.; Cui, X.; Ji, W.; Yuan, H.; Wang, J. Apple Grading Method Design and Implementation for Automatic Grader Based on Improved YOLOv5. Agriculture 2023, 13, 124. [Google Scholar] [CrossRef] [Scilit]
  69. Moya, V.; Guerra, M.; Pazmiño, K.; Abedrabbo, F.; Chicaiza, F.A.; Pozo-Espín, D. Tomato classification with YOLOv8: Enhancing automated sorting and quality assessment. Smart Agric. Technol. 2025, 12, 101221. [Google Scholar] [CrossRef] [Scilit]
  70. Cong, P.; Wang, K.; Liang, J.; Xu, Y.; Li, T.; Xue, B. TQVGModel: Tomato Quality Visual Grading and Instance Segmentation Deep Learning Model for Complex Scenarios. Agronomy 2025, 15, 1273. [Google Scholar] [CrossRef] [Scilit]
  71. Rybacki, P.; Bahcevandziev, K.; Jarquin, D.; Kowalik, I.; Osuch, A.; Osuch, E.; Niemann, J. Three-Dimensional Convolutional Neural Networks (3D-CNN) in the Classification of Varieties and Quality Assessment of Soybean Seeds (Glycine max L. Merrill). Agronomy 2025, 15, 2074. [Google Scholar] [CrossRef] [Scilit]
  72. Hu, G.; Zhang, E.; Zhou, J.; Zhao, J.; Gao, Z.; Sugirbay, A.; Jin, H.; Zhang, S.; Chen, J. Infield Apple Detection and Grading Based on Multi-Feature Fusion. Horticulturae 2021, 7, 276. [Google Scholar] [CrossRef] [Scilit]
  73. Xu, J.; Lu, Y. Design and Preliminary Evaluation of Automated Sweetpotato Sorting Mechanisms. AgriEngineering 2024, 6, 3058–3069. [Google Scholar] [CrossRef] [Scilit]
  74. Huynh, H.X.; Lam, B.H.; Le, H.V.C.; Le, T.T.; Duong-Trung, N. Design of an IoT ultrasonic-vision based system for automatic fruit sorting utilizing size and color. Internet Things 2024, 25, 101017. [Google Scholar] [CrossRef] [Scilit]
  75. Haggag, M.; Abdelhay, S.; Mecheter, A.; Gowid, S.; Musharavati, F.; Ghani, S. An Intelligent Hybrid Experimental-Based Deep Learning Algorithm for Tomato-Sorting Controllers. IEEE Access 2019, 7, 106890–106898. [Google Scholar] [CrossRef] [Scilit]
  76. You, J.; Wang, B.; Qin, C.; Wang, D.; Jin, N.; Zhao, X.; Li, F.; Dou, G.; Bai, H. Development of tomato nondestructive measurement system based on machine vision. J. Food Meas. Charact. 2025, 19, 3507–3525. [Google Scholar] [CrossRef] [Scilit]
  77. Wang, W.; Li, C. A multimodal machine vision system for quality inspection of onions. J. Food Eng. 2015, 166, 291–301. [Google Scholar] [CrossRef] [Scilit]
  78. Bharadwaj, S.; Arora, V.; Mulaveesala, R. Subsurface Bruise Detection in Fruits Using Active Infrared Imaging. Appl. Fruit Sci. 2025, 67, 472. [Google Scholar] [CrossRef] [Scilit]
  79. Huang, Y.; Xiong, J.; Li, Z.; Hu, D.; Sun, Y.; Jin, H.; Zhang, H.; Fang, H. Recent Advances in Light Penetration Depth for post-harvest Quality Evaluation of Fruits and Vegetables. Foods 2024, 13, 2688. [Google Scholar] [CrossRef] [Scilit]
  80. Romano, G.; Argyropoulos, D.; Nagle, M.; Khan, M.T.; Müller, J. Combination of digital images and laser light to predict moisture content and color of bell pepper simultaneously during drying. J. Food Eng. 2012, 109, 438–448. [Google Scholar] [CrossRef] [Scilit]
  81. Unay, D.; Gosselin, B.; Kleynen, O.; Leemans, V.; Destain, M.F.; Debeir, O. Automatic grading of Bi-colored apples by multispectral machine vision. Comput. Electron. Agric. 2011, 75, 204–212. [Google Scholar] [CrossRef] [Scilit]
  82. Gangadhar, B.H.; Mishra, R.K.; Pandian, G.; Park, S.W. Comparative Study of Color, Pungency, and Biochemical Composition in Chili Pepper (Capsicum annuum) Under Different Light-emitting Diode Treatments. HortScience 2012, 47, 1729–1735. [Google Scholar] [CrossRef] [Scilit]
  83. Peng, Y.; Sun, J.; Wu, Z.; Shi, L.; Ji, X.; Jia, Y.; Xie, Y. Apple maturity quantification based on fine-grained coloration analysis using deep learning and computer vision. J. Food Meas. Charact. 2025, 19, 9637–9653. [Google Scholar] [CrossRef] [Scilit]
  84. Campos, M.T.; Maia, L.F.; Edwards, H.G.M.; Cappa de Oliveira, L. Raman Analysis of Natural Pigments During the Ripening Process in Different Types of Peppers. J. Braz. Chem. Soc. 2025, 36, e-20250101. [Google Scholar] [CrossRef] [Scilit]
  85. Viveros Escamilla, L.D.; Gómez-Espinosa, A.; Escobedo Cabello, J.A.; Cantoral-Ceballos, J.A. Maturity Recognition and Fruit Counting for Sweet Peppers in Greenhouses Using Deep Learning Neural Networks. Agriculture 2024, 14, 331. [Google Scholar] [CrossRef] [Scilit]
  86. Appe, S.N.; Arulselvi, G.; Balaji, G.N. CAM-YOLO: Tomato detection and classification based on improved YOLOv5 using combining attention mechanism. PeerJ Comput. Sci. 2023, 9, e1463. [Google Scholar] [CrossRef] [Scilit]
  87. Legner, R.; Voigt, M.; Servatius, C.; Klein, J.; Hambitzer, A.; Jaeger, M. A Four-Level Maturity Index for Hot Peppers (Capsicum annum) Using Non-Invasive Automated Mobile Raman Spectroscopy for On-Site Testing. Appl. Sci. 2021, 11, 1614. [Google Scholar] [CrossRef] [Scilit]
  88. Ye, R.; Shao, G.; Gao, Q.; Zhang, H.; Li, T. CR-YOLOv9: Improved YOLOv9 Multi-Stage Strawberry Fruit Maturity Detection Application Integrated with CRNET. Foods 2024, 13, 2571. [Google Scholar] [CrossRef] [Scilit]
  89. Zhang, Z.; Wang, Y.; Chai, S.; Tian, Y. Detection and Maturity Classification of Dense Small Lychees Using an Improved Kolmogorov–Arnold Network–Transformer. Plants 2025, 14, 3378. [Google Scholar] [CrossRef] [Scilit]
  90. Zhao, S.; Fang, C.; Hua, T.; Jiang, Y. Detecting the Maturity of Red Strawberries Using Improved YOLOv8s Model. Agriculture 2025, 15, 2263. [Google Scholar] [CrossRef] [Scilit]
  91. Nithya, R.; Santhi, B.; Manikandan, R.; Rahimi, M.; Gandomi, A.H. Computer Vision System for Mango Fruit Defect Detection Using Deep Convolutional Neural Network. Foods 2022, 11, 3483. [Google Scholar] [CrossRef] [Scilit]
  92. Chen, J.; Fu, H.; Lin, C.; Liu, X.; Wang, L.; Lin, Y. YOLOPears: A novel benchmark of YOLO object detectors for multi-class pear surface defect detection in quality grading systems. Front. Plant Sci. 2025, 16, 1483824. [Google Scholar] [CrossRef] [Scilit]
  93. Xue, S.; Li, Z.; Wang, D.; Zhu, T.; Zhang, B.; Ni, C. YOLO-ALDS: An instance segmentation framework for tomato defect segmentation and grading based on active learning and improved YOLO11. Comput. Electron. Agric. 2025, 238, 110820. [Google Scholar] [CrossRef] [Scilit]
  94. Agarla, M.; Napoletano, P.; Schettini, R. Quasi Real-Time Apple Defect Segmentation Using Deep Learning. Sensors 2023, 23, 7893. [Google Scholar] [CrossRef] [Scilit]
  95. Gao, X.; Li, S.; Su, X.; Li, Y.; Huang, L.; Tang, W.; Zhang, Y.; Dong, M. Application of Advanced Deep Learning Models for Efficient Apple Defect Detection and Quality Grading in Agricultural Production. Agriculture 2024, 14, 1098. [Google Scholar] [CrossRef] [Scilit]
  96. Qiu, D.; Guo, T.; Yu, S.; Liu, W.; Li, L.; Sun, Z.; Peng, H.; Hu, D. Classification of Apple Color and Deformity Using Machine Vision Combined with CNN. Agriculture 2024, 14, 978. [Google Scholar] [CrossRef] [Scilit]
  97. Zhang, J.; Ma, L.; Gou, Y.; Xia, W.; Chang, X.; Liu, H.; An, T. Detection of green pepper impurities based on hyperspectral imaging technology. Spectrochim. Acta Part A Mol. Biomol. Spectrosc. 2025, 338, 126170. [Google Scholar] [CrossRef] [Scilit]
  98. Luo, J.; Yang, Z.; Cao, Y.; Wen, T.; Li, D. RT-DETR-MCDAF: Multimodal Fusion of Visible Light and Near-Infrared Images for Citrus Surface Defect Detection in the Compound Domain. Agriculture 2025, 15, 630. [Google Scholar] [CrossRef] [Scilit]
  99. Cordeiro, L.D.S.; Nääs, I.D.A.; Okano, M.T. Smart post-harvest Management of Strawberries: YOLOv8-Driven Detection of Defects, Diseases, and Maturity. AgriEngineering 2025, 7, 246. [Google Scholar] [CrossRef] [Scilit]
  100. Zhu, H.; Wang, D.; Wei, Y.; Wang, P.; Su, M. YOLOV8-CMS: A high-accuracy deep learning model for automated citrus leaf disease classification and grading. Plant Methods 2025, 21, 88. [Google Scholar] [CrossRef] [Scilit]
  101. Li, H.; Wang, X.; Bu, Y.; David, C.C.; Chen, X. YOLOv8-Orah: An Improved Model for post-harvest Orah Mandarin (Citrus reticulata cv. Orah) Surface Defect Detection. Agronomy 2025, 15, 891. [Google Scholar] [CrossRef] [Scilit]
  102. Singh, R.; Nisha, R.; Naik, R.; Upendar, K.; Nickhil, C.; Deka, S.C. Sensor fusion techniques in deep learning for multimodal fruit and vegetable quality assessment: A comprehensive review. J. Food Meas. Charact. 2024, 18, 8088–8109. [Google Scholar] [CrossRef] [Scilit]
  103. Yang, C.; Guo, Z.; Fernandes Barbin, D.; Dai, Z.; Watson, N.; Povey, M.; Zou, X. Hyperspectral Imaging and Deep Learning for Quality and Safety Inspection of Fruits and Vegetables: A Review. J. Agric. Food Chem. 2025, 73, 10019–10035. [Google Scholar] [CrossRef] [Scilit]
  104. Pham, Q.T.; Lu, S.E.; Liou, N.S. Development of sorting and grading methodology of jujubes using hyperspectral image data. Postharvest Biol. Technol. 2025, 222, 113406. [Google Scholar] [CrossRef] [Scilit]
  105. Ukwuoma, C.C.; Zhiguang, Q.; Bin Heyat, M.B.; Ali, L.; Almaspoor, Z.; Monday, H.N. Recent Advancements in Fruit Detection and Classification Using Deep Learning Techniques. Math. Probl. Eng. 2022, 2022, 1–29. [Google Scholar] [CrossRef] [Scilit]
  106. Dang, M.; Wang, H.; Li, Y.; Nguyen, T.H.; Tightiz, L.; Xuan-Mung, N.; Nguyen, T.N. Computer Vision for Plant Disease Recognition: A Comprehensive Review. Bot. Rev. 2024, 90, 251–311. [Google Scholar] [CrossRef] [Scilit]
  107. Wang, L.; Ye, R.; Chen, Y.; Li, T. YOLOv10-LGDA: An Improved Algorithm for Defect Detection in Citrus Fruits Across Diverse Backgrounds. Plants 2025, 14, 1990. [Google Scholar] [CrossRef] [Scilit]
  108. Zhang, W.; Chen, K.; Wang, J.; Shi, Y.; Guo, W. Easy domain adaptation method for filling the species gap in deep learning-based fruit detection. Hortic. Res. 2021, 8, 119. [Google Scholar] [CrossRef] [Scilit]
  109. Kumar Pothula, A.; Zhang, Z.; Lu, R. Evaluation of a new apple in-field sorting system for fruit singulation, rotation and imaging. Comput. Electron. Agric. 2023, 208, 107789. [Google Scholar] [CrossRef] [Scilit]
  110. Moallem, P.; Serajoddin, A.; Pourghassem, H. Computer vision-based apple grading for golden delicious apples based on surface features. Inf. Process. Agric. 2017, 4, 33–40. [Google Scholar] [CrossRef] [Scilit]
  111. Costa, C.; Antonucci, F.; Pallottino, F.; Aguzzi, J.; Sun, D.W.; Menesatti, P. Shape Analysis of Agricultural Products: A Review of Recent Research Advances and Potential Application to Computer Vision. Food Bioprocess Technol. 2011, 4, 673–692. [Google Scholar] [CrossRef] [Scilit]
  112. Li, Z.; Li, Y.; Zhao, H.; Huang, L.; Zhao, Z.; Liao, J.; Wang, M.; Wu, X.; Gong, M.; He, Z.; et al. A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM. Plants 2026, 15, 2103. [Google Scholar] [CrossRef] [Scilit]
  113. Alonso, J.; Aranda, M.; Córdoba, M.G.; Ruiz-Moyano, S.; Velázquez, R.; Aranda, E.; Martín, A. VIS–SWIR hyperspectral imaging for campaign-year verification and moisture screening in smoked paprika (Capsicum annuum L.). J. Food Compos. Anal. 2026, 153, 109097. [Google Scholar] [CrossRef] [Scilit]
  114. Rahman, A.; Faqeerzada, M.A.; Joshi, R.; Lohumi, S.; Kandpal, L.M.; Lee, H.; Mo, C.; Kim, M.S.; Cho, B.K. Quality Analysis of Stored Bell Peppers Using Near-Infrared Hyperspectral Imaging. Trans. ASABE 2018, 61, 1199–1207. [Google Scholar] [CrossRef] [Scilit]
  115. Zsom, T.; Zsom-Muha, V.; Le Nguyen, L.P.; Nagy, D.; Hitka, G.; Polgári, P.; Baranyai, L. Nondestructive detection of low temperature induced stress on post-harvest quality of kápia type sweet pepper. Prog. Agric. Eng. Sci. 2021, 16, 173–186. [Google Scholar] [CrossRef] [Scilit]
  116. Mei, M.; Li, J. An overview on optical nondestructive detection of bruises in fruit: Technology, method, application, challenge and trend. Comput. Electron. Agric. 2023, 213, 108195. [Google Scholar] [CrossRef] [Scilit]
  117. Aktaş, H.; Karagöz, Ö. Sorting and counting of almond kernels on conveyor belt using computer vision and deep learning techniques. Postharvest Biol. Technol. 2025, 228, 113654. [Google Scholar] [CrossRef] [Scilit]
  118. Sofu, M.M.; Er, O.; Kayacan, M.C.; Cetişli, B. Design of an automatic apple sorting system using machine vision. Comput. Electron. Agric. 2016, 127, 395–405. [Google Scholar] [CrossRef] [Scilit]
  119. Zhou, W.; Song, C.; Song, K.; Wen, N.; Sun, X.; Gao, P. Surface Defect Detection System for Carrot Combine Harvest Based on Multi-Stage Knowledge Distillation. Foods 2023, 12, 793. [Google Scholar] [CrossRef] [Scilit]
  120. Elkaoud, N.S.M.; Mahmoud, R.K. Design and implementation of sequential fruit size sorting machine. Rev. Bras. Eng. Agrícola E Ambient. 2022, 26, 722–728. [Google Scholar] [CrossRef] [Scilit]
  121. Chen, Y.; An, X.; Gao, S.; Li, S.; Kang, H. A Deep Learning-Based Vision System Combining Detection and Tracking for Fast On-Line Citrus Sorting. Front. Plant Sci. 2021, 12, 622062. [Google Scholar] [CrossRef] [Scilit]
  122. Liu, W.; Wang, S.; Gao, X.; Yang, H. A Tomato Recognition and Rapid Sorting System Based on Improved YOLOv10. Machines 2024, 12, 689. [Google Scholar] [CrossRef] [Scilit]
  123. Mendez, E.; Escobedo Cabello, J.A.; Gómez-Espinosa, A.; Cantoral-Ceballos, J.A.; Ochoa, O. Capsicum Counting Algorithm Using Infrared Imaging and YOLO11. Agriculture 2025, 15, 2574. [Google Scholar] [CrossRef] [Scilit]
  124. Sa, I.; Lehnert, C.; English, A.; McCool, C.; Dayoub, F.; Upcroft, B.; Perez, T. Peduncle Detection of Sweet Pepper for Autonomous Crop Harvesting—Combined Color and 3-D Information. IEEE Robot. Autom. Lett. 2017, 2, 765–772. [Google Scholar] [CrossRef] [Scilit]
  125. Jiang, D.; Gu, S.; Chu, Q.; Yang, Y.; Gu, M.; Yang, Y. Pneumatic separation system for collected seedlings using subdivided air streams. Int. J. Agric. Biol. Eng. 2022, 15, 84–92. [Google Scholar] [CrossRef] [Scilit]
  126. Cao, M.; Zhang, J.; Sun, Y.; Zhu, J.; Hu, Y. A novel air-suction classifier for fresh sphere fruits in pneumatic bulk grading. J. Food Meas. Charact. 2023, 17, 3390–3402. [Google Scholar] [CrossRef] [Scilit]
  127. Cardona, C.I.; Tinoco, H.A.; Perdomo-Hurtado, L.; Duque-Dussán, E.; Banout, J. Optimizing Harvesting Efficiency: Development and Assessment of a Pneumatic Air Jet Excitation Nozzle for Delicate Biostructures in Food Processing. Foods 2024, 13, 1458. [Google Scholar] [CrossRef] [Scilit]
  128. Hu, C.; Liu, D.; Ji, J.; Zheng, S.; Li, Y.; Li, J. Design and Simulation of an Underactuated End-Effector for Low-Damage Sweet Pepper Harvesting. INMATEH Agric. Eng. 2025, 77, 502–513. [Google Scholar] [CrossRef] [Scilit]
  129. Hui, Y.; Liao, Y.; Wang, D.; Li, X.; Lu, Z.; You, Y.; Jia, H. Design and Experiment of an End-Effector for Harvesting Sweet Pepper in Compliant Obstacle Environment. INMATEH Agric. Eng. 2023, 71, 271–281. [Google Scholar] [CrossRef] [Scilit]
  130. Xie, S.; Mei, J.; Liu, H.; Wang, Y. Hysteresis modeling and trajectory tracking control of the pneumatic muscle actuator using modified Prandtl–Ishlinskii model. Mech. Mach. Theory 2018, 120, 213–224. [Google Scholar] [CrossRef] [Scilit]
  131. Tian, Z.; Wei, Z.; Liu, Y.; Su, G.; Cui, Z.; Wang, F.; Zhang, X.; Wang, X.; Cheng, X.; Wang, X.; et al. Experimental study on flow field distribution and material separation characteristics in a pneumatic potato impurity removal device based on CFD-DEM coupling. Comput. Electron. Agric. 2026, 253, 112163. [Google Scholar] [CrossRef] [Scilit]
  132. Zournatzis, I.; Kalaitzakis, S.; Polygerinos, P. SoftER: A Spiral Soft Robotic Ejector for Sorting Applications. IEEE Robot. Autom. Lett. 2023, 8, 7098–7105. [Google Scholar] [CrossRef] [Scilit]
  133. Maier, G.; Pfaff, F.; Pieper, C.; Gruna, R.; Noack, B.; Kruggel-Emden, H.; Langle, T.; Hanebeck, U.D.; Wirtz, S.; Scherer, V.; et al. Experimental Evaluation of a Novel Sensor-Based Sorting Approach Featuring Predictive Real-Time Multiobject Tracking. IEEE Trans. Ind. Electron. 2021, 68, 1548–1559. [Google Scholar] [CrossRef] [Scilit]
  134. Zhang, R.; Lu, W.; Jian, X.; Luo, H. Intelligent sorting method for assembly line based on visual positioning and model predictive control of robotic arm. Int. J. Agric. Biol. Eng. 2023, 16, 206–214. [Google Scholar] [CrossRef] [Scilit]
  135. Arad, B.; Balendonck, J.; Barth, R.; Ben-Shahar, O.; Edan, Y.; Hellström, T.; Hemming, J.; Kurtser, P.; Ringdahl, O.; Tielen, T.; et al. Development of a sweet pepper harvesting robot. J. Field Robot. 2020, 37, 1027–1039. [Google Scholar] [CrossRef] [Scilit]
  136. Polic, M.; Tabak, J.; Orsag, M. Pepper to fall: A perception method for sweet pepper robotic harvesting. Intell. Serv. Robot. 2022, 15, 193–201. [Google Scholar] [CrossRef] [Scilit]
  137. Song, Z.; Du, C.; Chen, Y.; Han, D.; Wang, X. Development and test of a spring-finger roller-type hot pepper picking header. J. Agric. Eng. 2024, 55, 1562. [Google Scholar] [CrossRef] [Scilit]
  138. Nuño-Maganda, M.A.; Dávila-Rodríguez, I.A.; Hernández-Mier, Y.; Barrón-Zambrano, J.H.; Elizondo-Leal, J.C.; Díaz-Manriquez, A.; Polanco-Martagón, S. Real-Time Embedded Vision System for Online Monitoring and Sorting of Citrus Fruits. Electronics 2023, 12, 3891. [Google Scholar] [CrossRef] [Scilit]
  139. Hemamalini, V.; Rajarajeswari, S.; Nachiyappan, S.; Sambath, M.; Devi, T.; Singh, B.K.; Raghuvanshi, A. Food Quality Inspection and Grading Using Efficient Image Segmentation and Machine Learning-Based System. J. Food Qual. 2022, 2022, 5262294. [Google Scholar] [CrossRef] [Scilit]
  140. Mukhiddinov, M.; Muminov, A.; Cho, J. Improved Classification Approach for Fruits and Vegetables Freshness Based on Deep Learning. Sensors 2022, 22, 8192. [Google Scholar] [CrossRef] [Scilit]
  141. Schmitt, J.; Bönig, J.; Borggräfe, T.; Beitinger, G.; Deuse, J. Predictive model-based quality inspection using Machine Learning and Edge Cloud Computing. Adv. Eng. Inform. 2020, 45, 101101. [Google Scholar] [CrossRef] [Scilit]
  142. Pasache, H.; Tuesta, C.; Inga, C. Design of an Automated System for Classifying Maturation Stages of Erythrina edulis Beans Using Computer Vision and Convolutional Neural Networks. AgriEngineering 2025, 7, 277. [Google Scholar] [CrossRef] [Scilit]
  143. Ireri, D.; Belal, E.; Okinda, C.; Makange, N.; Ji, C. A computer vision system for defect discrimination and grading in tomatoes using machine learning and image processing. Artif. Intell. Agric. 2019, 2, 28–37, Erratum in Artif. Intell. Agric. 2021, 5, 301–302. https://doi.org/10.1016/j.aiia.2020.12.001. [Google Scholar] [CrossRef] [Scilit]
  144. Du, P.; Han, W.; Xu, Y.; Hu, W.; Xiang, Y. Real-time monitoring of pepper harvesting loss using a lightweight vision model. Smart Agric. Technol. 2026, 14, 102197. [Google Scholar] [CrossRef] [Scilit]
  145. Lu, L.; Lei, J.; Cheng, C.; Wang, S.; Wang, C.; Qin, X. Impurity rates detection for pepper harvesting based on YOLOv8n-Seg-ASB and random forest. Smart Agric. Technol. 2025, 12, 101224. [Google Scholar] [CrossRef] [Scilit]
  146. Wang, J.; Huo, Y.; Wang, Y.; Zhao, H.; Li, K.; Liu, L.; Shi, Y. Grading detection of “Red Fuji” apple in Luochuan based on machine vision and near-infrared spectroscopy. PLoS ONE 2022, 17, e0271352. [Google Scholar] [CrossRef] [Scilit]
  147. Ji, W.; Wang, J.; Xu, B.; Zhang, T. Apple Grading Based on Multi-Dimensional View Processing and Deep Learning. Foods 2023, 12, 2117. [Google Scholar] [CrossRef] [Scilit]
  148. Zhang, L.; Luo, P.; Ding, S.; Li, T.; Qin, K.; Mu, J. The grading detection model for fingered citron slices (citrus medica ‘fingered’) based on YOLOv8-FCS. Front. Plant Sci. 2024, 15, 1411178. [Google Scholar] [CrossRef] [Scilit]
  149. Xu, J.; Lu, Y. Development and systematic evaluation of an innovative multispectral vision-based sweetpotato grading and sorting system. Smart Agric. Technol. 2026, 14, 102204. [Google Scholar] [CrossRef] [Scilit]
  150. Romaniello, R.; Barrasso, A.E.; Perone, C.; Tamborrino, A.; Berardi, A.; Leone, A. Optimisation of an Industrial Optical Sorter of Legumes for Gluten-Free Production Using Hyperspectral Imaging Techniques. Foods 2024, 13, 404. [Google Scholar] [CrossRef] [Scilit]
  151. Wang, J.; Guo, Z.; Zou, C.; Jiang, S.; El-Seedi, H.R.; Zou, X. General model of multi-quality detection for apple from different origins by Vis/NIR transmittance spectroscopy. J. Food Meas. Charact. 2022, 16, 2582–2595. [Google Scholar] [CrossRef] [Scilit]
  152. Liu, Y.; Han, X.; Ren, L.; Ma, W.; Liu, B.; Sheng, C.; Song, Y.; Li, Q. Surface Defect and Malformation Characteristics Detection for Fresh Sweet Cherries Based on YOLOv8-DCPF Method. Agronomy 2025, 15, 1234. [Google Scholar] [CrossRef] [Scilit]
  153. Ortenzi, L.; Figorilli, S.; Costa, C.; Pallottino, F.; Violino, S.; Pagano, M.; Imperi, G.; Manganiello, R.; Lanza, B.; Antonucci, F. A Machine Vision Rapid Method to Determine the Ripeness Degree of Olive Lots. Sensors 2021, 21, 2940. [Google Scholar] [CrossRef] [Scilit]
  154. Zhang, Z.; Cheng, H.; Chen, M.; Zhang, L.; Cheng, Y.; Geng, W.; Guan, J. Detection of Pear Quality Using Hyperspectral Imaging Technology and Machine Learning Analysis. Foods 2024, 13, 3956. [Google Scholar] [CrossRef] [Scilit]
  155. Liu, Z.; Zhao, J.; Zheng, W.; Song, Q.; Zhang, X.; Liu, W.; Shan, F.; Xu, R.; Li, Z.; Dong, J.; et al. Machine vision-based detection of browning maturity in shiitake cultivation sticks. Front. Plant Sci. 2025, 16, 1676977. [Google Scholar] [CrossRef] [Scilit]
  156. Zhang, J.; Pu, J.; An, T.; Wu, P.; Zhou, H.; Niu, Q.; Li, C.; Wang, L. Monitoring of impurities in green peppers based on convolutional neural networks. Signal Image Video Process. 2024, 18, 63–69. [Google Scholar] [CrossRef] [Scilit]
  157. Soltani Firouz, M.; Sardari, H. Defect Detection in Fruit and Vegetables by Using Machine Vision Systems and Image Processing. Food Eng. Rev. 2022, 14, 353–379. [Google Scholar] [CrossRef] [Scilit]
  158. Sun, C.; Zhang, L.; Zhai, L.; Shen, T.; Cai, J.; Zou, X.; Guo, Z. Automatic early bruise detection in strawberry fruit by hyperspectral imaging and deep learning techniques. Postharvest Biol. Technol. 2026, 232, 113966. [Google Scholar] [CrossRef] [Scilit]
  159. Li, Q.; Chen, K.; Wang, Q.; Wang, F.; Deng, W. Enhanced YOLOv11n: A Method for Potato Peel Damage Detection. Food Sci. Nutr. 2025, 13, e70576. [Google Scholar] [CrossRef] [Scilit]
  160. Li, J.; Luo, W.; Han, L.; Cai, Z.; Guo, Z. Two-wavelength image detection of early decayed oranges by coupling spectral classification with image processing. J. Food Compos. Anal. 2022, 111, 104642. [Google Scholar] [CrossRef] [Scilit]
  161. Ren, J.; Xiong, Y.; Chen, X.; Hao, Y. Comparative Analysis of Machine Learning and Deep Learning Algorithms for Assessing Agricultural Product Quality Using NIRS. Sensors 2024, 24, 5438. [Google Scholar] [CrossRef] [Scilit]
  162. Garillos-Manliguez, C.A.; Chiang, J.Y. Multimodal Deep Learning and Visible-Light and Hyperspectral Imaging for Fruit Maturity Estimation. Sensors 2021, 21, 1288. [Google Scholar] [CrossRef] [Scilit]
  163. Guo, Z.; Xiao, H.; Dai, Z.; Wang, C.; Sun, C.; Watson, N.; Povey, M.; Zou, X. Identification of apple variety using machine vision and deep learning with Multi-Head Attention mechanism and GLCM. J. Food Meas. Charact. 2025, 19, 6540–6558. [Google Scholar] [CrossRef] [Scilit]
  164. Wu, X.; Wu, B.; Sun, J.; Li, M.; Du, H. Discrimination of Apples Using Near Infrared Spectroscopy and Sorting Discriminant Analysis. Int. J. Food Prop. 2016, 19, 1016–1028. [Google Scholar] [CrossRef] [Scilit]
  165. Wu, X.; Wu, B.; Sun, J.; Yang, N. Classification of Apple Varieties Using Near Infrared Reflectance Spectroscopy and Fuzzy Discriminant C-Means Clustering Model. J. Food Process Eng. 2017, 40, e12355. [Google Scholar] [CrossRef] [Scilit]
  166. Behera, S.K.; Rath, A.K.; Sethy, P.K. Maturity status classification of papaya fruits based on machine learning and transfer learning approach. Inf. Process. Agric. 2021, 8, 244–250. [Google Scholar] [CrossRef] [Scilit]
  167. Li, H.; Zhu, Q.; Huang, M.; Guo, Y.; Qin, J. Pose Estimation of Sweet Pepper through Symmetry Axis Detection. Sensors 2018, 18, 3083. [Google Scholar] [CrossRef] [Scilit]
  168. Guo, L.; Gao, Q.; Wang, Y.; Bao, Y.; Liu, Z.; Liu, H.; Huang, J.; Zheng, M.; Wang, L.; Yu, J.; et al. Design, optimization, and validation of a high-speed delta parallel system for soft fruit sorting. Smart Agric. Technol. 2026, 14, 102031. [Google Scholar] [CrossRef] [Scilit]
  169. Ma, P.; Sun, J.; Cong, S.; Dai, C.; Cai, Z.; Yao, K.; Zhou, X.; Wu, X.; Liu, J. Detection of Early Damage in Kiwifruit Based on Near-Infrared Technology. J. Food Process Eng. 2025, 48, e70130. [Google Scholar] [CrossRef] [Scilit]
  170. Naqvi, L.H.; Balasubramaniam, B.; Li, J.; Liu, L.; Li, B. Four-Dimensional Hyperspectral Imaging for Fruit and Vegetable Grading. Agriculture 2025, 15, 1702. [Google Scholar] [CrossRef] [Scilit]
  171. Rahman, K.S.; Rakib, M.R.I.; Rahman, A.; Khaled, A.A.; Salehin, M.M.; Amin, A.; Rahman, A. Rapid Assessment of Fruits and Vegetables Quality Using Spectroscopy and Imaging Techniques: Recent Advances and Perspectives. Food Eng. Rev. 2025, 17, 777–808. [Google Scholar] [CrossRef] [Scilit]
  172. Batista, C.B.; Soldateli, F.J.; Fehndrich, S.P.; Bittencourt, M.N.; Silva, L.S.; Ethur, L.Z. Maturation and determination of the harvest point of Capsicum chinense Jacq. (‘biquinho’ pepper). Rev. Bras. Ciências Agrárias—Braz. J. Agric. Sci. 2022, 17, e161. [Google Scholar] [CrossRef] [Scilit]
  173. Chu, Y.; Feng, D.; Liu, Z.; Zhao, Z.; Wang, Z.; Xia, X.G.; Quek, T.Q.S. Hybrid-Learning-Based Operational Visual Quality Inspection for Edge-Computing-Enabled IoT System. IEEE Internet Things J. 2022, 9, 4958–4972. [Google Scholar] [CrossRef] [Scilit]
  174. Li, J.; He, L.; Liu, M.; Chen, J.; Xue, L. Hyperspectral dimension reduction and navel orange surface disease defect classification using independent component analysis-genetic algorithm. Front. Nutr. 2022, 9, 993737. [Google Scholar] [CrossRef] [Scilit]
  175. Lun, Z.; Wu, X.; Dong, J.; Wu, B. Deep Learning-Enhanced Spectroscopic Technologies for Food Quality Assessment: Convergence and Emerging Frontiers. Foods 2025, 14, 2350. [Google Scholar] [CrossRef] [Scilit]
  176. Sun, H.; Chu, H.Q.; Qin, Y.M.; Hu, P.; Wang, R.F. Empowering Smart Soybean Farming with Deep Learning: Progress, Challenges, and Future Perspectives. Agronomy 2025, 15, 1831. [Google Scholar] [CrossRef] [Scilit]
  177. Li, X.; Wang, F.; Guo, Y.; Liu, Y.; Lv, H.; Zeng, F.; Lv, C. Improved YOLO v5s-based detection method for external defects in potato. Front. Plant Sci. 2025, 16, 1527508. [Google Scholar] [CrossRef] [Scilit]
  178. Aline, U.; Bhattacharya, T.; Faqeerzada, M.A.; Kim, M.S.; Baek, I.; Cho, B.K. Advancement of nondestructive spectral measurements for the quality of major tropical fruits and vegetables: A review. Front. Plant Sci. 2023, 14, 1240361. [Google Scholar] [CrossRef] [Scilit]
  179. Lorente, D.; Aleixos, N.; Gómez-Sanchis, J.; Cubero, S.; García-Navarrete, O.L.; Blasco, J. Recent Advances and Applications of Hyperspectral Imaging for Fruit and Vegetable Quality Assessment. Food Bioprocess Technol. 2012, 5, 1121–1142. [Google Scholar] [CrossRef] [Scilit]
  180. Wieme, J.; Mollazade, K.; Malounas, I.; Zude-Sasse, M.; Zhao, M.; Gowen, A.; Argyropoulos, D.; Fountas, S.; Van Beek, J. Application of hyperspectral imaging systems and artificial intelligence for quality assessment of fruit, vegetables and mushrooms: A review. Biosyst. Eng. 2022, 222, 156–176. [Google Scholar] [CrossRef] [Scilit]
  181. Walsh, K.B.; Blasco, J.; Zude-Sasse, M.; Sun, X. Visible-NIR ‘point’ spectroscopy in post-harvest fruit and vegetable assessment: The science behind three decades of commercial use. Postharvest Biol. Technol. 2020, 168, 111246. [Google Scholar] [CrossRef] [Scilit]
  182. Shu, Z.; Li, X.; Liu, Y. Detection of Chili Foreign Objects Using Hyperspectral Imaging Combined with Chemometric and Target Detection Algorithms. Foods 2023, 12, 2618. [Google Scholar] [CrossRef] [Scilit]
  183. Shao, Y.; Ji, S.; Shi, Y.; Xuan, G.; Jia, H.; Guan, X.; Chen, L. Growth period determination and color coordinates visual analysis of tomato using hyperspectral imaging technology. Spectrochim. Acta Part A Mol. Biomol. Spectrosc. 2024, 319, 124538. [Google Scholar] [CrossRef] [Scilit]
  184. Zhao, P.; Wang, X.; Zhao, Q.; Xu, Q.; Sun, Y.; Ning, X. Nondestructive Detection of External Defects in Potatoes Using Hyperspectral Imaging and Machine Learning. Agriculture 2025, 15, 573. [Google Scholar] [CrossRef] [Scilit]
  185. Chaudhary, R.K.; Neupane, A.; Wang, Z.; Walsh, K. Mango Quality Assessment Using Near-Infrared Spectroscopy and Hyperspectral Imaging: A Systematic Review. Agronomy 2025, 15, 2271. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Conceptual framework and technological evolution of intelligent post-harvest color-sorting equipment for peppers.
Figure 1. Conceptual framework and technological evolution of intelligent post-harvest color-sorting equipment for peppers.
Processes 14 02991 g001
Figure 2. Literature screening flowchart.
Figure 2. Literature screening flowchart.
Processes 14 02991 g002
Figure 3. Hierarchical representation of pepper color traits and the boundaries of machine-vision grading tasks. (a) Biological and optical determinants of overall fruit color. (b) Hierarchical description of pepper appearance based on variety or type, ripening color stage, overall color intensity, and surface gloss. (c) Color- and shape-based class separation under controlled imaging conditions. Here, “Labels” denotes categorical annotation values assigned to the five representative fruit-color classes (green, yellow/orange, orange, red, and dark red), which convey differences in color and ripening stage. (d) Distinction among object detection, candidate-class recognition, and commercial grading. Commercial grades require joint evaluation of ripeness, color intensity, and gloss. Conceptual synthesis based on Refs. [9,10,32,33,34,35,36,37,38].
Figure 3. Hierarchical representation of pepper color traits and the boundaries of machine-vision grading tasks. (a) Biological and optical determinants of overall fruit color. (b) Hierarchical description of pepper appearance based on variety or type, ripening color stage, overall color intensity, and surface gloss. (c) Color- and shape-based class separation under controlled imaging conditions. Here, “Labels” denotes categorical annotation values assigned to the five representative fruit-color classes (green, yellow/orange, orange, red, and dark red), which convey differences in color and ripening stage. (d) Distinction among object detection, candidate-class recognition, and commercial grading. Commercial grades require joint evaluation of ripeness, color intensity, and gloss. Conceptual synthesis based on Refs. [9,10,32,33,34,35,36,37,38].
Processes 14 02991 g003
Figure 4. Boundary between direct Capsicum evidence and required system-level validation. (a) Existing chili pepper research supports controlled image acquisition and real-time YOLO-based recognition [67]. (b) A complete pepper color-sorting system still requires validation of continuous target tracking, timed actuator diversion, and end-of-line outcome verification across cultivars, batches, occlusion, illumination, and production speeds. This original synthesis contains no performance values transferred from other crops.
Figure 4. Boundary between direct Capsicum evidence and required system-level validation. (a) Existing chili pepper research supports controlled image acquisition and real-time YOLO-based recognition [67]. (b) A complete pepper color-sorting system still requires validation of continuous target tracking, timed actuator diversion, and end-of-line outcome verification across cultivars, batches, occlusion, illumination, and production speeds. This original synthesis contains no performance values transferred from other crops.
Processes 14 02991 g004
Figure 5. Workflow of the in-field apple grading equipment. Adapted from Ref. [72] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Figure 5. Workflow of the in-field apple grading equipment. Adapted from Ref. [72] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Processes 14 02991 g005
Figure 6. Pneumatically powered mechanisms proposed for online sweet potato sorting: (a) a linear air cylinder driving a pivoting paddle through a linkage; (b) a linear air cylinder directly driving a paddle along a linear trajectory; (c) a rotary pneumatic actuator driving a paddle attached to its shaft. Dashed arrows indicate the paddle trajectories during actuation. Adapted from Ref. [73] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Figure 6. Pneumatically powered mechanisms proposed for online sweet potato sorting: (a) a linear air cylinder driving a pivoting paddle through a linkage; (b) a linear air cylinder directly driving a paddle along a linear trajectory; (c) a rotary pneumatic actuator driving a paddle attached to its shaft. Dashed arrows indicate the paddle trajectories during actuation. Adapted from Ref. [73] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Processes 14 02991 g006
Figure 7. Experimental pipeline used to evaluate sweet potato sorting mechanisms. Adapted from Ref. [73] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Figure 7. Experimental pipeline used to evaluate sweet potato sorting mechanisms. Adapted from Ref. [73] under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Processes 14 02991 g007
Figure 8. Output granularity and engineering trade-offs of methods for detecting visible color anomalies.
Figure 8. Output granularity and engineering trade-offs of methods for detecting visible color anomalies.
Processes 14 02991 g008
Figure 9. Hierarchical framework from complex-scene perception to actionable target-sorting information.
Figure 9. Hierarchical framework from complex-scene perception to actionable target-sorting information.
Processes 14 02991 g009
Figure 10. Functional sequence and evaluation interfaces for feeding, singulation, spreading, and pose control. Conceptual synthesis based on Refs. [109,110,111,112].
Figure 10. Functional sequence and evaluation interfaces for feeding, singulation, spreading, and pose control. Conceptual synthesis based on Refs. [109,110,111,112].
Processes 14 02991 g010
Figure 11. Pneumatic actuation routes and shared validation metrics for intelligent pepper sorting. Conceptual synthesis based on Refs. [125,126,127,128,129,130,131,132].
Figure 11. Pneumatic actuation routes and shared validation metrics for intelligent pepper sorting. Conceptual synthesis based on Refs. [125,126,127,128,129,130,131,132].
Processes 14 02991 g011
Table 1. Evidence types, research endpoints, and applicability boundaries for intelligent pepper color sorting in Section 1.
Table 1. Evidence types, research endpoints, and applicability boundaries for intelligent pepper color sorting in Section 1.
Evidence TypeResearch Object/
Question
Main Technical RouteCommon Evaluation EndpointValue and Limitation
for This Review
[E4] Cross-crop machine-vision reviews [1,13,14,19]Appearance detection, grading, and quality controlRGB imaging, image processing, and machine learningDetection or classification performance; scope of applicationDefine general methods, but do not directly demonstrate applicability to peppers or whole-system performance.
[E4] Post-harvest reviews of sweet pepper [2,22]Post-harvest management, intelligent identification, and production needsNondestructive detection, artificial intelligence, and automationTechnology coverage and research trendsProvide crop-specific context, but address online closed-loop operation and standardized equipment metrics only to a limited extent.
[E2] Pepper maturity/color studies [9,10,11]Maturity stages and color gradesRGB imaging, wavelength selection, and deep learningClassification accuracy, F1 score, mAP, or confusion matrixDirectly support pepper perception routes, but usually do not evaluate final product diversion.
[E2] Studies of pepper external traits and nondestructive information [16,24]External traits or non-visual sorting informationMachine vision and odor/sensor informationFeature extraction or class discriminationBroaden the information sources, but online speed, actuator interfaces, and cross-batch evidence still require verification.
Pepper/sweet-pepper sorting studies: [3,4,5] (E1); [6] (E4)Online grading and automatic diversionVision, conveying, control, and actuationReported sorting accuracy or physical throughput [3,4,5]; simulated circuit operation [6]Refs. [3,4,5] provide physical-system evidence under differing conditions; Ref. [6] is a simulation-only design and does not validate physical sorting.
[E4] Reviews of tomato, citrus, berry, and other crops [20,21,23]Quality detection or processing of other cropsVision, spectroscopy, and data fusionPredictive performance, quality dimensions, or process coverageOffer comparable technical routes, but conclusions require validation for pepper-specific objects and operating conditions.
Note: Evidence categories are defined in Table 4. E1, empirical pepper sorting systems; E2, pepper perception or component studies; E3, original transferable studies in other crops or engineering contexts; E4, reviews, background studies, or simulation-only designs. RGB, red–green–blue; mAP, mean average precision; F1, harmonic mean of precision and recall at a matched evaluation level.
Table 2. Database-specific literature search strategies.
Table 2. Database-specific literature search strategies.
ItemDetails
Web of Science Core Collection (Clarivate)
Date searched22–25 July 2026; last search/update: 25 July 2026
Coverage period1 January 2008 to 25 July 2026
Complete search stringMain query: TS = ((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color sorting” OR sorting OR grading OR maturity OR defect OR impurity) AND (“machine vision” OR “hyperspectral imaging” OR “multispectral imaging” OR “deep learning” OR “sensor fusion”) AND (“online sorting” OR conveyor OR singulation OR positioning OR tracking OR “pneumatic rejection” OR “system integration”))
Supplementary query: TS = ((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color recognition” OR “visual perception” OR “material handling” OR conveying OR positioning OR tracking OR “pneumatic actuation” OR “pneumatic rejection” OR “system integration”))
Search fieldsTopic (TS), comprising title, abstract, author keywords, and Keywords Plus
Language restrictionEnglish
Publication/document type restrictionNo publication/document-type filter was applied during database retrieval. Peer-reviewed journal articles and reviews were prioritized during eligibility screening; eligible full-text conference papers and doctoral dissertations were retained only as supplementary evidence.
Other filtersEnglish-language core evidence; duplicate removal and predefined eligibility screening were performed after export to EndNote.
Scopus (Elsevier)
Date searched22–25 July 2026; last search/update: 25 July 2026
Coverage period1 January 2008 to 25 July 2026
Complete search stringMain query: TITLE-ABS-KEY((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color sorting” OR sorting OR grading OR maturity OR defect OR impurity) AND (“machine vision” OR “hyperspectral imaging” OR “multispectral imaging” OR “deep learning” OR “sensor fusion”) AND (“online sorting” OR conveyor OR singulation OR positioning OR tracking OR “pneumatic rejection” OR “system integration”)) AND PUBYEAR > 2007 AND PUBYEAR < 2027 AND LIMIT-TO(LANGUAGE, “English”)
Supplementary query: TITLE-ABS-KEY((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color recognition” OR “visual perception” OR “material handling” OR conveying OR positioning OR tracking OR “pneumatic actuation” OR “pneumatic rejection” OR “system integration”)) AND PUBYEAR > 2007 AND PUBYEAR < 2027 AND LIMIT-TO(LANGUAGE, “English”)
Search fieldsTitle, abstract, and author/indexed keywords (TITLE-ABS-KEY)
Language restrictionEnglish (LIMIT-TO(LANGUAGE, “English”))
Publication/document type restrictionNo publication/document-type filter was applied during database retrieval. Peer-reviewed journal articles and reviews were prioritized during eligibility screening; eligible full-text conference papers and doctoral dissertations were retained only as supplementary evidence.
Other filtersEnglish-language core evidence; duplicate removal and predefined eligibility screening were performed after export to EndNote.
Agricultural & Environmental Science Collection (AESC; ProQuest)
Date searched22–25 July 2026; last search/update: 25 July 2026
Coverage period1 January 2008 to 25 July 2026
Complete search stringMain query: NOFT((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color sorting” OR sorting OR grading OR maturity OR defect OR impurity) AND (“machine vision” OR “hyperspectral imaging” OR “multispectral imaging” OR “deep learning” OR “sensor fusion”) AND (“online sorting” OR conveyor OR singulation OR positioning OR tracking OR “pneumatic rejection” OR “system integration”)) AND YR(2008–2026) AND LA(English)
Supplementary query: NOFT((“chili” OR “pepper” OR “bell pepper” OR capsicum) AND (“color recognition” OR “visual perception” OR “material handling” OR conveying OR positioning OR tracking OR “pneumatic actuation” OR “pneumatic rejection” OR “system integration”)) AND YR(2008–2026) AND LA(English)
Search fieldsAnywhere except full text (NOFT), including bibliographic and indexing fields
Language restrictionEnglish (LA(English))
Publication/document type restrictionNo publication/document-type filter was applied during database retrieval. Peer-reviewed journal articles and reviews were prioritized during eligibility screening; eligible full-text conference papers and doctoral dissertations were retained only as supplementary evidence.
Other filtersEnglish-language core evidence; duplicate removal and predefined eligibility screening were performed after export to EndNote.
PubMed (National Library of Medicine)
Date searched22–25 July 2026; last search/update: 25 July 2026
Coverage period1 January 2008 to 25 July 2026
Complete search stringMain query: (“chili”[Title/Abstract] OR “pepper”[Title/Abstract] OR “bell pepper”[Title/Abstract] OR capsicum[Title/Abstract]) AND (“color sorting”[Title/Abstract] OR sorting[Title/Abstract] OR grading[Title/Abstract] OR maturity[Title/Abstract] OR defect[Title/Abstract] OR impurity[Title/Abstract]) AND (“machine vision”[Title/Abstract] OR “hyperspectral imaging”[Title/Abstract] OR “multispectral imaging”[Title/Abstract] OR “deep learning”[Title/Abstract] OR “sensor fusion”[Title/Abstract]) AND (“online sorting”[Title/Abstract] OR conveyor[Title/Abstract] OR singulation[Title/Abstract] OR positioning[Title/Abstract] OR tracking[Title/Abstract] OR “pneumatic rejection”[Title/Abstract] OR “system integration”[Title/Abstract]) AND 2008/01/01:2026/07/25[Date-Publication] AND English[Language]
Supplementary query: (“chili”[Title/Abstract] OR “pepper”[Title/Abstract] OR “bell pepper”[Title/Abstract] OR capsicum[Title/Abstract]) AND (“color recognition”[Title/Abstract] OR “visual perception”[Title/Abstract] OR “material handling”[Title/Abstract] OR conveying[Title/Abstract] OR positioning[Title/Abstract] OR tracking[Title/Abstract] OR “pneumatic actuation”[Title/Abstract] OR “pneumatic rejection”[Title/Abstract] OR “system integration”[Title/Abstract]) AND 2008/01/01:2026/07/25[Date-Publication] AND English[Language]
Search fieldsTitle/Abstract field tags for concepts; Publication Date and Language tags for limits
Language restrictionEnglish (English[Language])
Publication/document type restrictionNo publication/document-type filter was applied during database retrieval. Peer-reviewed journal articles and reviews were prioritized during eligibility screening; eligible full-text conference papers and doctoral dissertations were retained only as supplementary evidence.
Other filtersEnglish-language core evidence; duplicate removal and predefined eligibility screening were performed after export to EndNote.
Table 3. Operational criteria and scoring rules for study quality assessment.
Table 3. Operational criteria and scoring rules for study quality assessment.
DimensionScore 1: Criterion MetScore 0: Criterion Not MetEvaluation Basis
Study design and operational contextThe research object, task, study setting, experimental or operating conditions, and comparison or validation design were all described sufficiently to interpret the findings.One or more essential elements were absent or too unclear to determine how the study was designed or under which conditions the findings were obtained.Study design, materials and methods, and reported operating conditions
Sample size and independence of experimental unitsThe sample size and experimental unit were reported, and sampling, replication, or data partitioning preserved independence between training, validation, and test observations where applicable.Sample size or the experimental unit was not reported, or non-independence, pseudoreplication, or leakage between data subsets could not be ruled out.Sample description, replication structure, and dataset-partitioning procedure
Replication or validationThe study reported repeated trials, independent batches, temporally or spatially separate validation, an independent test set, or a clearly defined continuous-operation test appropriate to its design.Only a single unreplicated trial was reported, or the duration, number of repetitions, validation set, or validation procedure was absent or unclear.Number and duration of trials, batch structure, and internal or external validation procedure
Completeness of methods and results reportingAcquisition conditions, processing or model procedures, outcome definitions, evaluation metrics, and quantitative results were reported sufficiently to understand and assess the study.Essential methodological information, endpoint definitions, evaluation metrics, or quantitative results were missing or incomprehensible.Methods, parameter descriptions, endpoint definitions, tables, figures, and reported results
Overall decision ruleUnweighted total = sum of the four binary scores (range 0–4). Total 3–4: eligible for core synthesis. Total 0–2: excluded from core synthesis; a study with a unique and clearly defined engineering insight could be retained only as background evidence. Missing or unclear information was scored 0.
Table 4. Evidence hierarchy used for synthesis and interpretation.
Table 4. Evidence hierarchy used for synthesis and interpretation.
LevelEvidence DefinitionTypical EndpointsConclusions SupportedMain Limitations
E1Online complete machines or prototypes for chili or sweet pepper, with classification results linked to physical diversionOverall final-bin sorting accuracy, system false-rejection rate, system missed-rejection rate, throughput, and end-to-end latencySupports the feasibility of pepper equipment under the reported objects and operating conditionsCannot be extrapolated to untested cultivars, batches, speeds, or long-term operation
E2Pepper or sweet-pepper perception, localization, actuation, or controlled component experimentsClassification, detection, segmentation, localization, or component-response metricsSupports information separability or local feasibility for a specific moduleCannot independently demonstrate complete-system sorting performance
E3Original studies of other crops or non-pepper engineering systems, including task-specific offline models, components, and integrated sorting systemsModel, component, or system metrics, with the evaluation endpoint identified separatelySupports transfer hypotheses for task-specific methods, engineering components, and evaluation designsRequires pepper-specific revalidation; offline or component results do not establish physical sorting performance
E4General reviews, background quality or mechanism studies, and simulation-only designs without empirical task-level validationPrinciple-related parameters, trends, mechanisms, or research boundariesUsed to explain principles, identify risks, and formulate questions requiring validationMust not be used as direct evidence of pepper-equipment performance
Note: E1–E4 describe evidence applicability, not the methodological-quality score in Table 3. E3 includes original offline and component studies but does not imply a validated sorting line. Simulation-only designs are E4. Each cited source is assigned separately in mixed-source rows. Across the evidence-summary tables, classification accuracy evaluates model labels; sorting accuracy evaluates final physical assignments. Separation efficiency retains the source-specific material, success criterion, and denominator. Model metrics are not converted to MRsys or FRRsys without matching final-bin counts and a declared positive class (Section 7.4).
Table 5. Cross-study comparison of evidence supporting post-harvest pepper color sorting.
Table 5. Cross-study comparison of evidence supporting post-harvest pepper color sorting.
EvidenceObject/StateResearch TaskSensingOperating ConditionValidation LevelMain Boundary
Ref. [5] (E1);
Ref. [29] (E2)
Red chili; whole fruitColor/grading and conveyor sortingRGB imageControlled or conveyor-basedPepper conveyor prototype [5]; offline model [29]Limited reporting of end-to-end errors
[E2] Refs. [9,10]Bell pepper; maturity stagesMaturity estimation/classificationRGB or hyperspectralControlled imagingPepper model evidenceNot equivalent to commercial grading
[E2] Refs. [43,44,45]Bell/sweet pepper; damage or fluorescence responseEarly damage/defect detectionFluorescence or hyperspectralControlled experimentsPepper sensing evidenceOnline rejection not validated
[E2] Refs. [39,40]Intact bell pepper; maturity/freshnessSensor comparison or fusionImage, fluorescence, Vis–NIRLaboratory/post-harvestPepper multimodal evidenceDeployment cost and speed unclear
[E3] Refs. [30,31]Tomato; sorting/surface defectsOnline system or defect detectionRGB imageOnline/complex backgroundCross-crop engineering evidenceTask and error costs differ from pepper
Refs. [49,52] (E2); Refs. [48,50,51] (E3)Potato/tomato/sweet pepper; occlusion or illuminationDetection, registration and visibilityRGB, NIR, thermal, depthField/greenhouse/online contextsMixed direct and cross-context evidenceNot jointly validated in pepper conveyor flow
Note: Evidence is classified by research object, task, operating condition and validation level. Cross-crop or harvest-scene studies are used only as engineering references and do not establish acceptance thresholds for Capsicum conveyor sorting. Abbreviations: RGB, red–green–blue; NIR, near-infrared; Vis–NIR, visible–near-infrared.
Table 6. Recommended functional requirements and system-level performance metrics for intelligent color-sorting equipment.
Table 6. Recommended functional requirements and system-level performance metrics for intelligent color-sorting equipment.
Functional LevelFunctional RequirementRecommended MetricsEvidence and Comparisons from the Selected Literature
Functional levelPerform object detection, color/maturity recognition, grade assignment, and sorting actuationCompleteness of the functional closed loop; traceability of class definitions and outputsRef. [6] (E4) simulates color sensing, object detection, and rejection; Refs. [3,4] (E1) evaluate integrated pepper sorting with maturity, size, or five-grade outputs.
Recognition levelDistinguish visible color defects and adjacent maturity gradesModel classification accuracy, precision, recall, specificity and F1, with the positive class, counting unit and class-averaging rule; separate model and final-bin confusion matricesRef. [3] (E1): reported accuracy 93.2%, sensitivity 84.0%. Ref. [4] (E1): five-grade overall accuracy 96.9%; separate accuracy 98.7%, precision 97.0%, sensitivity 96.9%, specificity 99.0%, F1 96.9%. Aggregation is not harmonized.
Detection/segmentation levelLocalize targets and extract valid regions under complex orientations and backgroundsTask-specific detection/segmentation metrics, mAP@0.5, and background-separation capabilityRef. [55] (E3) reports segmentation mAP@0.5; however, the study concerns fresh Sichuan pepper (Zanthoxylum), and the result requires validation on Capsicum datasets.
Real-time performance levelEnable continuous sensing and decision-making during conveyingPer-sample inference latency, end-to-end latency, camera frame rate, and actuator response timeRefs. [3,4] (E1) report approximately 0.2 s per sample and 4 ms per sample, respectively; the reported times represent different processing scopes and should not be compared without matching end-to-end definitions; both studies report approximately 3000 samples/h.
Throughput levelMaintain stable feeding, conveying, and sorting cyclesMeasured physical throughput in items/time or mass/time; specify output boundary, channel count, observation window and downtime; report nominal feed capacity separatelyRefs. [3,4] (E1) report approximately 3000 samples/h; Ref. [53] (E3) reports 13.95 t/h after optimization, indicating that throughput should be reported separately according to material and sorting mechanism.
Actuation levelAccurately map vision-based decisions to material-rejection positionsActuation success/failure per issued command; separately, final-bin MRsys and FRRsys and induced damage rate (Section 7.4)Refs. [3,4] (E1) demonstrate online integration, but recognition and actuation errors require separate reporting. Ref. [6] (E4) provides simulated control logic only.
Deployment levelMaintain operational performance under limited computing resources and varying operating conditionsModel size, frames per second (FPS), computing/memory requirements, and performance retention across batches and lighting conditionsRef. [54] (E4) emphasizes multi-view/spectral imaging, lightweight models, and real-time data; Ref. [55] (E3) reports a 5.84 MB model and 98.34% maturity-classification accuracy, but the study concerns Zanthoxylum rather than Capsicum and cannot directly define acceptance criteria for pepper sorting.
Note: The metrics are derived from task definitions, online validation results, and engineering optimization variables in the selected literature; they are not validated acceptance thresholds for Capsicum-sorting equipment. Cross-crop evidence, particularly results for potato, apple, and Sichuan pepper (Zanthoxylum), is included only as engineering context. Thresholds should be calibrated for the target cultivar, sample-size distribution, conveyor speed, imaging configuration, and actuator parameters. Abbreviations: F1, harmonic mean of precision and recall at a matched evaluation level; mAP, mean average precision; mAP@0.5, mAP at an intersection-over-union threshold of 0.5; FPS, frames per second; MRsys, system missed-rejection rate; FRRsys, system false-rejection rate. MRsys counts rejection-required fruit retained divided by all rejection-required fruit; FRRsys counts qualified fruit rejected divided by all qualified fruit. Both use final-bin outcomes (Section 7.4).
Table 7. Technical paradigms, representative evidence, and system-level implications for color sorting.
Table 7. Technical paradigms, representative evidence, and system-level implications for color sorting.
ParadigmDecision BasisRepresentative EvidenceSystem ContributionPrincipal Engineering Need
Human-integrated judgement [E4]Operator synthesis of color, shape, and visible conditionManual grading and reference labelling [56,58,59,60,61]Flexible handling of atypical appearanceInter-rater agreement and stable grade boundaries
Rule-based 2D vision [E3/E4]Controlled color channels, handcrafted features, and thresholdsRobot and conveyor prototypes [56,57,62,63]Repeatable sensing and explicit control logicRobust acquisition and final-bin validation
Learned 2D vision [E2/E3]Data-driven features for grading, detection, and segmentationCapsicum and related-crop models [64,65,66,67,68,69,70]Multi-attribute decisions and object localizationBatch generalization and actuator coupling
Multispectral sensing [E3/E4]Selected spectral responses and temporal signaturesCultivar-dependent measurements [61]Access to weakly visible quality informationCalibration, acquisition speed, and cultivar transfer
Three-dimensional representation [E3]Learned geometric structureSB3D-NET soybean study [71]Shape information beyond projected colorContinuous acquisition and industrial throughput
Table 8. Integrated evidence for maturity, color-grade, anomaly, and multimodal perception methods.
Table 8. Integrated evidence for maturity, color-grade, anomaly, and multimodal perception methods.
Perception TaskTechnical RouteRepresentative EvidenceOperational ValueRequired Validation
Maturity and color gradeVisual indices and explicit featuresSweet-pepper and pepper studies [9,83,87]Interpretable grade thresholdsMapping to commercial grade and rejection action
Maturity and color gradeDeep detection and sensor fusionPepper and related-crop studies [10,11,39,40,84,85,88,89,90]Localization and multi-stage recognitionIndependent batches, latency, and final-bin errors
Visible anomalyImage classificationMango classification [91]Low-cost fruit-level screeningFruit-level partitioning and external validation
Visible anomalyObject detection and active learningPear and tomato studies [92,93]Actionable location with scalable classesLong-tail recall and annotation benefit
Visible anomalySemantic or instance segmentationApple studies [94,95,96]Area, boundary, and severityLabel consistency, computation, and multi-view coverage
Low-contrast stateMultispectral and multimodal sensingPepper, citrus, and related studies [10,39,40,44,46,47,97,98]Complementary tissue and defect informationRegistration, calibration, cycle time, and hardware cost
Table 9. System-chain evidence for material handling, localization, timing, and pneumatic execution.
Table 9. System-chain evidence for material handling, localization, timing, and pneumatic execution.
System StageCritical VariableRepresentative EvidenceEngineering ImplicationPriority Endpoint
Feeding and poseSeparation, rotation, visible areaMultichannel conveying and shape studies [109,110,111,112]Throughput reduces images and coverage per targetSeparation rate, pose dispersion, coverage, damage
ConditioningMoisture, dust, temperature, surface stateSpectral, multimodal, and thermal studies [40,113,114,115,116]Input state must be defined before imagingResidual contamination and post-cleaning image stability
Speed and spacingv, s, Δt = s/v, shared resourcesParallel, edge, and mechanical systems [117,118,119,120]Local processing speed does not define line capacitySpeed-accuracy curve and sustained per-channel throughput
Localization and trackingTarget ID, trajectory, coordinate mappingTracking, RGB-D, and execution studies [121,122,123,124]Visual objects must become executable statesID switches, localization error, arrival-window error
Pneumatic executionPressure, geometry, frequency, valve delayArray, pulse, suction, and CFD evidence [125,126,127,128,129,130,131,132]Actuation parameters operate as a coupled systemCommand-to-impact latency, displacement, and damage
Timing and outcomePrediction, compensation, contact, releaseControl and compliant-mechanism studies [133,134,135,136,137]Errors propagate across the complete chainTask completion, final-bin outcome, throughput, damage
Table 10. Interface coverage and deployment routes for vision-execution integration.
Table 10. Interface coverage and deployment routes for vision-execution integration.
Integration RouteInterface CoverageRepresentative EvidenceDeployment ValueMain Unresolved Link
Vision workflowAcquisition to model decisionFood inspection framework [139]Reusable perception pipelineNo physical-interface validation
Online sortingVision, controller, conveyor, and actuatorCitrus and pepper systems [3,4,138]Direct evidence of physical sortingComplete timing and outcome feedback
Mobile-serverWireless transfer and remote inferenceImproved YOLOv4 platform [140]Assistive identificationIndustrial actuation and network jitter
Edge-cloudTCP/IP data, database, and edge inferenceCeleron N2930 platform [141]Data and model managementActuator closure and end-to-end timing
Fully localCamera, local inference, GPIO/PWM controlEmbedded prototype and Sweet Pepper Online [4,142]Low network exposure and direct actuationThermal load, durability, energy, and damage
Table 11. (a) System-level evidence and reporting completeness for intelligent sorting equipment. (b) Cross-study sources of performance variation, applicable conditions, and principal limitations.
Table 11. (a) System-level evidence and reporting completeness for intelligent sorting equipment. (b) Cross-study sources of performance variation, applicable conditions, and principal limitations.
(a)
System or TaskReported EndpointThroughput or TimingOther System DimensionsUse in the Synthesis
Apple, vision + NIR [146]Reported overall grading accuracy 96.67%Physical throughput not reportedFinal-bin MRsys/FRRsys unavailableMulti-source discrimination evidence
Apple, four-lane physical sorting [147]Final-bin accuracy 282/300 (94.00%)32 FPS; four fruits/s; speed-sensitiveEnergy, damage, and sustained reliability not reportedDirect speed-accuracy and final-bin evidence
Fingered citron detector [148]Precision 96.1%; recall 94.9%; mAP@0.5 98.1%130.3 FPS; material throughput not reportedNo physical rejection outcomeModel-level timing evidence
Sweet-pepper online sorter [4]Five-grade in-line overall accuracy 96.9%About 3000 samples/h/channelComplete latency, energy, damage, and availability not reportedMost direct crop-specific system evidence
Integrated multispectral sorter [149]Overall accuracy 94.1% to 91.4% across tested speeds15–35 cm/s; two items/s/channelEnergy, damage, and long-duration reliability not reportedDynamic operating-point evidence
Industrial optical sorter [150]Source-reported CCR 99.36% to 99.67%Repeated trials; end-to-end latency not decomposedTotal energy and material damage not reportedIndustrial repeatability evidence
Green-pepper damage detection [43]Experimentally defined damage classificationNo integrated sorting throughputSorting-induced damage not measuredDamage-detection method, not equipment damage evidence
(b)
Evidence GroupCrop/Cultivar and Task/ClassSample Size and PartitionIllumination and CalibrationPose, Speed, and HardwareEndpoint and Reported PerformanceLikely Variation, Applicability, and Limitations
Traditional RGB/HSV classification [34]Chili; five image categories; commercial reject mapping NR210 training and 90 test images; fruit-level independence, balance, and external batch NRControlled background; light source, calibration, and pose NROffline images; conveyor speed and hardware NRImage-level accuracy 90/90; precision and recall 1.0; averaging NRStable background and bounded classes reduce within-domain variation. Useful as a low-cost baseline, but not evidence of cultivar, batch, or line transfer.
Active and multispectral evidence [43,45,61]Green or sweet pepper damage and apple cultivar response; class boundaries vary by taskSample size and partitioning NR in the present synthesisUV fluorescence or multispectral acquisition; signals depend on pigmentation, geometry, isolation, and calibrationStatic or line speed and processing hardware NRDamage separability or cultivar-dependent spectral overlap; no common endpointAdditional channels help when RGB contrast is weak, but cultivar overlap, calibration, and acquisition time limit transfer and online use.
Integrated apple and 3D model evidence [71,72]Apple multi-feature grading and soybean five-cultivar classification; task and crop differSample size or partition details NR here; [71] reports training and validation endpointsThree-camera imaging [72]; lighting details NR; cultivar overlap and geometric coverage remain relevant1.2 s per fruit with a bottom blind area [72]; system hardware NR; [71] has no line throughput95.49% mean multi-feature and 94.12% field accuracy [72]; 95.54% training and 90.74% validation [71]Differences reflect crop morphology, surface coverage, cultivar diversity, validation level, and endpoint. Neither value supports a direct route ranking.
Classification, detection, segmentation, and active learning [91,92,93,94]Mango, pear, apple, and tomato anomaly tasks; image, box, and pixel labels are not equivalentSample size, fruit-level partitioning, and external batches NR in this synthesisRGB or selected RGB-NIR bands; detailed lighting and calibration conditions vary or are NRPrimarily model-level tests; conveyor speed and deployment hardware NRAccuracy/AUC, mAP@0.5, F-score, and annotation reduction are retained in source-defined formsMetric variation mainly reflects output granularity, label structure, class difficulty, and annotation design. Choose the route by the actuator information required.
Architecture optimization and domain adaptation [107,108]Citrus disease detection and orange-to-apple/tomato transfer; chili classes not testedWithin-task validation [107] and cross-crop transfer [108]; cross-batch and cross-device evidence NRComplex backgrounds or translated domains; matched illumination and calibration NRConveyor speed and hardware platform NRWithin-domain metric gains [107] and target-domain mAP gains [108]Architecture changes address within-domain representation, whereas adaptation addresses domain shift. Neither establishes chili line generalization without independent deployment tests.
Material presentation and edge-speed evidence [109,119]Multi-view material handling and edge sorting; class definitions differSample composition and partitioning NRMotion imaging; calibration details NR; pose and surface coverage change with feedingOne to three targets/s reduced views from 24 to 9 [109]; 0.6 m/s and 100 mm spacing give an ideal 0.167 s interval [119]Views per fruit and derived arrival interval; sustained final-bin output NRSpeed reduces coverage and timing margin. Edge computation helps processing latency but cannot remove motion blur, overlap, or actuator recovery limits.
Tracking and geometric localization [121,122,123,124]Identity tracking, weak-light counting, RGB-D localization, and grasping; state outputs differEight to twenty targets [121] and 70 targets [123]; broader partitioning NRVisible, infrared, or RGB-D sensing; weak-light and geometric conditions differ23 ms/frame [121]; conveyor integration, processor, and full timing boundary NRMOTA/MOTP, target counts, 2D tracks, or 3D action points; no common final-bin endpointInfrared favors weak-light continuity, while RGB-D supplies geometry. Calibration, synchronization, reconstruction, and latency constrain online use.
Online and industrial sorting systems [4,147,149,150]Sweet pepper, apple, multispectral, and industrial optical sorting; grades and positive classes differApple final-bin test 282/300 [147]; other sample and partition details vary or are NROptical configurations differ; calibration and pose reporting are incompleteAbout 3000 samples/h/channel [4]; four fruits/s [147]; 15–35 cm/s and two items/s/channel [149]; hardware details incomplete96.9% five-grade in-line accuracy [4], 94.00% final-bin accuracy [147], 94.1% to 91.4% across speeds [149], and source-defined CCR [150]Variation combines class definition, speed, channel count, hardware, and endpoint. These systems support operating-point evidence, not an overall ranking of technical routes.
NR, not reported in the evidence summarized in this review. Metrics are retained in their source-defined forms and are not used for direct cross-task ranking.
Table 12. Qualitative comparison of evidence boundaries across sensing and modeling routes; values from heterogeneous tasks should not be used for direct performance ranking.
Table 12. Qualitative comparison of evidence boundaries across sensing and modeling routes; values from heterogeneous tasks should not be used for direct performance ranking.
Study and TaskRouteKey EvidenceBoundary and Value for Pepper Sorting
[E3] [161] Agricultural-product quality assessmentNIRS; machine learning vs. deep learningDeep models had higher accuracy and lower spectral-noise sensitivity.Noise robustness was tested, not cross-origin or cross-batch transfer; domain-held-out validation is required.
[E3] [162,166] Papaya maturity classificationVisible-light and hyperspectral fusionBest F1 ≈ 0.90; top-2 error ≈ 1.45%.Random splitting cannot establish domain robustness; test modality complementarity on identical held-out domains.
[E3] [163] Eleven-cultivar apple identificationRGB, MobileNetV2, GLCM texture and attentionBest reported cultivar-classification accuracy was 98.25% (aggregation not verified); one attention variant underperformed the texture model.Shows within-dataset separability, not transfer; complexity does not ensure generalization.
[E3] [164] Two-cultivar apple discriminationNIR with discriminant projectionClass-oriented projection improved separation of the tested cultivars.A useful baseline, but controlled two-cultivar evidence cannot support broad transfer.
[E3] [165] Four-cultivar apple classificationNIR, PCA and fuzzy clusteringReported four-cultivar classification accuracy ≈ 97%; aggregation not verified.Supports screening of known classes; external cultivar and batch validation remains necessary.
Note: NIRS, near-infrared spectroscopy; NIR, near-infrared; RGB, red–green–blue; GLCM, gray-level co-occurrence matrix; PCA, principal component analysis; F1, harmonic mean of precision and recall at a matched evaluation level. Top-2 error means the correct class is absent from the two highest-ranked predictions.
Table 13. Qualitative comparison of sensing and validation pathways for chili sorting under heterogeneous experimental conditions.
Table 13. Qualitative comparison of sensing and validation pathways for chili sorting under heterogeneous experimental conditions.
Sensing RouteComparative EvidenceDeployment AdvantageDecision Boundary for Chili Sorting
[E4] Full-spectrum HSI discovery [179,180]Rich spatial-spectral information supports maturity, pigment and visible-defect analysis; 601–950 nm is common.Suited to band discovery and failure analysis.Data and calibration burdens limit line deployment.
Compact multispectral imaging: [180] (E4); [184] (E3)Selected bands can retain task-specific discrimination with fewer variables.Reduces optical, data and inference complexity while retaining spatial localization.Compare with full HSI on the same samples and held-out domains.
[E4] Point Vis–NIR sensing [181]Commercial systems integrate optical geometry, calibration and multi-lane line-speed operation.Strong operational maturity and maintainability evidence.Weak spatial localization may miss local defects; geometry must match chili size and pose.
[E4] AI and data-standardization support [180,185]Shared metadata, domain-aware splits and real-time validation recur across routes.Supports reproducible comparison, calibration transfer and controlled updating.Deployment claims require external-domain tests, hardware latency and system metadata.
Proposed staged hybrid pathway; synthesis of [179,180,181,185] (E4) and [184] (E3)Laboratory information richness and commercial maturity come from different routes.Proposes linking HSI discovery, multispectral localization and selective point sensing.The combined pathway has not been validated as a complete system; select modules by tested marginal benefit.
Note: HSI, hyperspectral imaging; Vis–NIR, visible–near-infrared; AI, artificial intelligence. E3 and E4 refer to individual sources under Table 4; the staged pathway is a proposed synthesis, not a separately validated system.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Cao, J.; Wu, Y.; Zhang, L.; Zhang, Y.; Tang, Z. Research Progress on Intelligent Color-Sorting Equipment for Post-Harvest Chili Peppers: Machine Vision, Pneumatic Actuation, and System Integration. Processes 2026, 14, 2991. https://doi.org/10.3390/pr14182991

AMA Style

Cao J, Wu Y, Zhang L, Zhang Y, Tang Z. Research Progress on Intelligent Color-Sorting Equipment for Post-Harvest Chili Peppers: Machine Vision, Pneumatic Actuation, and System Integration. Processes. 2026; 14(18):2991. https://doi.org/10.3390/pr14182991

Chicago/Turabian Style

Cao, Junhao, Yapeng Wu, Liming Zhang, Yu Zhang, and Zhong Tang. 2026. "Research Progress on Intelligent Color-Sorting Equipment for Post-Harvest Chili Peppers: Machine Vision, Pneumatic Actuation, and System Integration" Processes 14, no. 18: 2991. https://doi.org/10.3390/pr14182991

APA Style

Cao, J., Wu, Y., Zhang, L., Zhang, Y., & Tang, Z. (2026). Research Progress on Intelligent Color-Sorting Equipment for Post-Harvest Chili Peppers: Machine Vision, Pneumatic Actuation, and System Integration. Processes, 14(18), 2991. https://doi.org/10.3390/pr14182991

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop