1. Introduction
In large-scale grain storage and purchasing, throughput, inspection costs, and labor availability constrain the number of kernels that can be inspected in detail. Low-cost, non-destructive screening based on machine vision and basic physical measurements can process large numbers of samples rapidly, whereas more intensive re-inspection is limited by the associated time, cost, and labor requirements [
1,
2,
3,
4]. Therefore, a practical inspection system should not only predict quality labels, but also determine which kernels should be inspected first. Under a fixed re-inspection budget, the operational objective is to rank abnormal or high-priority kernels near the top of the screening list so that limited verification capacity is allocated efficiently.
Recent advances in machine vision, deep learning, and non-destructive sensing have substantially improved automated cereal-quality assessment. Convolutional neural networks [
5], visual Transformers [
6], and non-destructive imaging methods [
7,
8] have shown strong performance in detecting kernel damage, discoloration, pest attack, and other visible defects. Beyond RGB appearance, complementary sensing channels such as hyperspectral or near-infrared measurements and lightweight physical attributes can provide additional information on composition and structural condition, supporting more robust quality assessment across heterogeneous samples [
9,
10].
Most existing approaches still optimize classification accuracy or defect-identification performance and implicitly treat samples as equally important, without considering the practical constraint of limited re-inspection resources [
11,
12]. Under a fixed budget, however, screening utility depends on whether high-priority kernels are ranked early enough to enter the restricted re-inspection subset [
13,
14].
Also, the complexity of risk priority judgment is increased by the multimodal nature of low-cost screening cues. The visual appearance principally indicates phenomena like mold growth, surface discoloration, and pest activity, whereas lightweight physical characteristics might reveal the loss of integrity, degrees of deformation, and associations with impurities which cannot be readily observed with the naked eye. Despite the popularity of multimodal modeling in the field of agricultural perception, the current literature mostly assumes that all modalities are equally reliable and tries to diminish the differences between modalities during the fusion process [
15,
16]. However, in practice, intermodal consistency highly depends on specific sample conditions. An example is that the appearance of some grains with normal appearance may show abnormal weight or size variations simultaneously because of shooting angles or lighting. These are not data collection noise but can often correspond to potential quality risks [
17]. When the sample-level modality reliability is not constant and cross-modal conflict structures cannot be explicitly modeled, the ability of the model to remain stable and to discriminate will be significantly affected in the risk ranking task.
Alongside this, the two-stage pipeline modeling paradigm is also widely used in current research on quality management of cereals such as risk stratification: firstly, quality groups or pseudo-labels are constructed using unsupervised clustering or heuristic rules, and then independent prediction models are fitted to the grouping outcomes [
18,
19,
20]. Even though this strategy is convenient for establishing hierarchical division, its operation is extremely reliant on the stability of the initial grouping, and the representation learning procedure is separated from the downstream screening decision, readily introducing the issue of error accumulation and structural mismatch [
21,
22]; thus, it is not simple to establish a stable and budget-aware multi-level risk ranking framework.
To respond to these difficulties, this paper introduces a multimodal end-to-end risk screening framework, RiMS-FiLM, that aims to prioritize the discovery of high-risk samples under a fixed re-inspection budget instead of merely pursuing traditional classification accuracy. This framework forms the multimodal fusion process upon the Feature-wise Linear Modulation (FiLM) mechanism that adaptively modulates the visual features using the physical properties at the sample level and integrates conflict-aware modeling strategies to explicitly depict cross-modal inconsistencies, and creates a representation space that is more suitable for the needs of risk screening tasks. It is based on this that a prototype-guided risk structure learning mechanism is presented to arrange multimodal embeddings into relatively compact and separable low-, medium-, and high-priority regions, to allocate multimodal embeddings by stable priority ranking without post hoc clustering. The model directly outputs continuous risk scores to serve real-world screening decisions through a unified end-to-end optimization framework. Extensive experimental results using the GrainSet-Maize multimodal maize kernel dataset reveal that RiMS-FiLM improves the high-risk discovery rate under limited re-inspection capacity constraints, outperforming traditional two-stage hierarchical methods and representative selective screening paradigms. Accordingly, the methodological contribution of RiMS-FiLM lies in adapting and integrating these established components into an end-to-end fixed-budget re-inspection prioritization framework for maize kernels.
The primary contributions of this work can be outlined as follows:
(1) We define the quality inspection of maize kernels as a risk screening problem under a fixed re-inspection budget constraint, shifting the evaluation focus from global classification accuracy to decision-oriented screening effectiveness under real re-inspection capacity constraints, making the model optimization objective consistent with the real quality management needs.
(2) We present a FiLM-based multimodal fusion mechanism and introduce a conflict-aware modulation strategy to capture the sample-dependent modality reliability variations and also the cross-modal discrepancies between visual features and physical attributes, hence learning risk-aware representations which are more suitable for the risk priority ranking task.
(3) We design an end-to-end joint optimization framework that jointly learns multimodal embed-ding representations, prototype-guided structured risk representations, and continuous risk scores to make screening decisions, achieving stable and interpretable priority ranking without staged clustering.
3. Results and Discussion
3.1. Experimental Setup and Baseline Methods
3.1.1. Implementation Details
All experiments are implemented using the PyTorch (Python 3.12) framework. The visual encoder adopts the Swin-T architecture (swin_tiny_patch4_window7_224) pretrained on ImageNet and fine-tuned on the GrainSet-Maize dataset. The physical encoder, fusion modules, prototype parameters, and prediction head are trained from scratch.
Models are optimized using the AdamW optimizer with an initial learning rate of 2 × 10−4 and a weight decay of 1 × 10−2. Training is conducted for up to 30 epochs with a batch size of 32. Early stopping is applied based on validation Macro-F1 with a patience of 7 epochs to prevent overfitting. Automatic mixed-precision training is enabled to improve computational efficiency. The principal comparisons and ablation experiments were repeated with five fixed random seeds (42, 123, 2024, 2025, and 2026); quantitative tables report mean ± standard deviation. For the CCL-SC baseline, five cross-entropy-only warm-up epochs preceded the contrastive stage, and the same validation Macro-F1/patience principle was applied after contrastive training began.
The hidden dimensions of the physical encoder and fusion layers are both set to 256, and cross-attention fusion employs eight attention heads. The prototype structure regularization weight is fixed at , while conflict-aware modulation is activated with a scaling factor of 1.0.
All experiments are performed on an NVIDIA RTX 3090 GPU (NVIDIA Corporation, Santa Clara, CA, USA). Comparable backbone, optimizer, batch-size, and evaluation settings are used across methods where applicable, while method-specific training components required by individual baselines are retained.
3.1.2. Baseline Screening Paradigms
To comprehensively evaluate the proposed RiMS-FiLM framework under fixed-budget screening scenarios, four representative screening paradigms are selected as comparative baselines.
(1) Two-stage KMeans screening (TSK)
This method applies unsupervised KMeans clustering on multimodal embeddings to partition maize samples into several latent risk groups [
19,
20]. High-risk clusters are subsequently prioritized for reinspection according to cluster-level risk statistics.
(2) Selective classification network (SCN)
SelectiveNet-style models jointly learn prediction and confidence estimation, allowing uncertain samples to be rejected [
28]. In this study, confidence scores are adapted as screening priorities for fixed-budget reinspection.
(3) Training-dynamics screening (TDS)
This approach leverages prediction stability and learning difficulty across training epochs as indicators of risk, ranking samples with unstable training dynamics as high-priority candidates [
29].
(4) Confidence-aware contrastive screening (CCS)
We implement the confidence-aware contrastive learning for selective classification (CCL-SC) formulation [
30] and adapt its output to fixed-budget screening; this adapted baseline is denoted CCS in the tables and figures for brevity.
All baselines are implemented under the same multimodal encoder backbone where applicable and converted into ranking-based screening strategies for fair comparison under identical budget constraints.
3.1.3. Complementary Classification Performance Metrics
In addition to the budget-aware screening metrics defined in
Section 2.2.3, conventional classification metrics are reported to provide complementary evaluation of risk prediction performance.
Specifically, Macro-F1 and Micro-F1 scores are computed across the three risk levels. Macro-F1 equally weights each risk category and is particularly suitable for assessing robustness under class imbalance, while Micro-F1 reflects overall prediction accuracy dominated by majority classes. These metrics serve as auxiliary indicators and do not directly represent operational screening effectiveness, which is primarily assessed through , , and .
Specifically, Macro-F1 and Micro-F1 scores are computed across the three priority levels. Macro-F1 equally weights each priority category and is particularly suitable for assessing robustness under class imbalance, while Micro-F1 reflects overall predictive accuracy. These metrics serve as complementary indicators and do not directly represent operational screening effectiveness, which is primarily assessed through , , and .
3.2. Fixed-Budget Screening Performance Comparison
3.2.1. High-Risk Prioritization Under Fixed Budgets
We first compare RiMS-FiLM against four representative screening paradigms under the fixed-budget setting defined in
Section 2.2.
Table 4 summarizes high-priority recall and missed-high rates at representative inspection budgets B.
Across five random seeds, RiMS-FiLM retrieves 23.75% of high-priority kernels at a 5% budget and 47.50% at a 10% budget, reaching the theoretical maximum at both operating points.
At a 20% re-inspection budget, RiMS-FiLM achieves a mean high-priority recall of 0.9465 ± 0.0021, compared with 0.9420 ± 0.0023 for CCS, and reaches 0.9995 ± 0.0007 at a 30% budget.
These results indicate that RiMS-FiLM is able to front-load high-priority kernels within limited inspection budgets, a property relevant to real-world grain storage quality control.
3.2.2. Risk–Coverage Behavior and Screening Efficiency
Figure 4 summarizes the five-seed mean missed-high rate at representative inspection budgets. RiMS-FiLM remains near the lower envelope of missed-high rates over the operational range, consistent with its lowest mean
among the screening baselines.
Figure 5 reports the corresponding five-seed mean high-priority recall at the same representative budgets, with error bars showing the standard deviation across seeds. RiMS-FiLM reaches the theoretical ceiling at 5% and 10% budgets and remains close to the attainable maximum at 20% and 30%.
(missed-high) represents cumulative missed detection over the opera-tional budget range
; a lower value indicates faster high-priority sample discovery. As presented in
Table 5, RiMS-FiLM achieved the lowest mean
among the compared screening baselines (0.1056 ± 0.0001), compared with 0.1059 ± 0.0002 for CCS, and substantially lower values than SCN and TDS. Under the same test-set prevalence, the theoretical continuous lower bound is approximately 0.1053, so the observed RiMS-FiLM value is close to the attainable lower bound over this budget range.
3.2.3. Joint Analysis with Classification Robustness
Beyond operational screening efficiency,
Table 5 also reports Macro-F1 and Micro-F1 scores to evaluate the robustness of risk prediction under class imbalance.
RiMS-FiLM achieves the highest Macro-F1 among all methods, indicating balanced recognition capability across low-, medium-, and high-risk levels. This improvement is particularly important in grain storage scenarios where high-risk samples constitute a minority yet carry the greatest operational importance. Meanwhile, Micro-F1 remains competitive, reflecting strong overall predictive accuracy.
The combination of the lowest mean among the screening baselines and the highest Macro-F1 indicates that RiMS-FiLM provides both efficient early retrieval and balanced three-class prediction. These quantitative results are consistent with the intended representation design, while the diagnostic visualizations below are interpreted descriptively rather than as stand-alone causal proof of the internal mechanism.
As an additional grouped robustness evaluation, RiMS-FiLM retained its performance when identical (DU_grain, weight) groups were prevented from crossing partitions: over five seeds, Macro-F1 was 0.9940 ± 0.0014 and Recall@20% was 0.9485 ± 0.0010. Under the same grouped protocol, CCS achieved 0.9850 ± 0.0025 and 0.9428 ± 0.0030, respectively, while the vision-only ablation achieved 0.9849 ± 0.0023 and 0.9433 ± 0.0029. This result indicates that the main finding is retained under the more conservative group-disjoint split.
3.2.4. Discussion: Implications for Operational Grain Risk Management
The findings demonstrate that, under a fixed re-inspection budget, screening effectiveness depends more on how quickly the model ranks high-risk samples near the top than on overall classification accuracy. In addition, there may be substantial differences in the actual value of the different types of methods when used for actual screenings. This is particularly evident within the lower budget range; therefore, even small differences in rank position can produce significant missed opportunities for high-risk kernels.
Similarly, predictive assessment tools based on either confidence-based or dynamic training do not actually characterize risk by prediction uncertainty. While these methods may have some advantages in an overall abstention-based behavior model, the ranking structures of these methods are typically not stable, especially when budgets are limited when the highest-risk samples will be primarily distributed throughout the midrange and low-end of the sample rankings while limiting the effectiveness of earlier screening methods. Additionally, representative-space detection strategies will provide smoother risk coverage curves; however, misalignment of screening priorities can still occur, at the borderline between medium-and high-risk samples.
RiMS-FiLM does not utilize post hoc adjustments to the confidence in the ranking like these methods but rather shapes the ranking space through structured representation learning. To avoid relying too heavily on one modality, the combination of conflicting multimodal feature fusion mechanisms allows the model to modify feature responses based on the inconsistency of both the visual and physical domains. Additionally, the prototype-based risk structure constraints increase the separation of different risk types within the embedding space thereby allowing higher risk instances to be located more stably in front of the ranking system. The mechanism of creating structural representations at the representation level and then using that structural representation to reflect a ranking at the decision level provides for consistent behavior in the screening process regardless of the budget.
With respect to being able to manage storage from a practical standpoint there are very direct implications for operation. When storing or moving large amounts of grain, the limiting factor will almost always be the resources used to re-inspect product. Thus, the key area for making decisions will not be ‘Is my decision correct?’, but rather ‘Which samples have the highest risk and should be inspected sooner than later?’. So even if a screening type of model can just raise the rank of those high-risk samples reliably, there may be great economic savings and a significant reduction in risk as it relates to quality due to practical applications.
Further studies should evolve from a “prediction-based” model development paradigm to a “decision-based” model development paradigm, as discussed above. Using classification confidence or error minimizer objectives in isolation does not demonstrate the actual utility of the models when resources are limited. Therefore, developing the risk structure in the representation and aligning a rank objective with the risk would appear to be the best way to maximize utility in these types of problems. This study’s performance of RiMS-FiLM accordingly supports the above argument through evidential validation of its potential.
3.3. Mechanistic Comparison with CCS
Because CCS is the strongest screening baseline in
Table 4 and
Table 5, we further compare it with RiMS-FiLM using the final five-seed results. This comparison focuses on reproducible classification and budget-aware screening metrics rather than inferring internal mechanisms from a single confusion matrix or discrete score plot.
3.3.1. Five-Seed Classification Comparison
Figure 6 compares Macro-F1 and Micro-F1 across five seeds. RiMS-FiLM achieves 0.9917 ± 0.0027 Macro-F1 and 0.9927 ± 0.0024 Micro-F1, compared with 0.9858 ± 0.0018 and 0.9877 ± 0.0015 for CCS, respectively.
The advantage is therefore observed in both balanced class-wise performance and overall predictive accuracy, while the standard deviations remain small for both methods.
These five-seed results provide the quantitative basis for comparing the two methods; they replace reliance on the earlier single-run confusion-matrix interpretation.
3.3.2. Fixed-Budget Ranking Comparison
Figure 7 compares high-priority recall at the four reported re-inspection budgets. Both methods reach the theoretical ceilings at 5% and 10%. At a 20% budget, RiMS-FiLM achieves 0.9465 ± 0.0021 compared with 0.9420 ± 0.0023 for CCS; at 30%, the two methods are both close to complete high-priority coverage.
Together with the
values in
Table 5 (0.1056 ± 0.0001 for RiMS-FiLM and 0.1059 ± 0.0002 for CCS), the results show a small but consistent advantage for RiMS-FiLM in early high-priority retrieval under the primary stratified protocol.
3.4. Effectiveness of Multimodal Fusion Strategies
To analyze the impact of different multimodal interaction mechanisms on risk representation and screening stability, this paper compares the FiLM fusion method with three typical strategies: feature concatenation, gated weighted fusion, and attention-based cross-modal interaction. Under the premise of maintaining consistent encoder structure and training settings, the above methods are all embedded into the unified risk screening framework for evaluation.
In addition to the five-seed metrics,
Figure 8 provides a single-run diagnostic view of boundary-sensitive classification errors for the fusion variants. This visualization is descriptive only; quantitative conclusions are based on the five-seed summaries in
Table 6. The integrated screening metric is
as defined in
Section 2.2.3.
Figure 8 presents the performance of different fusion strategies on three critical error types from a screening decision perspective, including high-risk samples misclassified as low risk, classification deviation of medium-risk samples, and non-high-risk samples misclassified as high risk.
On the most critical severe miss metric, the differences are most pronounced in this diagnostic run. FiLM controls this error rate at 0.005, while Concat, Gated-Sum, and Cross-Attention reach 0.028, 0.010, and 0.014, respectively. Notably, Concat’s error rate exceeds FiLM’s by more than five times, indicating that in the absence of effective cross-modal modulation, high-risk features are more likely to be attenuated during fusion, thereby being incorrectly suppressed to the lowest priority region.
Based on the medium-risk error metric, FiLM has the lowest risk score at 0.001. The other methods were at risk scores of 0.015 for Concat, 0.006 for Gated-Sum, and 0.007 for Cross-Attention. The effect of this error type is limited on overall accuracy, but it could create an issue with continuity in the ranking structure based on the cumulative effect in terms of risk stratification.
In this single-run diagnostic, FiLM shows lower boundary-sensitive error rates than the alternative fusion variants. These patterns provide qualitative context for
Table 6 but are not used to claim statistical dominance on every screening metric.
The five-seed results in
Table 6 provide the primary evidence: RiMS-FiLM attains the highest mean Macro-F1 and Micro-F1 among the fusion variants while maintaining near-ceiling fixed-budget screening performance.
Table 6 compares different multimodal fusion strategies from two perspectives: overall screening performance and classification stability. Overall, the differences between methods in
are relatively close, indicating that at the global risk-coverage curve level, different fusion methods are all able to form relatively similar ranking trends.
At the 20% re-inspection budget, all strong fusion variants remain close to the theoretical ceiling, with mean Recall@20% values ranging from 0.9455 to 0.9475; RiMS-FiLM achieves 0.9465 ± 0.0021.
RiMS-FiLM achieves the highest mean Macro-F1 (0.9917 ± 0.0027) and Micro-F1 (0.9927 ± 0.0024) among the fusion variants, while maintaining near-ceiling fixed-budget screening performance.
Accordingly, the boundary-error visualization should be read as a descriptive complement to the five-seed results. The principal quantitative conclusion is that FiLM provides the strongest mean classification performance among the tested fusion strategies without sacrificing screening efficiency near the operational ceiling.
3.5. Ablation Study and Structural Interpretation of RiMS-FiLM
To further analyze the specific roles of each core module of RiMS-FiLM in risk screening, this paper conducts ablation experiments by progressively removing key components, including physical feature inputs, prototype-guided risk structure, and conflict-aware modulation mechanism. The corresponding quantitative results are shown in
Table 7, while
Figure 9 and
Figure 10 provide single-run diagnostic visualizations of the changes in representation structure and screening behavior under different component removal conditions. Under the premise of maintaining consistent training settings and the remaining model structure, the above ablation experiments aim to evaluate the contribution of each module to risk ranking stability and screening effectiveness from two levels: overall performance and structural performance. Compared with analysis methods that rely solely on metric changes, combining the comparison of embedding space distribution and risk ranking behavior can more intuitively reveal the impact of different components on the model’s internal risk organization mechanism.
3.5.1. Impact of Prototype-Guided Risk Structure
As shown in
Table 7, removing the prototype-guided structure produces a small decrease in mean Macro-F1, from 0.9917 ± 0.0027 to 0.9908 ± 0.0030, while the fixed-budget ranking metrics remain close to saturation. The five-seed results therefore support the prototype term primarily as a representation-structure regularizer rather than as a component that must improve every individual screening metric.
Figure 9 provides a descriptive view of how the prototype regularizer changes the geometry of the learned representation. In the with-prototype condition, the Top-B samples show a more concentrated distance pattern relative to the class reference centers, and the margin distribution shifts toward larger separation values compared with the no-prototype condition. This is consistent with the explicit margin constraint in Equation (16).
Figure 9c shows that the margin distribution for the selected Top-B samples is generally shifted toward larger values when prototype regularization is used.
Figure 9d further illustrates the margin behavior of missed high-priority samples. These plots are descriptive diagnostics of representation geometry and are not interpreted as independent causal proof of screening performance.
Combining
Table 7 and
Figure 9, it can be seen that the role of the prototype-guided structure in RiMS-FiLM is mainly reflected in two aspects: on one hand, it provides explicit risk-level reference for multimodal representations, enabling samples to form a more ordered organization in the embedding space; on the other hand, it enlarges the geometric margin near the screening boundary, thereby improving the stability of high-risk samples being preferentially selected under limited budget conditions. Therefore, this module provides an explicit structural regularizer for the multimodal representation. These representation-space visualizations are descriptive and are not used as causal proof of the mechanism.
3.5.2. Effect of Multimodal Physical Feature Integration
We further analyze the role of physical features in multimodal fusion. Removing the physical branch produces the clearest ablation effect:
Table 7 shows that Macro-F1 decreases from 0.9917 ± 0.0027 to 0.9859 ± 0.0014, Recall@20% decreases from 0.9465 ± 0.0021 to 0.9420 ± 0.0026, and
increases from 0.1056 ± 0.0001 to 0.1060 ± 0.0001. This indicates that physical attributes provide complementary information beyond visual appearance.
This transition illustrates that using only the visual modality cannot provide all the characteristics necessary for assessing the operational re-inspection priority of maize kernels; some high-risk samples may be visually indistinguishable, but have previously identified weight/size anomalies. Information about the sample will be unusable or inadequate without a physical representation, which leads to bias in the model’s determination of sample risk.
With regard to the screening process, this bias manifests as an overall backward shift in the ranking positions of high-priority abnormal samples. Due to the lack of supporting visual evidence, the model is more likely to classify a sample as “visually normal” and place it in the low-risk area even if the sample warrants a higher re-inspection priority. This has important implications for re-inspection under limited budgets, where ranking position directly impacts whether a sample will be included in the re-inspection set.
In general, physical features complement visual information and help stabilize risk judgment when integrating different modalities. Removing physical features from the model reduces the model’s ability to identify potentially high-priority abnormal samples due to lack of information, ultimately resulting in lower overall screening effectiveness.
3.5.3. Contribution of Conflict-Aware Modulation
Finally, we analyze the role of the conflict-aware modulation mechanism. Across five seeds, removing this mechanism changes the mean Macro-F1 from 0.9917 ± 0.0027 to 0.9910 ± 0.0027, Recall@20% from 0.9465 ± 0.0021 to 0.9463 ± 0.0027, and
from 0.1056 ± 0.0001 to 0.1057 ± 0.0001. The numerical effect is modest;
Figure 10 is therefore used to describe how the conflict score behaves rather than to claim a large independent performance gain.
Figure 10 shows the conflict-score distributions for selected and missed high-priority samples in a single diagnostic run, with and without the conflict-aware branch. The distributions differ across these groups, indicating that the explicit discrepancy signal participates in the model’s decision pathway. However, the conflict score is not monotonic with prediction error and should not be interpreted as a calibrated uncertainty measure.
The comparison without the conflict branch further shows that the distributional pattern changes when the explicit discrepancy-based logit adjustment is removed. Because this visualization is single-run and descriptive, it is used only to illustrate how the conflict signal behaves, while the five-seed ablation values in
Table 7 provide the quantitative evidence for the component’s contribution.
Accordingly, the conflict-aware branch is best interpreted as an additional cross-modal discrepancy adjustment rather than a stand-alone uncertainty detector. Its mean quantitative effect is modest, but the full model retains slightly higher Macro-F1 and comparable near-ceiling screening performance across the five seeds.
4. Conclusions
Under a fixed re-inspection budget, operational grain screening is not only a classification task but also a prioritization problem. RiMS-FiLM combines non-destructive RGB appearance with lightweight physical metadata to rank maize kernels for follow-up verification. Across five random seeds on the primary stratified split, the model achieved Recall@20% = 0.9465 ± 0.0021 and Macro-F1 = 0.9917 ± 0.0027. Because the risk tiers are operational labels derived from defect categories rather than laboratory-validated toxicological endpoints, these results should be interpreted as re-inspection-priority performance for quality control.
Ablation and diagnostic analyses are consistent with complementary roles for the three main components: prototype-guided learning provides class-structured regularization, physical metadata supplement visual cues, and FiLM-based conditional modulation with the discrepancy branch provides sample-dependent cross-modal adjustment. Together, these components support prioritization near the re-inspection cutoff, where ranking errors have the greatest operational consequence.
Overall, the study demonstrates the potential of low-cost multimodal sensing and machine learning for non-destructive grain quality screening under resource constraints. The present study is limited to one public dataset and does not include independent external validation, real warehouse testing, laboratory-confirmed endpoints, or a dedicated sensor-variability analysis. Future work should validate the framework against laboratory-measured physicochemical or safety endpoints, evaluate external batches and distribution shifts, and investigate dynamic inspection budgets, additional sensing modalities, and uncertainty-calibrated decision rules.