Next Article in Journal
Contribution of Protein, Starch, and Fiber Composition to the Prediction of Dough Rheology and Baking Quality in U.S. Hard Red Spring Wheat
Previous Article in Journal
Effect of Combined Pretreatments on Yield and Quality of Cold-Pressed Pomegranate Seed Oil
 
 
Article
Peer-Review Record

Development and Interpretability Analysis of Near-Infrared Spectroscopy Models for Fat and Protein Prediction in Foxtail Millet [Setaria italica (L.) Beauv.]

by Anqi Gao 1,2,3,4, Erhu Guo 1,2,3, Bin Wang 2, Dongxu Zhang 2, Kai Cheng 2, Xiaofu Wang 5, Aiying Zhang 1,2,3,* and Guoliang Wang 1,2,3,*
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3: Anonymous
Reviewer 4: Anonymous
Submission received: 10 January 2026 / Revised: 5 February 2026 / Accepted: 9 February 2026 / Published: 11 February 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

Dear authors,

The present manuscript presents a near-infrared spectroscopy–based study for predicting fat and protein contents in foxtail millet using a combination of spectral preprocessing, feature selection, and machine-learning models. While the individual methods employed (SSA, PLS, RF, SHAP) are well established in chemometrics and spectral analysis, the study offers a meaningful contribution through their systematic integration.

In particular, the use of repeated SSA runs to identify stable key wavelengths based on selection frequency, together with SHAP-based interpretation of wavelength contributions, enhances both the robustness and interpretability of the proposed models.

The work does not introduce novel algorithms; however, it provides a methodologically sound and application-oriented framework for component-specific modeling of millet quality, with practical implications for non-destructive testing and portable sensor development. Therefore, the novelty of the study needs to be highlighted and better explained.

The title is generally informative but should be revised to improve clarity and avoid overstating methodological innovation, particularly regarding the use of the term “innovation.” The title would benefit from simplification and clearer emphasis on model development and interpretability rather than innovation. Revising the title to improve readability and accuracy is recommended.

The sample size of 214 is adequate for NIR chemometric modeling and allows robust calibration and validation. However, the authors should provide more information on the range and distribution of fat and protein contents and clarify whether the sample set adequately represents typical variability in production conditions, and in relation with already available relevant studies.

Regression model is not clearly described and presented, as it is important for such study.

The strengths and limitations section is good structured but would be good to benefit from more wording regarding methodological novelty and generalizability, to better align strengths with admitted limitations. Besides, the manuscript tends to overstate methodological innovation in parts of this section. The Sparrow Search Algorithm and SHAP are established techniques, and the authors should clarify that the contribution lies in their integration and stability-based application rather than in the novelty of the algorithms themselves. Additionally, claims regarding improved generalizability should be moderated, given the acknowledged homogeneity of the sample set.

Several sentences are understandable but not written in natural academic English, and it could be improved for clarity and professionalism. (e.g. ''This work took "Changnong No. 47" foxtail millet as the research object, with 214 samples collected.'')

 

Comments on the Quality of English Language

Should be revised for standard academic English.

Author Response

Dear reviewer

Thank you very much for reviewing our manuscript. We would also like to express our gratitude to the reviewers for their efforts in helping to improve our manuscript titled “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (ID: foods-4118117). The comments were all valuable and very helpful for revising and improving our paper. Our research team have studied each reviewers’ comments point by point and have made necessary modifications and supplements, which we hope will meet with your approval. Revised portion are mark by “Track Change” in bold font throughout the revised manuscript with track change.

 

The title is generally informative but should be revised to improve clarity and avoid overstating methodological innovation, particularly regarding the use of the term “innovation.” The title would benefit from simplification and clearer emphasis on model development and interpretability rather than innovation. Revising the title to improve readability and accuracy is recommended.

Thank you for your constructive feedback regarding the title of our manuscript. We agree that the original title could be simplified to more accurately reflect the core methodological contributions of the work, without overstating innovation. As recommended, we have revised the title to improve clarity, readability, and accuracy. The new title now explicitly emphasizes the two key aspects of our study: 1) model development and 2) interpretability analysis. We have removed the term “innovation” to ensure a more precise and modest description of our work.

The revised title is: Development and Interpretability Analysis of Near-Infrared Spectroscopy Models for Predicting Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.]

We believe this title more clearly and accurately conveys the paper's content, focusing squarely on the construction and explanation of the predictive models. Thank you again for this valuable suggestion.

The sample size of 214 is adequate for NIR chemometric modeling and allows robust calibration and validation. However, the authors should provide more information on the range and distribution of fat and protein contents and clarify whether the sample set adequately represents typical variability in production conditions, and in relation with already available relevant studies.

Thank you for your constructive feedback and for acknowledging the adequacy of our sample size for NIR modeling. We appreciate your suggestion to provide more detailed information on the composition ranges and the representativeness of our sample set.

We have performed detailed descriptive and inferential statistics on the reference data. To ensure robustness, we employed the interquartile range (IQR) method for outlier detection.

Fat Content: No outliers were detected among the 214 samples. The content ranged from 3.30% to 4.21%, with a mean of 3.74% and a standard deviation (SD) of 0.35%. The distribution was approximately normal, symmetric, and concentrated, as detailed in the  Figure 1(b).

Protein Content: After removing 22 outliers (10.28%), 192 samples were retained for analysis. The protein content showed a mean of 10.43% with a notably low SD of 0.39%, indicating high sample-to-sample stability. Its distribution showed a moderate negative skew (Figure 1(c)).
The significant difference in the coefficient of variation (CV) between fat (9.34%) and protein (3.72%) confirms that fat content exhibits greater natural variability among foxtail millet samples, which our modeling successfully captures.
For most foxtail millet cultivars, the fat content ranges from 3% to 4%, and the protein content from 9% to 12%. The observed statistical dispersion (particularly for fat) and the near-normal distributions demonstrate that the set encompasses the intrinsic biological variation found in production. The use of IQR-based outlier removal further ensures that the subsequent model calibration is based on a representative core population, avoiding distortion from anomalous specimens.
The compositional ranges established in our study align with the known nutritional profile of foxtail millet. Our reported fat range (3.30–4.21%) falls within the broader 1.7–5.0% range commonly cited in the literature for this crop. Similarly, our mean protein content (~10.43%) is consistent with typical values (often 10–14%) reported in prior nutritional studies. The fact that our statistically refined data lies centrally within these well-established global ranges strongly supports that our sample set is representative of the species' typical variability.

Thank you again for this valuable suggestion.

Regression model is not clearly described and presented, as it is important for such study.

Thank you for your valuable feedback. We agree that a clear description and thorough analysis of the regression models are crucial for this study. In response, we have substantially expanded the manuscript to provide a detailed account of the model development, evaluation, and comparative analysis for both fat and protein content.

All models demonstrated favorable performance on the prediction set based on the key wavelength set, as indicated by their R2 values. Among these, the RF model exhibited the best overall predictive performance (RP2 = 0.797, RMSEP = 0.218%, RPDP = 2.219), highlighting its advantage in handling nonlinear spectral relationships. The PLS and SVM models also delivered robust results, with prediction set R2 values of 0.742 and 0.737, and RPD values of 1.969 and 1.950, respectively. This indicates that the linear approach (PLS) and the kernel-based nonlinear method (SVM) remain effective, though slightly less accurate than RF for this specific component. A detailed comparison of the evaluation metrics for each model is presented in Table 1. In summary, the wavelength selection conducted via the SSA, combined with multi-strategy candidate set optimization, enables the stable and effective identification of key spectral features associated with fat content in foxtail millet.

Among these, the PLS model demonstrated the best overall predictive performance (RP2 = 0.695, RMSEP = 0.268%, RPDP = 1.811). In contrast, the nonlinear models, RF and SVM, exhibited significantly lower predictive accuracy, with RP2 values of only 0.472 and 0.538, respectively. A detailed comparison of the evaluation metrics for all models is provided in Table 1

We discussed the model in the discussion section. 

The strengths and limitations section is good structured but would be good to benefit from more wording regarding methodological novelty and generalizability, to better align strengths with admitted limitations. Besides, the manuscript tends to overstate methodological innovation in parts of this section. The Sparrow Search Algorithm and SHAP are established techniques, and the authors should clarify that the contribution lies in their integration and stability-based application rather than in the novelty of the algorithms themselves. Additionally, claims regarding improved generalizability should be moderated, given the acknowledged homogeneity of the sample set.

Thank you for your constructive and insightful feedback regarding the Strengths and Limitations section. We sincerely appreciate your positive note on its structure and your specific suggestions for improving the wording on methodological novelty and generalizability. We fully agree that the contribution of our work lies more in the integrated application and stability-focused framework of established techniques rather than in the inherent novelty of the algorithms themselves, and that claims of generalizability should be carefully moderated in light of our sample characteristics.

As you recommended, we have thoroughly revised the section to better align the stated strengths with the admitted limitations, and to provide a more accurate and moderated description of our methodological contribution. The primary strengths of this study lie in its systematic methodological integration and interpretability-driven approach. Firstly, it integrates the SSA with a frequency statistics strategy for robust feature wavelength selection in near-infrared spectroscopy. This combination and its application within a stability-analysis framework enhance the global optimization capability of the selection process and improve the robustness of the results, achieving effective data dimensionality reduction with clear physicochemical grounding. Secondly, beyond constructing quantitative prediction models, the study incorporates SHAP values to provide post-hoc interpretability of the model decisions, thereby improving the transparency of the machine learning pipeline for this application. Thirdly, by conducting a detailed chemical bond attribution analysis for the key wavelengths identified by SSA, the study directly links spectral features to the molecular structures of fat and protein, which strengthens the mechanistic foundation and supports the broader relevance of the findings.

This study also has several limitations, which help appropriately contextualize its findings. The primary limitation lies in the genetic and geographical homogeneity of the sample set. All millet samples were sourced from a single ecological region and consisted of only one cultivar. While this helped control for extraneous variables, it limits the generalizability of the calibration model to a wider range of cultivars, growing environments, and potential genotype-by-environment interactions. Secondly, although the sample size was sufficient for preliminary modeling, it may not fully capture the extreme compositional variations encountered in actual production, which could affect the model's robustness in broader applications. Thirdly, while the SSA identified key wavelengths, the inherent spectral overlap and collinearity between fat and protein information in the NIR region may impose certain constraints on the model's specificity and accuracy. Fourthly, regarding model validation strategy, this study primarily relied on a single random data split for final performance evaluation, which may lead to an insufficient assessment of model robustness. Fifthly, this study has not yet explicitly defined the Applicability Domain for the developed predictive models. Additionally, the research did not delve into the effects of variations in sample physical states (such as particle size and moisture content) on spectra and model performance. Therefore, future studies should employ more diverse and larger sample sets, coupled with more rigorous validation strategies and uncertainty analysis, to develop and validate non-destructive detection models with greater generalizability and metrological reliability.

We believe these revisions successfully address your concerns by providing a more precise, balanced, and academically rigorous discussion of the study's value and its boundaries. Thank you again for your valuable guidance, which has undoubtedly improved the quality and clarity of our manuscript.

Several sentences are understandable but not written in natural academic English, and it could be improved for clarity and professionalism. (e.g. ''This work took "Changnong No. 47" foxtail millet as the research object, with 214 samples collected.'')

We appreciate your constructive suggestions. The full text has undergone meticulous revision in response.

 

Reviewer 2 Report

Comments and Suggestions for Authors

The authors present a study on the prediction of protein and fat contents in foxtail millet using NIR hyperspectral imaging coupled with different modeling techniques. The manuscript is well written, and the authors appropriately discuss the limitations of their study. Nevertheless, several issues should be addressed to strengthen the manuscript and improve its scientific rigor.

  • Lines 222–223: The authors state that a nitrogen-to-protein conversion factor of 6.25 was used for foxtail millet. However, this factor represents a general average and is known to overestimate true protein content for many cereals. Previous studies (e.g., Mariotti et al., 2008) have demonstrated that foxtail millet has a crop-specific conversion factor, with reported values around 5.8 depending on nitrogen content, which is lower than 6.25. Given the unusually high nitrogen content relative to true protein in foxtail millet, the use of a universal factor (6.25) should be carefully justified, or alternatively, a species-specific conversion factor should be applied. The authors are encouraged to revise the calculation accordingly and/or discuss the potential impact of using 6.25 on protein estimation accuracy.

François Mariotti, Daniel D. Tomé, Philippe Patureau Mirand. Converting Nitrogen into Protein – Beyond 6.25 and Jones’ Factors. Critical Reviews in Food Science and Nutrition, 2008, 48 (2), pp.177-184. ff10.1080/10408390701279749ff. ffhal-02105858f

  • The manuscript does not clearly state whether intact or ground kernels were used for NIR hyperspectral analysis. This information is important, as sample preparation can substantially affect spectral variability, scattering effects, and model robustness. Please clarify this point.

 

  • The calculation and interpretation of RPD values are unclear. The authors should explicitly describe how RPD was computed and discuss whether the reported values (1.376 to 2.219) demonstrate adequate predictive performance. According to commonly accepted criteria, RPD values below ~2 generally indicate limited predictive capability. Please clarify whether the presented models can be reliably used for protein and fat prediction, and support this claim with appropriate references.

Author Response

Dear reviewer

Thank you very much for reviewing our manuscript. We would also like to express our gratitude to the reviewers for their efforts in helping to improve our manuscript titled “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (ID: foods-4118117). The comments were all valuable and very helpful for revising and improving our paper. Our research team have studied each reviewers’ comments point by point and have made necessary modifications and supplements, which we hope will meet with your approval. Revised portion are mark by “Track Change” in bold font throughout the revised manuscript with track change.

 

Lines 222–223: The authors state that a nitrogen-to-protein conversion factor of 6.25 was used for foxtail millet. However, this factor represents a general average and is known to overestimate true protein content for many cereals. Previous studies (e.g., Mariotti et al., 2008) have demonstrated that foxtail millet has a crop-specific conversion factor, with reported values around 5.8 depending on nitrogen content, which is lower than 6.25. Given the unusually high nitrogen content relative to true protein in foxtail millet, the use of a universal factor (6.25) should be carefully justified, or alternatively, a species-specific conversion factor should be applied. The authors are encouraged to revise the calculation accordingly and/or discuss the potential impact of using 6.25 on protein estimation accuracy.

Thank you for this insightful and highly relevant comment regarding the nitrogen-to-protein conversion factor. We sincerely appreciate your expertise on this matter and apologize for the oversight in our manuscript description. We have verified this with our testing service provider and confirm that this was an error in our writing. The protein content in the foxtail millet samples was determined in accordance with the latest Chinese National Standard GB 5009.5–2025, which stipulates a conversion factor of 5.83 for foxtail millet.

Thank you again for this critical correction, which significantly improves the accuracy and reliability of our methodology section.

The manuscript does not clearly state whether intact or ground kernels were used for NIR hyperspectral analysis. This information is important, as sample preparation can substantially affect spectral variability, scattering effects, and model robustness. Please clarify this point.

Thank you for raising this important point regarding sample preparation for NIR analysis. We agree that the physical state of the sample is critical for interpreting spectral data and model performance. In lines 185-187, we described the mechanical hulling of 214 samples to obtain clean millet grains. A total of 214 samples were collected, each weighing 250 g. The sample size and randomness met the requirements for later lab analysis. After harvesting, the samples were sun-dried naturally to below 13% moisture. then mechanically hulled to obtain clean, intact millet kernels.

The calculation and interpretation of RPD values are unclear. The authors should explicitly describe how RPD was computed and discuss whether the reported values (1.376 to 2.219) demonstrate adequate predictive performance. According to commonly accepted criteria, RPD values below ~2 generally indicate limited predictive capability. Please clarify whether the presented models can be reliably used for protein and fat prediction, and support this claim with appropriate references.

We thank the reviewer for the valuable comment regarding the clarity of RPD calculation and interpretation. 

In Section 2.6, we have included an explanation of RPD, with its calculation method detailed in the cited references. "Model performance was evaluated using the coefficient of determination (R2), RMSE, and relative percent deviation (RPD). An RPD > 2.0 indicates good to excellent predictive performance; an RPD between 1.4 and 2.0 suggests the model has moderate predictive capability; and an RPD < 1.4 signifies that the model lacks sufficient predictive accuracy [5,17,20]. All data processing and analyses were performed using MATLAB software (Version 2023b, MathWorks, Natick, MA, USA)."

In Section 4.1, we have included a discussion on how RPD is applied to evaluate models for fat and protein content prediction. Therefore, for the prediction of fat content, the RF model achieved an RPDP of 2.219. The PLS and SVM models also attained RPDP values of 1.969 and 1.950, respectively. This indicates that the RF model possesses excellent predictive capability and is suitable for direct use in practical applications. In contrast, the prediction models for protein content generally yielded lower RPDP values. The PLS model achieved the highest RPDp of 1.811, which is sufficient for rough estimation and can be applied to distinguish between high and low sample concentration ranges. Furthermore, models with RPD values between 1.8 and 2.0 still hold considerable potential for practical application [5].

Reviewer 3 Report

Comments and Suggestions for Authors

This study presents a near-infrared (NIR) hyperspectral approach for the rapid, non-destructive prediction of fat and protein content in foxtail millet (Setaria italica). The authors employ the Sparrow Search Algorithm (SSA) to identify stable key wavelengths (13 for fat and 15 for protein) and compare three machine learning models: Partial Least Squares (PLS), Random Forest (RF), and Support Vector Machine (SVM). Their main finding is that RF yields the best performance for fat prediction (RP2 =0.797, RMSEP = 0.218%), while PLS performs best for protein (RP2 =0.695, RMSEP = 0.268%). The use of SHAP (SHapley Additive exPlanations) to interpret feature contributions adds valuable transparency to the otherwise opaque “black-box” nature of RF.

The experimental design is robust and well-executed. Samples were collected under controlled agronomic conditions, reference values were obtained in triplicate using standardized methods (Soxhlet and Kjeldahl), and spectral preprocessing (Savitzky–Golay smoothing followed by Standard Normal Variate correction) was appropriately applied. The resulting models exhibit small gaps between calibration and prediction statistics (e.g., ΔR2 < 0.04 for both optimal models). This indicates that the wavelength selection strategy, based on frequency analysis across 50 independent SSA runs, effectively mitigates randomness and enhances reproducibility.

However, the validation strategy lacks a rigorous assessment of model robustness through internal validation techniques such as repeated k-fold cross-validation or bootstrap resampling. Although 5-fold cross-validation was used during wavelength selection, the final model performance is evaluated solely on a single hold-out partition (3:1 split). This approach does not quantify the variability of performance metrics due to data partitioning and may mask the influence of atypical or high-leverage samples. A repeated cross-validation scheme would provide estimates of the standard uncertainty in R2, RMSEP, and RPDP, thereby offering a more reliable basis for comparing models and assessing their generalization capability. Internal-external comparison alone is insufficient to establish model robustness in chemometric practice.

Furthermore, the manuscript does not address the applicability domain (AD) of the developed models, a critical omission for any predictive model intended for real-world deployment. For the PLS model, a Williams plot (leveraging leverage vs. residual distance) or Mahalanobis distance in the reduced wavelength space would readily define the boundaries of reliable prediction. For the RF model, given its non-linear nature, a k-nearest neighbors (k-NN) distance in the selected wavelength space offers a model-agnostic and interpretable AD criterion. Notably, k-NN distance can serve as a unified AD metric for both models, enabling consistent quality control in practical applications. Real users cannot determine whether a new sample falls within the chemical or spectral space represented by the calibration set, risking unreliable extrapolations.

Finally, the relatively modest prediction coefficients of determination (RP2 < 0.80) raise concerns about the measurement uncertainty associated with individual predictions. According to metrological best practices (GUM, ISO/IEC Guide 98-3), the total uncertainty of a NIR prediction should account for contributions from the reference method, the instrument, and the model itself. The authors should estimate how much additional uncertainty the modeling process introduces into the predicted value. This could be achieved through a sensitivity analysis (e.g., Monte Carlo propagation of the reference method’s uncertainty through the model) or by computing prediction intervals via bootstrap or Bayesian methods. Clarifying this would strengthen the metrological traceability and practical utility of the proposed models.

Author Response

Dear reviewer

Thank you very much for reviewing our manuscript. We would also like to express our gratitude to the reviewers for their efforts in helping to improve our manuscript titled “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (ID: foods-4118117). The comments were all valuable and very helpful for revising and improving our paper. Our research team have studied each reviewers’ comments point by point and have made necessary modifications and supplements, which we hope will meet with your approval. Revised portion are mark by “Track Change” in bold font throughout the revised manuscript with track change.

 

However, the validation strategy lacks a rigorous assessment of model robustness through internal validation techniques such as repeated k-fold cross-validation or bootstrap resampling. Although 5-fold cross-validation was used during wavelength selection, the final model performance is evaluated solely on a single hold-out partition (3:1 split). This approach does not quantify the variability of performance metrics due to data partitioning and may mask the influence of atypical or high-leverage samples. A repeated cross-validation scheme would provide estimates of the standard uncertainty in R2, RMSEP, and RPDP, thereby offering a more reliable basis for comparing models and assessing their generalization capability. Internal-external comparison alone is insufficient to establish model robustness in chemometric practice.

We sincerely thank the reviewer for their insightful and constructive comments. We fully agree with the point raised regarding the need to go beyond a single data split for model robustness assessment and to employ more rigorous internal validation techniques to quantify the uncertainty in performance metrics. This is crucial for comprehensively evaluating a model's generalizability, reliably comparing different models, and adhering to the best practices in chemometrics.

In this study, we chose to report the final model performance based on a single random hold-out split (3:1 ratio for training/prediction sets) after careful consideration of the following factors:

In the field of near-infrared spectroscopy quantitative analysis, particularly in applied research on non-destructive quality detection of agricultural products, it is a widely accepted and standard reporting practice to use a completely independent prediction set to report final performance metrics (R², RMSE). This approach aims to simulate the model's predictive performance when encountering entirely new, unseen samples in a real-world deployment scenario, providing an intuitive and unbiased estimate of its practical potential. Our reporting method aligns with numerous similar high-level studies in the field, ensuring the comparability of our results.

The effective sample sizes for fat and protein in this study are 214 and 192, respectively. To ensure that the evaluation using the independent prediction set has sufficient statistical power (it is generally recommended that the prediction set contains no fewer than 50 samples), we allocated approximately one-quarter of the samples to the prediction set. If we were to perform multiple repeated k-fold cross-validation (e.g., repeated 10-fold CV) on the remaining three-quarters designated as the training set, the size of each training subset would be further reduced. For ensemble models like Random Forest (RF), which require sufficient data to build stable and diverse decision trees, an overly small training subset may not adequately reflect the model's true performance on larger datasets and could even introduce a performance underestimation bias due to insufficient training data. Therefore, given the current sample size, we prioritized allowing the models—especially RF—to learn adequately on a training set of sufficient size, while using a substantial independent set for a single, clear performance validation.

The primary methodological innovations and contributions of this study are: (a) proposing a strategy that combines the Sparrow Search Algorithm with selection frequency statistics, aimed at screening stable and reliable key wavelengths from high-dimensional spectral data, thereby overcoming the randomness inherent in results from a single run; and (b) innovatively introducing the SHAP interpretability framework after model construction to quantify and elucidate the contribution patterns and decision logic of key wavelengths from both global and local perspectives. The design of our validation strategy aimed to clearly and directly demonstrate the final predictive efficacy of the models built upon these innovative methodologies, serving the argumentation of the core methodology.

Furthermore, the manuscript does not address the applicability domain (AD) of the developed models, a critical omission for any predictive model intended for real-world deployment. For the PLS model, a Williams plot (leveraging leverage vs. residual distance) or Mahalanobis distance in the reduced wavelength space would readily define the boundaries of reliable prediction. For the RF model, given its non-linear nature, a k-nearest neighbors (k-NN) distance in the selected wavelength space offers a model-agnostic and interpretable AD criterion. Notably, k-NN distance can serve as a unified AD metric for both models, enabling consistent quality control in practical applications. Real users cannot determine whether a new sample falls within the chemical or spectral space represented by the calibration set, risking unreliable extrapolations.

We thank the reviewer for their important comments. We understand that defining the Applicability Domain (AD) is crucial for the reliability and interpretability of predictive models in real-world applications. The methods indicated by the reviewer, such as the Williams plot, Mahalanobis distance, or k-NN distance, are indeed effective means to delineate the boundaries of reliable prediction for the models.

In the present study, our primary focus has been on model development and performance evaluation, and the explicit definition of the applicability domain has not been included within the scope of the current manuscript. We fully agree that in practical deployment, users require reliable criteria to determine whether a new sample falls within the chemical or spectral space covered by the calibration set to avoid unreliable extrapolations. The reviewer's suggestion to employ the k-NN distance as a unified AD metric, particularly for non-linear models like Random Forest, is a valuable and practical recommendation.

We will give full consideration to incorporating applicability domain assessment methods in subsequent research and applications to ensure the robustness and transparency of the models in practical scenarios. As this content is not included in the current manuscript, we will address it in the Limitations section and suggest it as a key aspect for future work. We once again thank the reviewer for their insightful and constructive comments.

Finally, the relatively modest prediction coefficients of determination (RP2 < 0.80) raise concerns about the measurement uncertainty associated with individual predictions. According to metrological best practices (GUM, ISO/IEC Guide 98-3), the total uncertainty of a NIR prediction should account for contributions from the reference method, the instrument, and the model itself. The authors should estimate how much additional uncertainty the modeling process introduces into the predicted value. This could be achieved through a sensitivity analysis (e.g., Monte Carlo propagation of the reference method’s uncertainty through the model) or by computing prediction intervals via bootstrap or Bayesian methods. Clarifying this would strengthen the metrological traceability and practical utility of the proposed models.

We thank the reviewer for their insightful comments. We understand that the coefficient of determination for the model predictions (R2<0.80) rightly raises reasonable concerns regarding the uncertainty of individual predicted values. As the reviewer pointed out, evaluating the total uncertainty in accordance with metrological best practices (e.g., GUM, ISO/IEC Guide 98-3) is a critical step in establishing the metrological traceability of the model and enhancing its practical utility.

In the present study, our primary focus has been on model development and fundamental performance validation. A detailed uncertainty evaluation and decomposition analysis for the predicted values has not been conducted. We agree that a comprehensive total uncertainty for an NIR prediction should systematically account for contributions from multiple sources, including the reference method, the instrument, and the model itself. The methods suggested by the reviewer—such as sensitivity analysis (e.g., Monte Carlo propagation of the reference method's uncertainty through the model) or computing prediction intervals via bootstrap or Bayesian methods—are indeed effective and rigorous approaches for quantifying the additional uncertainty introduced by the modeling process.

Clarifying this aspect would undoubtedly significantly enhance the metrological quality of the model results. We will address this point in the Limitations section of the manuscript and propose that systematic uncertainty evaluation be implemented as a core step in subsequent in-depth research and practical application deployment. We once again thank the reviewer for this highly constructive and important feedback.

Reviewer 4 Report

Comments and Suggestions for Authors

The ms. “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (Ms. Ref. No. foods-4118117-v1) presents a comprehensive study on the rapid, non-destructive prediction of fat and protein content in foxtail millet using near-infrared (NIR) hyperspectral imaging combined with machine learning and model interpretability (SHAP). The authors integrate the Sparrow Search Algorithm (SSA) for key wavelength selection with PLS, RF, and SVM models, and further provide interpretability analysis, which is a notable strength. There is a lot of work involved and the ms. has evident merit.

The topic falls within the aims and scopes of the Foods journal.

However, several clarifications and improvements are required.

Major issues:

  1. All samples come from a single ecological region and a single variety (“Changnong 47”). This significantly limits the generalizability of the developed models. This is a limitation it should be emphasized more clearly in the Abstract and Conclusions. The authors should explicitly state that the models are currently variety- and region-specific. In the Conclusion section, future work should be proposed to include multi-variety and multi-location datasets to enhance robustness.
  2. 22 protein outliers (10.28%) were removed, which is a relatively large proportion and may bias the dataset. Therefore, please provide a clearer justification for removing these outliers (e.g., analytical errors, biological plausibility). It is not clear if models were tested with and without outliers to assess robustness.
  3. Please consider including a brief sensitivity analysis.
  4. The study relies on a single random hold-out split (3:1), which may not fully reflect model stability. Please consider adding repeated random splits or external validation. Alternatively, you can justify more explicitly why hold-out was preferred over full cross-validation for final model evaluation.
  5. Important hyperparameters are not fully reported. For reproducibility, please clarify:
  • SSA population size, number of discoverers/followers/sentinels, mutation probabilities, and numbr of iterations.
  • RF parameters (number of trees, max depth, feature sampling).
  • SVM kernel type and hyperparameter tuning strategy.
  • PLS latent variable selection method.

Minor issues:

  1. Some sentences are very long and could be shortened for readability.
  2. Please ensure consistent formatting of °C, %, nm, and equations.
  3. Please clarify whether RMSEP values are in % or absolute units consistently.

Given the completed score sheet and the comments above, after careful evaluation, the ms. “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (Ms. Ref. No. foods-4118117-v1) needs Major Revision according to comments before being considered for publication in Foods journal.

Comments on the Quality of English Language
  1. Some sentences are very long and could be shortened for readability.
  2. Minor English corrections are needed.

The ms. “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (Ms. Ref. No. foods-4118117-v1) presents a comprehensive study on the rapid, non-destructive prediction of fat and protein content in foxtail millet using near-infrared (NIR) hyperspectral imaging combined with machine learning and model interpretability (SHAP). The authors integrate the Sparrow Search Algorithm (SSA) for key wavelength selection with PLS, RF, and SVM models, and further provide interpretability analysis, which is a notable strength. There is a lot of work involved and the ms. has evident merit.

The topic falls within the aims and scopes of the Foods journal.

However, several clarifications and improvements are required before the manuscript ccould be considered for publication.

Major issues:

  1. All samples come from a single ecological region and a single variety (“Changnong 47”). This significantly limits the generalizability of the developed models. This is a limitation it should be emphasized more clearly in the Abstract and Conclusions. The authors should explicitly state that the models are currently variety- and region-specific. In the Conclusion section, future work should be proposed to include multi-variety and multi-location datasets to enhance robustness.
  2. 22 protein outliers (10.28%) were removed, which is a relatively large proportion and may bias the dataset. Therefore, please provide a clearer justification for removing these outliers (e.g., analytical errors, biological plausibility). It is not clear if models were tested with and without outliers to assess robustness.
  3. Please consider including a brief sensitivity analysis.
  4. The study relies on a single random hold-out split (3:1), which may not fully reflect model stability. Please consider adding repeated random splits or external validation. Alternatively, you can justify more explicitly why hold-out was preferred over full cross-validation for final model evaluation.
  5. Important hyperparameters are not fully reported. For reproducibility, please clarify:
  • SSA population size, number of discoverers/followers/sentinels, mutation probabilities, and numbr of iterations.
  • RF parameters (number of trees, max depth, feature sampling).
  • SVM kernel type and hyperparameter tuning strategy.
  • PLS latent variable selection method.

Minor issues:

  1. Some sentences are very long and could be shortened for readability.
  2. Please ensure consistent formatting of °C, %, nm, and equations.
  3. Please clarify whether RMSEP values are in % or absolute units consistently.

Given the completed score sheet and the comments above, after careful evaluation, the ms. “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (Ms. Ref. No. foods-4118117-v1) needs Major Revision according to comments before being considered for publication in Foods journal.

Author Response

Dear reviewer

Thank you very much for reviewing our manuscript. We would also like to express our gratitude to the reviewers for their efforts in helping to improve our manuscript titled “Development and Interpretability Innovation Analysis for Fat and Protein in Foxtail Millet [Setaria italica (L.) Beauv.] Using Near-Infrared Spectroscopy” (ID: foods-4118117). The comments were all valuable and very helpful for revising and improving our paper. Our research team have studied each reviewers’ comments point by point and have made necessary modifications and supplements, which we hope will meet with your approval. Revised portion are mark by “Track Change” in bold font throughout the revised manuscript with track change.

 

All samples come from a single ecological region and a single variety (“Changnong 47”). This significantly limits the generalizability of the developed models. This is a limitation it should be emphasized more clearly in the Abstract and Conclusions. The authors should explicitly state that the models are currently variety- and region-specific. In the Conclusion section, future work should be proposed to include multi-variety and multi-location datasets to enhance robustness.

We thank the reviewer for their important comment. We fully acknowledge that the samples used in this study originate from a single ecological region and a single variety (“Changnong 47”). This limitation does indeed affect the generalizability and extrapolation capability of the developed models. We agree that this point should be emphasized more clearly in the Abstract and Conclusions.

Following your suggestion, we will explicitly state in the Abstract that the models developed in this study are currently applicable to a specific variety and ecological region. Furthermore, in the Conclusions section, we will more prominently highlight this limitation and specifically propose future research directions—namely, to validate and extend the model's robustness and applicability by incorporating datasets from multiple varieties and geographical locations. While this limitation has been mentioned in the Discussion, we will strengthen and sharpen its presentation in the Conclusions as advised.

We appreciate the reviewer's insightful suggestion, which will help enhance the rigor and transparency of our study.

22 protein outliers (10.28%) were removed, which is a relatively large proportion and may bias the dataset. Therefore, please provide a clearer justification for removing these outliers (e.g., analytical errors, biological plausibility). It is not clear if models were tested with and without outliers to assess robustness.

We thank the reviewer for raising this important point. We fully acknowledge the significance of the outlier removal proportion and its potential impact on the dataset, and we will provide a clearer explanation in the revised manuscript.

Regarding the specific justification for removing these 22 protein outliers (10.28% of the total samples), we will add the following clarification in the Methods (or Data Preprocessing) section:

These samples were identified as statistical extremes because their measured values significantly deviated from the normal distribution assumption of the main dataset. In building robust regression models, retaining such extreme values could excessively distort variable relationships, leading the model to capture individual noise rather than general patterns.

As correctly noted by the reviewer, these identified outliers did not participate in the training or validation processes of any subsequent models. Therefore, all model performance metrics reported (such as the coefficient of determination and errors) were calculated based on the dataset after outlier removal.

Regarding the reviewer's concern about model robustness assessment, we acknowledge that the study did not systematically compare model performance between datasets "with outliers included" and "with outliers excluded" at the design stage. This is a valuable perspective that directly tests a model's sensitivity to data anomalies. We will explicitly state this point in the Discussion or Limitations section and suggest incorporating such an analysis as a standard step in future work for model validation and robustness evaluation.

We thank the reviewer again for the rigorous review, which has prompted us to more thoroughly examine the transparency of our data preprocessing steps and the robustness of our models.

The study relies on a single random hold-out split (3:1), which may not fully reflect model stability. Please consider adding repeated random splits or external validation. Alternatively, you can justify more explicitly why hold-out was preferred over full cross-validation for final model evaluation.

We thank the reviewer for raising this important point. We fully understand that evaluating the model with a single random hold-out split (3:1) may not sufficiently reflect model stability, and we appreciate the suggestion to employ repeated random splits or external validation.

We recognize that methods such as repeated cross-validation typically provide more robust performance estimates. In this study, the decision to use a single random split for final evaluation was primarily based on practical computational resource considerations, as well as preliminary analyses indicating relatively consistent model performance across multiple random subsets. Nevertheless, we agree that this remains a methodological limitation and may not fully capture potential fluctuations in model performance.

As emphasized by the reviewer, to enhance the reliability of the conclusions, adopting repeated random splits or validating on a completely independent external dataset with temporal or spatial differences would be a more rigorous and preferable approach in future research or practical deployment. We will more explicitly clarify this limitation of the current validation approach in the Discussion or Limitations section of the manuscript and clearly propose implementing more robust validation strategies as a focus of subsequent work.

We once again thank the reviewer for this insightful suggestion, which will help improve the rigor of our research methodology.

 

Important hyperparameters are not fully reported. For reproducibility, please clarify:

SSA population size, number of discoverers/followers/sentinels, mutation probabilities, and numbr of iterations. RF parameters (number of trees, max depth, feature sampling). SVM kernel type and hyperparameter tuning strategy. PLS latent variable selection method.

We thank the reviewer for this important comment. We fully agree that reporting all key hyperparameters in detail is crucial for ensuring the reproducibility of the study. As requested, we have clearly supplemented and specified the core hyperparameters for each algorithm in the Methods section of the manuscript, as detailed below:

Sparrow Search Algorithm (SSA): The population size, the number of discoverers/followers/sentinels, mutation probabilities, and the number of iterations are now explicitly stated.

Random Forest (RF) Model: The core parameters of the RF model have been detailed, including the number of trees, maximum depth, and the feature sampling strategy considered for splitting at each node.

Support Vector Machine (SVM) Model: The kernel type used for the SVM (e.g., Radial Basis Function kernel) is now clearly stated, along with an overview of the strategy employed for tuning its key hyperparameters (such as the regularization parameter C and kernel coefficient gamma), for example, grid search or Bayesian optimization.

Partial Least Squares (PLS) Regression Model: The method used to determine the optimal number of latent variables (e.g., based on cross-validation error minimization) has been clearly explained.

We believe these additions adequately address your concern regarding model reproducibility. We sincerely thank the reviewer once again for their meticulous and rigorous attention, which has greatly enhanced the completeness and scientific rigor of our manuscript.

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

Dear authors,

the quality of the manuscript has been improved, and can be endorsed for publication now. Thank you.

The tile sounds better but needs to be revised:

Development and Interpretability Analysis of Near-Infrared Spectroscopy Models for Fat and Protein Prediction in Foxtail Millet (Setaria italica [L.] Beauv.)

or

Development, Innovation, and Interpretability of Near-Infrared Spectroscopy Models for Fat and Protein Assessment in Foxtail Millet (Setaria italica [L.] Beauv.)

Author Response

Thank you for your valuable comments. We have revised the article title to: Development and Interpretability Analysis of Near-Infrared Spectroscopy Models for Fat and Protein Prediction in Foxtail Millet [Setaria italica (L.) Beauv.]

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

The reviewer sincerely appreciate the authors’ detailed reply and the effort invested in addressing the concerns. However, after careful consideration of the methodological issues and based on recent scientific evidence, the current justification remains insufficiently convincing regarding the assessment of model robustness in machine learning applications within chemometrics and near-infrared spectroscopy.

A recent systematic review published in Journal of Chemometrics underscores that the reliability of chemometric models, particularly those applied to near-infrared spectroscopy (NIRS) data, fundamentally depends on rigorous validation protocols that go beyond single external splits [10.1002/cem.70036]. The manuscript examined validation strategies across a large sample of published regression models for NIRS analysis and found that while external validation is desirable, many studies rely solely on cross-validation methods. Such reliance, in the absence of properly designed internal validation, can lead to over-optimistic performance estimates and hinder meaningful comparisons of predictive metrics across studies. The review emphasises that cross-validation strategies (such as k-fold, repeated resampling, or nested cross-validation) should form part of a comprehensive validation framework that provides more reliable and interpretable figures of merit.

Another review in Chemosensors, reinforces these principles from a practical perspective [10.3390/chemosensors10080323]. The authors suggested alternative internal validation strategies such as bootstrapping and random subsampling, which can provide less biased estimates of predictive accuracy than naive use of hold-out splits alone, particularly when sample sizes are limited.

Taken together, these two bodies of literature suggest that internal validation should not be treated as an optional step but as an integral part of model development and assessment in chemometrics, especially when machine learning techniques are used with NIRS data. By quantifying the variability of performance metrics under different resampling schemes and ensuring that model tuning and evaluation do not inadvertently use the same information, internal validation provides a more realistic characterization of model robustness and generalisability. Accordingly, the reviewer respectfully suggest that the manuscript be revised to include an internal validation analysis.

Author Response

Regarding your question about the model internal validation mechanisms, thank you for your review and valuable feedback. We have already described the internal validation strategies for the three models in the manuscript. To provide further clarification, we would like to elaborate on the relevant content and design rationale as follows:

In lines 315–316, we state that for the PLS model, "The main hyperparameters were optimized using a random search strategy." We used this method to systematically evaluate the prediction error of the model with different numbers of latent variables during training. This allows the model to automatically determine the optimal complexity that captures the underlying data signals without overfitting to noise. This process is conducted internally on the training set, which helps ensure that the final model has better generalization ability.

For the Random Forest model, lines 323–325 describe its parameter settings and mention that internal performance is evaluated using out-of-bag error. “The main parameters were configured as follows: the number of trees is set to 100, the minimum leaf size is 5, and the maximum depth is determined based on the minimum leaf node count, thus controlling model complexity.” When building each tree in the forest, about one-third of the data is not sampled. These "out-of-bag" samples serve as a validation set for that tree. By aggregating the out-of-bag prediction errors from all trees, we obtain an efficient and unbiased internal validation estimate without needing to split the data further. We set the number of trees to 100 and constrained the minimum leaf size and depth to maintain predictive performance while using out-of-bag error to monitor and control overfitting risk.

Regarding the Support Vector Machine, line 333 states that "The main hyperparameters were optimized using a random search strategy." During this process, we randomly sampled combinations of hyperparameters from a predefined search space. This strategy allows an efficient search over a broad range of values and avoids the computational burden of grid search.

In summary, the internal validation mechanisms used for each model—including cross-validation, out-of-bag estimation, and random search—are designed to optimize key model parameters and objectively evaluate model performance during the training phase. This provides a basis for selecting the final predictive model and enhances the reliability and reproducibility of the research results.

Thank you again for your valuable comments.

Author Response File: Author Response.pdf

Round 3

Reviewer 3 Report

Comments and Suggestions for Authors

The authors have addressed the main concern regarding internal validation. The manuscript is acceptable for publication.

Back to TopTop