Review Reports
- Oybek Tukhtamishov 1,
- Mohamed Fawzy 2,3,* and
- Zokhid Mamatkulov 1
- et al.
Reviewer 1: Heba Bedair Reviewer 2: Anonymous Reviewer 3: Faisal Mumtaz
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsUsing numerous vegetation indicators and multi-sensor satellite datasets (Landsat-8, Sentinel-1, and Sentinel-2), the manuscript offers a thorough framework for crop classification that is assessed using a number of machine learning classifiers. The study is pertinent and current, especially for semi-arid agricultural areas like Uzbekistan. The integration of vegetation indices with optical and SAR data is highly motivated, and the outcomes are encouraging. To improve the manuscript, a few points need to be clarified, improved, and discussed in greater detail.
Although the workflow (Fig. 3) is helpful, some of the steps are not fully explained. For instance:
In Sentinel-2 composites, how were cloud masks applied?
Before integration, were SAR backscatter readings adjusted or calibrated?
In what way were vegetation indices aggregated over time (monthly composites versus phenological stages)?
Reproducibility will increase with greater methodological clarity.
Method of Sampling
Spatial autocorrelation between training and validation sets is known to be a concern associated with the random sampling strategy. This is a crucial restriction. To guarantee more reliable accuracy estimations, think about talking about different approaches (such as field-level partitioning and spatial cross-validation).
Comparing Classifiers
There is no explanation of why CART and GBT performed marginally worse consistently, despite Random Forest outperforming other classifiers. Please elaborate on the causes of overfitting or noise sensitivity, as well as whether adjusting the parameters could enhance their functionality.
Accuracy Metrics
The manuscript reports overall accuracy and Kappa, but producer’s and user’s accuracy are only briefly mentioned. A more detailed breakdown of class-specific accuracies (cotton vs. wheat vs. other crops) would provide stronger evidence of the model’s reliability.
Figure 1 and Figure 2 are informative but could benefit from clearer legends and higher resolution.
Tables 1 and 2 are well-structured, but please ensure consistency in band naming (e.g., Sentinel-2 B8 vs. B8A).
Author Response
Dear reviewers,
Many thanks for your time and efforts during reviewing our manuscript entitled by “A Comprehensive Machine Learning Approach for Crop Classification Using Multi-Sensor Satellite Datasets and Multiple Vegetation Indices”.
All reviews, corrections and recommendations are appreciated as they add more value to the work. We hope the paper, after careful revisions, meets your high standards. The authors would really welcome any additional constructive comments.
The attached file "Response-to-reviewers" contains point-by-point responses for reviewer’s comments in blue with the authors’ response in green.
The attached file "Revised manuscript" includes the applied modifications highlighted in green.
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsDear authors, thank you for submitting your work, which appears to be interesting and addresses a relevant aspect of the literature. However, before it can be accepted, some significant changes must be made to the manuscript. Please find specific comments below.
Major comments
1) The abstract should be slightly shortened and made more concise.
2) The introduction justifies the use of machine learning, vegetation indices and Sentinel-2 data for agricultural monitoring, but the literature cited in this key section is markedly skewed towards studies on crop type classification, and completely omits the substantial body of literature that employs the exact same methodological paradigm — Satellite + vegetation indices + Random Forest — to estimate biophysical and physiological variables in crops in semi-arid environments. I strongly recommend that the authors cite and discuss the work of Giannico et al. (2024), Remote Sensing 16(24), 4784, https://doi.org/10.3390/rs16244784 . This study is highly pertinent for at least four reasons: (i) it uses exactly your sensor/algorithm combination (Sentinel-2 bands and VIs as predictors, RF benchmarked against regularized linear models); (ii) it operates in a semi-arid environment ecologically analogous to Urta Chirchik; (iii) it provides quantitative evidence that RF on multi-VI inputs outperforms both linear models and raw spectral bands, which is precisely the rationale underpinning your Scenario-2; and (iv) it identifies the red-edge region as the most informative spectral domain via permutation importance, directly supporting your repeated claim about NDRE and the red-edge bands.
3) Spatial autocorrelation: state the limitation more clearly. You describe the bias from random pixel-wise splitting as "slightly optimistic". This understates the issue: with multiple pixels per parcel potentially falling in both training and test sets, OA can be inflated by several percentage points. Please remove the "slightly" qualifier, state explicitly that no parcel-level or spatial block cross-validation was performed, note that small absolute differences between datasets and scenarios should therefore be interpreted with caution, and flag spatially independent validation as a recommended next step.
4) Hyperparameter choices. Some of your settings (lines 267–273) are problematic:
-KNN with k = 100 is essentially a global-majority vote given 833 training points on 7 classes (~119 per class). This provides a simple explanation for the poor performance of KNN and undermines the "curse of dimensionality" interpretation at lines 511–512. Please re-run with k tuned via internal cross-validation.
-CART with 100 leaf nodes is fixed a priori without justification. Tree complexity should be controlled via cost-complexity pruning, not a hard cap.
-RF (100 trees) and GBT (100 iterations) are plausible defaults but not optimised. Given the manuscript's "comprehensive" claim, please tune at least mtry for RF and learning rate × n_estimators for GBT.
-Reference [32] (Delpisheh et al. on porous media) is unrelated and should be replaced.
5) Line 569 states cotton UA reaches 100% "through RF and GBT classifiers", but Table 7 shows RF cotton UA = 88.89%; only GBT reaches 100%. Correct the wording.
6) Differences of 2–3% on n = 357 test points may not be statistically significant. Please add at least pairwise McNemar tests on the most relevant confusion matrices, or bootstrap confidence intervals on OA and Kappa
7) Sentinel-1 contribution is overstated relative to evidence. You use only the seasonal median of VV, VH and VV/VH in DS-3, foregoing the temporal richness that is SAR's main asset. The marginal gain of DS-3 over DS-2 is 0.22–3.92% in Scenario-2 — and just 0.22% for the headline RF + All VIs comparison, well within sampling noise. Please tone down the claims about SAR–optical fusion in the abstract and conclusions, and acknowledge the limitation of using only seasonal medians.
Minor comments
1) The manuscript needs thorough English editing: tense agreement, missing articles, and awkward constructions.
2) Please verify all figure numbering.
3) Reference [37] seems to be a thesis; replace with a peer-reviewed equivalent if available.
4) Figures 9–13 are visually almost indistinguishable. Consider keeping only the best (RF on DS-3) and worst (MD on DS-1) maps in the main text, with a disagreement map highlighting where classifiers diverge; move the rest to supplementary.
The work has clear potential. Addressing the bibliographic gap, the machine learning design, and the internal inconsistencies would, in my view, substantially raise both the rigour and the credibility of the contribution.
Author Response
Dear reviewers,
Many thanks for your time and efforts during reviewing our manuscript entitled by “A Comprehensive Machine Learning Approach for Crop Classification Using Multi-Sensor Satellite Datasets and Multiple Vegetation Indices”.
All reviews, corrections and recommendations are appreciated as they add more value to the work. We hope the paper, after careful revisions, meets your high standards. The authors would really welcome any additional constructive comments.
The attached file "Response-to-reviewers" contains point-by-point responses for reviewer’s comments in blue with the authors’ response in green.
The attached file "Revised manuscript" includes the applied modifications highlighted in green.
Author Response File:
Author Response.docx
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript "A Comprehensive Machine Learning Approach for Crop Classification Using Multi-Sensor Satellite Datasets and Multiple Vegetation Indices" addresses an applied and relevant topic: crop classification in Uzbekistan using Landsat-8, Sentinel-2, and Sentinel-1 imagery, vegetation indices, and machine-learning classifiers. The study may be useful as a local case study. However, the manuscript currently overstates its novelty and conclusions, while several methodological and validation weaknesses limit the reliability of the reported accuracies.
Major Issues,
- Lines: 251–253 and 429–431. The manuscript uses June 10 to July 22, but the temporal VI results indicate wheat peaks in April and cotton peaks in late August. The authors should justify this window and also use time-series features.
- Lines: 284–320. The preprocessing is not reproducible enough. The manuscript does not provide exact GEE collections, scale factors, cloud-mask bits, cloud thresholds, image counts, compositing rules, SAR units, or full resampling details.
- Line 121–128. The novelty there is overstated. The manuscript presents a "comprehensive multi-sensor and multi-index machine learning framework," but the workflow mainly combines commonly used sensors, vegetation indices, and standard classifiers. All the claims made there indicate that the authors haven't consulted sufficient literature and that are mainly based on the authors' perceptions.
- Lines: 213–214 and 265–275. The manuscript uses a random 70/30 split of pixel samples and acknowledges the potential spatial autocorrelation between neighboring training and validation points. This likely inflates accuracy. Field-level or spatial block validation is needed.
- Section 3.5. Line 400-402. The authors have used the Kappa coefficient, which is already dead and is not considered a valid indicator. Please have a careful check at https://www.tandfonline.com/doi/full/10.1080/01431161.2011.552923
- Lines: 186–205. The manuscript mentions 100 cotton fields, 100 wheat fields, 76 cultivated areas, and 30 non-vegetated regions. Still, it does not adequately report verification dates, boundary accuracy, field-survey protocol, or the independence of the reference data, all of which are critical.
- Lines: 206–208, 244–250, and 276–278. DS-1, DS-2, and DS-3 differ in resolution, spectral bands, red-edge availability, revisit frequency, and SAR content. Therefore, the superiority of DS-2 or DS-3 cannot be attributed to a single factor without controlled ablation tests.
- Lines: 236–238, 547–550, and Table 4, Lines 517–519. Table 4 shows that not all VIs consistently improve accuracy. For example, RF decreases for DS-1 and DS-3, and CART decreases for DS-3 when all VIs are used. The claim should be revised.
- Lines: 44–46 and Table 4, Lines 517–519. The statement that RF consistently outperformed other classifiers is inaccurate. Table 4 shows that GBT equals RF for DS-3 with all VIs and gives the highest DS-2 all-VI accuracy. The following claims should be thoroughly corrected throughout the manuscript's
- Lines: 531–532 and 217–221. DS-3 improves over DS-2 by only 0.22% to 3.92% in scenario 2, which may not be meaningful given the small validation set and spatial dependence. Confidence intervals or significance tests should be added.
- Lines: 265–272. RF, GBT, CART, and KNN parameters are fixed without reported tuning, cross-validation, sensitivity analysis, or random seed. The use of KNN with 100 neighbors is especially questionable.
- Lines: 153–158 and 284–288. The Landsat preprocessing description is inconsistent. The manuscript first states that Landsat-8 surface reflectance channels were used, but later describes the processing of TOA radiance and reflectance. The product type and scaling procedure should be clarified.
- Lines: 531–541. DS-3 gains are attributed to SAR sensitivity to canopy structure and moisture, but no Sentinel-1-only model, feature-importance analysis, or SAR ablation is provided.
- The class scheme is unclear. The manuscript uses seven classes, but definitions for "other crops," "vacant land," and non-vegetated classes are limited. Full class-wise confusion matrices and PA/UA values for all seven classes should be provided.
- Lines: 447–454. The manuscript reports CFS results but does not state whether CFS was applied only to training data or whether validation leakage was avoided. Relevant
- Lines: 384–390 and 512–516. The manuscript describes MD as Euclidean centroid-based classification, but later attributes its poor performance to normal-distribution and equal-variance assumptions, which are not central to simple MD classification.
- There is a figure cross-reference error. The methodology text refers to the workflow as Figure 4, but the workflow diagram is labeled Figure 3.
- Table 2 has a Sentinel-2 band-labeling error.
- The conclusions are too strong. The manuscript claims a robust and transferable framework, but the study lacks independent validation, statistical uncertainty analysis, controlled ablations, and transfer testing. The conclusions should be limited to the local case study.
- Line 193. satellite imagers???
- Line 297. What is zero sigma???
- lines 286–287. radiance& reflectance" should be "radiance and reflectance,"
- The manuscript requires mandatory professional English editing before it can be considered further. There are frequent grammatical errors, awkward sentence constructions, incorrect word choices, and inconsistent technical terminology that reduce clarity and weaken the scientific presentation.
Author Response
Dear reviewers,
Many thanks for your time and efforts during reviewing our manuscript entitled by “A Comprehensive Machine Learning Approach for Crop Classification Using Multi-Sensor Satellite Datasets and Multiple Vegetation Indices”.
All reviews, corrections and recommendations are appreciated as they add more value to the work. We hope the paper, after careful revisions, meets your high standards. The authors would really welcome any additional constructive comments.
The attached file "Response-to-reviewers" contains point-by-point responses for reviewer’s comments in blue with the authors’ response in green.
The attached file "Revised manuscript" includes the applied modifications highlighted in green.
Author Response File:
Author Response.docx
Round 2
Reviewer 3 Report
Comments and Suggestions for AuthorsAuthors' revisions are satisfactory
Author Response
Comments and Suggestions for Authors: Authors' revisions are satisfactory.
We sincerely thank the Reviewer for evaluating our revised manuscript and for the positive assessment. We appreciate the valuable comments and guidance provided throughout the review process, which helped us improve the quality and clarity of the manuscript.