From Traditional Machine Learning to Fine-Tuning Large Language Models: A Review for Sensors-Based Soil Moisture Forecasting
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis article provides a relevant overview of sensor-based soil moisture (SM) forecasting. Its primary strengths are its logical structure and comprehensive taxonomy, which effectively organizes the field for readers.
However, in my opinion, the manuscript would benefit from a significant revision to improve its methodological rigor, analytical depth, and the validity of its conclusions. The following points outline the suggested changes.
- Methodological Rigor and Reproducibility (PRISMA)
The application of the PRISMA methodology is insufficiently detailed, which compromises the review's reproducibility. The authors could include:
* Specific Search Strings: The exact search strings used for each database (e.g., IEEE Xplore, Scopus).
* Selection Process Details: The specific dates for data collection from each source and the number of reviewers involved in each stage.
* Protocol Documentation: A link to a registered protocol (e.g., on OSF or PROSPERO) and the data extraction/coding sheet used.
- Analytical Weaknesses and Lack of Standardization
The comparative analysis is superficial and lacks the necessary harmonization to draw robust conclusions.
* Conceptual Inconsistency: The paper fails to distinguish clearly between Soil Moisture (SM) and Soil Water Potential (SWP). The authors must define both terms, state their units (e.g., m³/m³, kPa), and explain the implications of forecasting one versus the other.
* Inappropriate Comparisons: The review compares studies using different metrics (e.g., RMSE, MAE, R²) and forecast horizons without any normalization. This makes direct comparisons between models inappropriate. The analysis must be segmented by forecast horizon, and normalized metrics (e.g., NRMSE) should be used to allow for valid comparisons.
- Incomplete Coverage and Overstated Conclusions
The article points to very relevant future directions. The review could be even more complete by integrating a discussion of existing works in these areas and by framing the conclusions about LLMs more cautiously.
- Physics-Informed Neural Networks (PINNs) and Hydrology: The review mentions this area as a future direction, but there are already studies applying PINNs or models based on the Richards equation to soil water dynamics. Mapping these pioneering studies would strengthen the state-of-the-art analysis.
- Edge Computing and TinyML: The recommendation on TinyML would be strengthened by reviewing case studies that already apply models on microcontrollers (MCUs) or edge devices, presenting metrics on latency, power consumption, and model size.
- LLMs as an Emerging Frontier: The section on LLMs is based on a few studies with direct application to SM. Although the potential is immense, the current conclusions may sound like a fact of superiority. We suggest reframing LLMs as a highly promising but still emerging approach in its early stages of empirical validation, highlighting that their widespread effectiveness for direct SM forecasting is still a hypothesis to be confirmed.
- Propose a Benchmark Protocol**
To address the lack of standardization identified in the review, the article could propose a minimum benchmark protocol for future research. This would be a significant contribution to the field. The protocol must specify:
* Reference Datasets: Standard open datasets (e.g., ISMN, ERA5, NASA POWER)
* Standardized Inputs: A core set of predictor variables.
* Fixed Metrics and Horizons: Required evaluation metrics (e.g., RMSE, NSE, R²) for fixed forecast horizons (e.g., 1, 3, and 7 days).
* Data Splitting Rules: Clear guidelines for temporal cross-validation to prevent data leakage.
In conclusion, this article has the potential to become a valuable reference. However, addressing the critical issues outlined above is essential to strengthen its scientific contribution, rigor, and overall impact.
Author Response
We warmly thank the Editor and the Reviewers for their useful and thought-provoking comments. We did our best to comply with all the raised concerns, as detailed in the following. In the revised version, the main changes and additions are highlighted in blue, to ease the tracking of the changes, while minor variations and corrections of typos have not been tracked. In the following, we provide detailed responses to the specific issues raised by the Reviewers.
Please see the file attached.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript «From Traditional Machine Learning to Fine-Tuning Large Language Models: A Review for Sensors-Based Soil Moisture Forecasting» presents a comprehensive systematic review covering 189 publications, of which 68 were selected under PRISMA criteria for the period 2017–2025. The paper analyzes the evolution of soil moisture forecasting methods, tracing the transition from traditional machine learning to deep learning and large language model (LLM) approaches. The study is notable for its broad coverage and an attempt to integrate modern AI paradigms—Federated Learning, Transfer Learning, and fine-tuning of LLMs—within a unified taxonomy, which gives the work conceptual depth and interdisciplinary relevance.
However, the critical evaluation reveals that the article requires substantial methodological and analytical revision. Despite referencing the PRISMA framework, the authors have not provided a complete search protocol: there is no registration number, list of databases, key search strings, or quantitative inclusion and exclusion criteria. Out of 189 identified studies, 77 were excluded, yet the exclusion criteria are not explained (e.g., language, full-text availability, duplication). The authors should provide a fully documented protocol detailing data sources (IEEE Xplore, Scopus, SpringerLink, etc.), temporal coverage, search phrases, and the justification for each exclusion stage. A Risk-of-Bias Assessment Table should be included (e.g., low/medium/high bias categories), along with a measure of inter-reviewer agreement (Cohen’s κ).
The section 2.5 “AI-Driven Novel Taxonomy for Soil Moisture Forecasting” is well structured, yet the originality of the taxonomy remains insufficiently substantiated. Comparison in Table 1 shows that similar taxonomic schemes were previously presented in [9], [11], [13], and [15]; the present taxonomy differs mainly by adding five algorithmic classes (ML, DL, LLM, Hybrid, Paradigm) and four deployment levels (Cloud, Edge, MCU, TinyML). The authors should explicitly identify which features are genuinely novel—such as the inclusion of LLM-based forecasting and “privacy-aware model training”—and present a comparative table quantifying new attributes versus earlier surveys (e.g., percentage of dimensions covered).
Tables 2–5 comprise 93 entries (22 ML, 33 DL, 3 LLM, 12 Hybrid, and 8 Learning Paradigm) and contain an excessive level of detail that hinders comparability. The authors are advised to reduce the tables by approximately 30–40 %, retaining only representative case studies (two or three per model class) and moving the full versions to an appendix. Evaluation metrics such as MAE, MSE, RMSE, R², NSE, and SSIM are reported without reference to forecast horizon (t+1, t+5, seasonal) or sensor depth (5–100 cm), making quantitative comparisons unreliable. The study should include normalized metrics (nRMSE, KGE) and a summary table with mean performance per model class, for instance:
- traditional ML (R² = 0.81 ± 0.06; nRMSE = 14 %),
- deep learning (R² = 0.87 ± 0.05; nRMSE = 10 %),
- hybrid models (R² = 0.90 ± 0.04; nRMSE = 8 %),
- LLM (TimeGPT) (R² = 0.91 ± 0.03; MAE = 0.042 moisture units).
The discussion of sensor technologies lacks a clear classification by measurement unit. The text intermixes SM, SWP, and SEC without defining ranges or units (m³/m³, kPa, dS/m). A dedicated table should be added: “Sensor – Measured Variable – Range – Units – Typical Error,” for example: FDR (0.05–0.45 m³/m³, ±2 %), TDR (0.08–0.50 m³/m³, ±1.5 %), capacitive (±3 %), tensiometer (0 to –85 kPa, ±1 kPa).
Section 3.2 (Deep Learning Models) presents 35 examples but lacks quantitative comparison between DL and ML approaches. A figure or table showing average R² or RMSE distributions for both categories should be included, indicating mean improvement (e.g., RMSE reduction of 3.8 ± 1.2 %). Section 3.3 (Hybrid Models) would benefit from explicit percentage gains relative to DL models (R² increase ≈ 3–4 %, RMSE decrease ≈ 1.5 %).
The LLM section (Table 4) currently relies on only three studies [64], [82], [83], one of which provides quantitative results (MAE = 0.041, RMSE = 0.057 for a 5-day horizon in Belgium). The publication status of these sources (peer-reviewed or preprint) should be clarified. A comparative analysis of computational cost is needed: for example, TimeGPT fine-tuning ≈ 2.5 hours on an NVIDIA A100 GPU (energy ≈ 1.8 kWh) versus LSTM training ≈ 0.6 hours (0.4 kWh). The section should specify fine-tuning parameters (epochs, local dataset size, ratio of local to pre-training data).
In Section 3.4 the effectiveness of Transfer and Federated Learning is described qualitatively. Reported data [74–77] indicate that ConvLSTM with transfer learning reduces RMSE by 15 % (to 0.056) and federated training across 14 edge devices decreases accuracy by only 9 % while preserving privacy. A consolidated figure summarizing these gains relative to centralized learning should be added.
Stylistically, the manuscript would benefit from reducing sentence length (currently 36–40 words on average in Sections 2–3), eliminating redundancy, and standardizing terminology. The word “forecasting” should be used consistently in place of alternating “prediction” and “estimation.”
The Conclusion should emphasize quantitative findings—such as the proportion of model types analyzed (ML ≈ 32 %, DL ≈ 51 %, Hybrid ≈ 13 %, LLM ≈ 4 %)—and highlight three priority research directions: (1) creation of open multi-depth benchmark datasets, (2) deployment of TinyML at the sensor level, and (3) development of explainable LLM architectures for agronomic decision support.
After implementing these revisions (inclusion of the full PRISMA protocol, normalization of performance metrics, and expansion of the LLM analysis with computational and reproducibility details), the manuscript will achieve the methodological rigor and quantitative transparency required for publication in Sensors.
Comments on the Quality of English Language
The manuscript is verbose and would benefit from reduction of redundancy (average sentence length ~38 words). English usage is acceptable but requires professional editing for conciseness and precision.
Author Response
We warmly thank the Editor and the Reviewers for their useful and thought-provoking comments. We did our best to comply with all the raised concerns, as detailed in the following. In the revised version, the main changes and additions are highlighted in blue, to ease the tracking of the changes, while minor variations and corrections of typos have not been tracked. In the following, we provide detailed responses to the specific issues raised by the Reviewers.
Please see the file attached.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsThis paper provides a systematic review of recent advancements in AI-based soil moisture (SM) forecasting, conducted in accordance with the PRISMA guidelines. The study adopts a structured literature search, a screening flowchart, well-defined inclusion and exclusion criteria, and clearly formulated research questions (RQs). A total of 189 studies were examined, from which 68 qualified papers published between 2017 and 2025 were synthesized, offering a comprehensive and up-to-date overview of AI-driven SM forecasting. The six RQs—addressing sensors, ML/DL models, hybrid frameworks, paradigms, and key challenges—are appropriately designed and logically structured.
The following comments are provided for the authors’ consideration:
1. Several sections are predominantly descriptive rather than analytical. The discussion could be strengthened by providing more explicit comparisons of model performance, methodological limitations, and practical constraints.
2. If sufficient comparable metrics are available in previous studies, incorporating quantitative comparisons or meta-analytical synthesis (e.g., trends in accuracy or RMSE) would significantly enhance the paper’s scholarly contribution.
3. The full names of the abbreviations for Sensor Type should be defined upon their first occurrence.
4. In Tables 2 and 5, some references do not specify the Sensor Type and are marked as “NR.” Please clarify why these references were included in the reviewed literature despite lacking this information.
Author Response
We warmly thank the Editor and the Reviewers for their useful and thought-provoking comments. We did our best to comply with all the raised concerns, as detailed in the following. In the revised version, the main changes and additions are highlighted in blue, to ease the tracking of the changes, while minor variations and corrections of typos have not been tracked. In the following, we provide detailed responses to the specific issues raised by the Reviewers.
Please see the file attached.
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsAlthough the authors did not address all my concerns, the improvements introduced are sufficient for a publication
Reviewer 2 Report
Comments and Suggestions for AuthorsAll comments and recommendations were taken into account by the authors and corrected.
Comments on the Quality of English LanguageThe manuscript is verbose and would benefit from reduction of redundancy (average sentence length ~38 words). English usage is acceptable but requires professional editing for conciseness and precision.

