Review Reports
- Christopher M. Ardohain 1,2,
- Dennis H. Choi 1,3 and
- Songlin Fei 1,*
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsGeneral Comments
This review is valuable and looks at many kinds of sensors, forest types, and inventory tasks. It is well organized and cites a lot of literature. But several important problems must be fixed before it can be published.
Major Revisions Needed
1. Present a literature search methodology.
There is no information in the manuscript about how the 163 studies were selected. It is essential information that readers require in order to determine if the search was thorough and fair.
Suggestion: Include a short section on “Search Strategy”. Mention the databases searched, keywords employed and time period covered, as well as inclusion/exclusion criteria. It is strongly recommended that a PRISMA flow diagram be used to describe the way the literature was searched. If it has not been done in a systematic manner, describe it as a limitation.
2. Examine the validity of the study results.
Review the results of the study's validity. All the cited works are considered equally in the review, but some may be based on small numbers, have not been validated externally or have reported only positive results.
Suggestion: Include short comments on the quality of the study. Mark the studies that were based on independent data from a test versus those that were made using internal validation.
3. Make sure to update keywords to make sure they don't get repeated.
The words deep learning, remote sensing and forest inventory are already in the title. The repetition of them as keywords can hinder discoverability.
Suggestion: Use more specific words than overlapping words, e.g., point cloud deep learning, individual tree detection, segmentation of tree crowns, and estimation of above-ground biomass. Keep the terms “LiDAR” and “data fusion.”
4. Reorganize the abstract
The abstract does not mention the literature search strategy, nor does it contain the quantitative results; it contains qualitative results.
Suggestion: Rewrite the abstract to clearly present: (1) background and objectives, (2) your search/selection methodology, (3) key quantitative results (e.g., number of papers analyzed), and (4) your main conclusions.
5. Avoid overstating facts or ideas
There is too much hyperbole or when there are interpretations, they are presented as facts. Use softer language. Note that most of the inventory programs are being implemented traditionally, and that deep learning is not used.
The assertion that deep learning is already changing the face of everyday forest inventories is still too early since the technology is not known to have been widely adopted in the field. The fact that the annotation of object detection is more common than others is not a fact, but your interpretation. Please, rephrase.
6. Make a comparison to previous reviews.
The introduction provides a list of the previous reviews, but lacks a discussion on how this paper is building on or diverging from these reviews.
Include a paragraph in the Introduction to compare results with the most relevant prior reviews. Describe what this review is changing, confirming or updating.
7. Verify 2026 references
Ensure that these papers published in 2026 are available. If they are still in press, edit the citations to include "In Press".
Minor Corrections
Grammar & Typos:
Line 60 (Keywords): Change the colon in “artificial intelligence:” to a comma or semicolon.
Line 198: Correct “a marque of transformers” to “a marquee feature of transformers.”
Line 780: Correct the section numbering from “4.1.2. LiDAR” to “4.2.2. LiDAR.”
Line 865: Change “ad reached” to “and reached.”
Line 1412: Change “has began” to “has begun.” sentence with the correct spelling.
Acronyms:
Define all acronyms at first use, then use them throughout the paper.
Recommendation: Reconsider after major revision.
The paper is timely and well-scoped, but resolving the issues with the methodology, study bias, abstract structure, and terminology is necessary before publication
Comments for author File:
Comments.pdf
The overall English-language quality is acceptable but would benefit from a professional proofread.
Author Response
Please see the attachment
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsThe author systematically reviews the application of deep learning in remote sensing-based forest resource surveys, focusing on three core tasks: tree counting and localization, tree species identification, and tree measurement. The manuscript seems original, and the authors conducted a detailed analysis in this regard. However, before being accepted, I suggest the following modifications:
1.The authors note that this paper is a review that is not limited by the type of data acquisition platform, covering a broader range of applications and exploring various remote sensing data products. However, the text appears to focus primarily on optical and LiDAR data, without addressing synthetic aperture radar (SAR). In fields such as forest inventory and biomass estimation, SAR is an indispensable core data source. The multimodal fusion of optical, LiDAR, and synthetic aperture radar (SAR) data is also an important research direction. Therefore, the following recommendations are suggested:
1) In the introduction or methods section, clearly state the scope of this review: focusing on optical and LiDAR data
2) In the “Future Directions” or “Challenges” section, add a paragraph discussing the potential of SAR
3) Revise the claims in the abstract or conclusions: Avoid using absolute statements such as “not limited by data type”
2.Line136, “Deep learning (or artificial intelligence) is a subset of machine learning…..” this should be corrected, because deep learning is not the same as artificial intelligence.
- There is an issue where abbreviations are used without prior definition, such as "FPS" in line 244, “BEV” in line 373.
4). The citation format is inconsistent.
1) Different spellings in the same journal, such as “Remote Sensing of Environment ”vs. “Remote sensing of environment”.
2) Mixed use of abbreviations and full forms.
3) Some references lack a DOI or page numbers.
Author Response
Please see the attachment
Author Response File:
Author Response.docx
Reviewer 3 Report
Comments and Suggestions for AuthorsThis review provides a comprehensive systematic review of deep learning applications in three core tasks of forest inventory: tree counting and localization, tree species identification, and tree attribute measurement. The authors synthesize existing research across multiple commonly used remote sensing platforms (including terrestrial, UAV, airborne, and satellite systems) and different data modalities (LiDAR, optical imagery, and fused multi-source data), highlighting key open challenges facing the current research community, such as the scarcity of high-quality annotated reference data and the limited cross-scenario transferability of deep learning models, and proposed a structured perspective on future research directions. Nevertheless, current version requires substantial revisions to improve its presentation clarity, analytical depth, and critical assessment of the current research progress.
- Although the manuscript is presented as a review article, it does not provide a detailed description of a reproducible methodology for literature selection. As a result, readers are unable to evaluate whether the included studies are representative of the existing research body, or whether selection bias may have been introduced during the screening process. Please add an independent "Methodology" section (typically placed after the Introduction) that explicitly clarifies the following information: (a) the academic databases used for literature search (e.g., Web of Science, Scopus, Google Scholar), (b) the search strings and keywords adopted in the retrieval process, (c) the inclusion and exclusion criteria applied for study screening (e.g., restriction to peer-reviewed publications, language requirements, specified publication date range), and (d) the number of papers retrieved at initial search, screened for eligibility, and finally included in the review (a PRISMA flow diagram is recommended to visualize this process). This requirement follows the standard practice for high-quality review papers in the field of remote sensing.
- The depth of analysis varies substantially across different sections of this review. For instance, in Section 3.3 focused on urban forests, the authors claim that they “did not find any study that combines multiple data sources into a fused product for individual tree localization in urban environments” (Lines 457–458). However, the authors themselves subsequently cite the work of Li et al. (2024), which uses MLS data, in Section 3.3.2 (Line 440), along with the study by Guo et al. (2019) within the same section. Moreover, the existing reference list also includes fusion-relevant studies such as Ottoy et al. (2025) and Wang et al. (2026), which further contradicts the afore mentioned statement that no relevant fusion studies were identified. This inconsistency requires careful re-examination and clarification by the authors.
- Bark-based species identification (Lines 587-600): This work addresses an interesting and specialized research niche, but its integration into the overall narrative flow of this review is currently poor. A key question arises regarding its current positioning: why is this content placed within the "Research Trends" subsection of the general Species Identification section, rather than as an independent subsection under Section 4.2 (Temperate Forests), where the vast majority of existing bark-based identification studies are conducted? Restructure the species identification section to improve structural consistency across the manuscript. This can be achieved in one of two ways: either reclassify bark-based identification research as a dedicated standalone subsection (4.2.4) under Temperate Forests, or provide a clearer, more robust justification for its current placement. Additionally, please verify that claims regarding the lack of fusion-oriented bark studies in specific forest types are factually accurate, and if the claimed gap is confirmed, include a discussion on the potential factors driving this underrepresentation in the literature.
- Most existing narrative sections currently present a sequential listing of individual study summaries, rather than a rigorous critical synthesis of existing work. While this structure clearly communicates “what” each prior investigation accomplished, it rarely addresses why specific methodological approaches outperform alternatives, or articulates the broader implications of these comparative outcomes for the research field. We recommend expanding the analytical depth by incorporating additional meta-analytic synthesis. For each thematic subsection, a dedicated paragraph that integrates cross-study findings should be added. For instance, after concluding the summary of LiDAR-based individual tree detection research in natural forests (Section 3.2.2), you could add a comparative summary table or a critical synthesis paragraph structured along these lines:"Across the studies reviewed, anchor-free detection architectures such as CenterNet (Xi & Hopkinson) demonstrate strong performance for individual tree crown delineation, but remain highly sensitive to variation in stand stem density. In contrast, region-based detectors including Faster R-CNN achieve superior accuracy in low-density stands, but suffer substantial performance degradation under conditions of high canopy occlusion. No one detection architecture universally outperforms all alternatives across all site conditions, indicating that method selection should be guided by the structural complexity of the target forest and the point density of the input point cloud."
- Although the manuscript discusses multiple accuracy metrics including F1-score, recall, precision, and RMSE, it lacks a systematic, high-level quantitative synthesis of existing results. To date, it remains unclear how typical F1-scores for tree counting tasks differ across stand types—for example, whether performance reaches above 0.95 in homogeneous plantation stands while falling in the range of 0.60 to 0.80 in heterogeneous natural forests. Adding a summary table or figure of published results would significantly strengthen the contribution of this work. Develop a summary table (e.g., Table 2) that stratifies existing studies according to five core attributes: (a) the specific computer vision task addressed, (b) the forest type where the method was tested, (c) the sensor modality used to collect data, (d) the deep learning architecture implemented, and (e) the reported accuracy (presented as both the mean value and the range of reported values across test cases). Following this table, add a paragraph in the Discussion section that interprets the sources and implications of the observed performance variation across the categorized study features.
- Line 78: "the measurement at individual tree level may be limited" Change to "measurement at the individual tree level is often limited".
- Line 116 (Figure 1): The figure resolution is insufficient, resulting in poor overall quality. Specifically, the text labels within the Venn diagram are illegible. Please replace the current figure with a high-resolution vector graphic. In addition, the categorization rationale underlying the grouping "33 other review papers" is not clearly presented. We ask that you clarify how these 33 papers were selected in the main text.
- Lines 174-175: "YOLO (You Only Look Once), Faster Region-Based Convolutional Neural Network (Faster R-CNN), and SSD..." – Define YOLO and SSD after first use. Already done, but ensure consistency.
- Line 299, 305: "F-1 scores" Change to "F1-scores" for consistency.
- Lines 323-325: "The authors achieved an R² score of 0.93 on their held-out independent test set... This performance represents a substantially higher accuracy than typically reported for individual-tree localization tasks relying on optical imagery from a single data source, particularly for imagery at the relatively coarse 0.8 m spatial resolution used in this work." A growing body of literature confirms that R² scores between 0.65 and 0.85 are the most common range for single-source optical imagery-based tree localization at comparable spatial resolutions. For example, in a study of individual tree position estimation using 1 m resolution Sentinel-2 optical imagery, reported an R2 of 0.72 for mixed temperate forests, which is over 20% lower than the 0.93 obtained in the work under review. This comparison contextualizes the strong performance reported here.
- Lines 621-622: "remainder are based solely on optical imagery, and we found a noticeable lack of studies that use only LiDAR in tropical and subtropical forests." – This is an interesting finding. Move this statement from the end of the introductory paragraph (4.1) to a dedicated "Research Gaps" paragraph within Section 4.1.
- Line 1122: "Accuracies metrics of minimum, maximum, and average relative height error were 1.92%, 4.87%, and 3.2%..." – This phrasing is awkward. It could change: "The relative height errors ranged from 1.92% (minimum) to 4.87% (maximum), with an average of 3.2%."
- Section 5.3, this section is overly brief (spanning Lines 1316–1328) and fails to present a coherent, comprehensive synthesis of existing research trends. We recommend one of two revisions: either integrate this section into the main Discussion chapter, or expand it to explicitly elaborate on two to three well-defined, prominent trends in the field. Illustrative examples of such trends include: the ongoing paradigm shift from hand-engineered, handcrafted features to data-driven learned feature representations for tree characterization; the rapidly growing adoption of multi-modal sensing data for comprehensive tree measurement; and the persistent performance gap between research-scale prototype methods and operational-scale deployment-ready approaches.
- Line 1384 (Table 1): The table appears to be missing. It is referenced in the Conclusions. This is a critical omission. Provide the table or remove the reference.
- Lines 1412-1413: "began to transform" Change to "is transforming" or "has begun to transform".
- Appendix A (Table A1): The content of this table is informative but poorly formatted for readability. It is recommended to reformat the table to landscape orientation, and consider splitting it into two separate tables (for example, grouped by forest type). Multiple text entries in the current table are truncated (e.g., the reference "Carpen- tier et al. 2018" has an erroneous line break splitting the author name). Please conduct a careful full proofreading of all table content to resolve such issues.
Author Response
Please see the attachment
Author Response File:
Author Response.docx
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThank you for your hard work in revising this manuscript. You have addressed the previous comments carefully and thoughtfully.
Adding the new search methodology section and the tables highlighting the sample sizes and limitations of each study makes this review much more reliable for readers. The updated keywords and the rewritten abstract also make the paper easier to find and read. All of the minor typos, numbering issues, and grammatical errors have been successfully corrected.
The paper is now in a very strong position, and I am pleased to support its publication.
Comments for author File:
Comments.pdf
Author Response
Thank you for all the comments and suggestions which made the paper stronger!
Reviewer 3 Report
Comments and Suggestions for AuthorsThe revisions are improved and many concerns have been addressed appropriately. However, some issues need further clarification.
- For comment 2:The authors have clarified "explicit deep learning-based optical-LiDAR fusion for urban individual tree localization." However, the distinction remains somewhat subtle. Please add a brief explanatory note stating that while LiDAR-only and optical-only urban studies are common, studies that fuse the two modalities specifically for individual tree localization remain scarce, and that this represents an identified research gap.
For comment 4:The synthesis paragraphs added to the close of numerous subsections now strengthen the analytical depth. Yet, the inserted text remains relatively general and does not fully resolve the underlying issues. For instance, the synthesis in Section 4.2.2 (LiDAR-based natural forest detection) would gain from a sharper contrast between anchor-free and region-based architectures, along with an explicit discussion of the specific conditions favoring each. Please revise each synthesis paragraph to deliver clear, comparative, domain-relevant insights instead of providing a broad summary.
For comment 5:Incorporating the performance comparison table (currently Appendix B, Table A5) into the main text as a dedicated table (e.g., Table 2) would significantly enhance the manuscript. This table provides a valuable cross-task, cross-forest-type, and cross-sensor perspective. We recommend promoting it from the appendix and pairing it with a dedicated Discussion Section. Furthermore, the interpretive text in the Conclusion should be expanded to more explicitly analyze the sources of performance variation. For instance, addressing why plantation forests achieve higher F1-scores than natural forests, and how LiDAR-derived improvements in measurement accuracy are contingent on the quality of reference data.
Author Response
For comment 2:The authors have clarified "explicit deep learning-based optical-LiDAR fusion for urban individual tree localization." However, the distinction remains somewhat subtle. Please add a brief explanatory note stating that while LiDAR-only and optical-only urban studies are common, studies that fuse the two modalities specifically for individual tree localization remain scarce, and that this represents an identified research gap.
Response: We have added the brief explanatory note as suggested.
For comment 4:The synthesis paragraphs added to the close of numerous subsections now strengthen the analytical depth. Yet, the inserted text remains relatively general and does not fully resolve the underlying issues. For instance, the synthesis in Section 4.2.2 (LiDAR-based natural forest detection) would gain from a sharper contrast between anchor-free and region-based architectures, along with an explicit discussion of the specific conditions favoring each. Please revise each synthesis paragraph to deliver clear, comparative, domain-relevant insights instead of providing a broad summary.
Response: We revised each synthesis paragraph per the reviewer’s recommendation.
For comment 5:Incorporating the performance comparison table (currently Appendix B, Table A5) into the main text as a dedicated table (e.g., Table 2) would significantly enhance the manuscript. This table provides a valuable cross-task, cross-forest-type, and cross-sensor perspective. We recommend promoting it from the appendix and pairing it with a dedicated Discussion Section. Furthermore, the interpretive text in the Conclusion should be expanded to more explicitly analyze the sources of performance variation. For instance, addressing why plantation forests achieve higher F1-scores than natural forests, and how LiDAR-derived improvements in measurement accuracy are contingent on the quality of reference data.
Response: We moved Table A5 to the new Discussion Section and we expanded the interpretive text in the Conclusion Section as suggested by the reviewer. In addition, we highlighted these points in the Abstract.