Review Reports
- Jin-Ye Guan 1,†,
- Xing Zou 1,† and
- Dan Deng 1,*
- et al.
Reviewer 1: Lalit Kumar Reviewer 2: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors compare quantitative evaluation methods with conventional scar scale analysis, focusing on pediatric pathological scars. Scar characteristics were assessed using the Vancouver Scar Scale (VSS). Quantitative assessment was performed using dermoscopy and the Antera 3D® system. The topic falls within the scope of Biomedicines; however, the following comments and concerns should be addressed:
- The authors claim that 3D cameras and dermoscopy are relatively objective methods that provide unbiased and accurate evaluation. This statement is overstated. No assessment method is entirely unbiased, and imaging-based techniques are also subject to technical and operator-related variability.
- In this study, 12 patients were included in the treatment subgroup. This sample size is very small for inferential statistical analysis, raising concerns about statistical power.
- The study includes 36 scar regions derived from 12 patients. These observations may not be independent, as multiple scars from the same patient introduce clustering. This should be accounted for in the statistical analysis.
- A paired Student’s t-test was used for statistical comparison. However, if the VSS is treated as an ordinal scale, a nonparametric test such as the Wilcoxon signed-rank test may be more appropriate.
- The manuscript states that “Dermoscopy presented the lowest variability, indicating limited discriminative power.” This interpretation is questionable and requires clarification. Low variability may reflect high measurement consistency or a homogeneous baseline sample rather than limited discriminative ability.
- The interpretation of the standardized response mean (SRM) requires further justification. Claims such as “outperforming traditional scales” based solely on SRM values should be reconsidered or supported with stronger comparative statistical evidence.
- For a meaningful comparison of evaluation tools, validation against a reference or gold standard is necessary. The absence of such validation limits the ability to conclude superiority or accuracy.
- Study limitations are not clearly discussed and should be explicitly acknowledged.
- The manuscript would benefit from careful grammatical revision and professional language editing.
Author Response
Comments 1:
The authors claim that 3D cameras and dermoscopy are relatively objective methods that provide unbiased and accurate evaluation. This statement is overstated. No assessment method is entirely unbiased, and imaging-based techniques are also subject to technical and operator-related variability.
Response 1:
Thank you for your comments and we fully agree with you that no technique is absolutely unbiased. This might be caused by our improper expression. What we originally intended to express is that scar scale evaluation is judged by physician or patient that is relatively subjected without measurable parameters. In contrast, 3D cameras and dermoscopy are relatively objective when comparing the scar scale evaluation. According to your suggestions, we have made the revision.
(line 18-20)
The original sentence:
“three-dimensional (3D) camera and dermoscopy are unbiased and accurate methods that can provide unbiased and accurate evaluation data”
has been revised to:
“When comparing different scar scale evaluation methods, “three-dimensional (3D) camera and dermoscopy may provide relatively objective measurable parameters to avoid possible subjective bias created by the observers.”
Comments 2:
In this study, 12 patients were included in the treatment subgroup. This sample size is very small for inferential statistical analysis, raising concerns about statistical power.
Response 2:
Thank you for this important comment. We are sorry to find that that the treatment subgroup actually included 18 patients rather than 12, and this has now been corrected throughout the manuscript. (line 23\128\207\226\233). Should the original data be required, please do not hesitate to contact us.
We agree that the sample size remains relatively limited and admitted this might be one shortcoming for future improvement. Technically, recurring a large sample of pediatric scar patient wound be difficult, because this needs to recruit eligible patients and the longitudinal follow-up, which requires children and their patients’ cooperation for a long time.
Although realizing this shortcoming, we tried to publish this study because:
(1) It is extremely important to inform pediatric doctors that quantitative evaluation is the key to predict prognosis and achieve desired therapeutic outcome as we experienced in our clinical work.
(2) There are quite some reports that have relative small sample size and involve imaging-based quantitative evaluation of hypertrophic scars in adult scar evaluation, please see the references.
1.Tawfik A A, Ali R A. Evaluation of botulinum toxin type A for treating post burn hypertrophic scars and keloid in children: An intra-patient randomized controlled study[J]. Journal of Cosmetic Dermatology, 2023, 22(4): 1256-1260.
2.Bray R, Forrester K, Leonard C, et al. Laser Doppler imaging of burn scars: a comparison of wavelength and scanning methods[J]. Burns: Journal of the International Society for Burn Injuries, 2003, 29(3): 199-206.
(3) To overcome this issue, we have actually adopted a way of multiple analysis in the scar of the same patient, thus creating 36 tested scar regions from 18 patients, and therefore the sample size was increased to 36. The rationale behind this is that we observed heterogeneous characters at different regions of the same scar, multiple measurements actually can better represent what has changed in the whole scar before and after the treatment.
Even so, we did acknowledge that the relatively small sample size represents a limitation of the study. This limitation has been clarified in the Discussion section (see below), and future studies with larger cohorts are warranted to further validate these findings.
(line 307-308)
“In terms of study design, limitations include the relatively small sample size.”
(line 310-311)
Additional sentence has been included:
“Future studies should include larger and more diverse patient cohorts to further validate the performance of different scar assessment tools.”
Comments 3:
The study includes 36 scar regions derived from 12 patients. These observations may not be independent, as multiple scars from the same patient introduce clustering. This should be accounted for in the statistical analysis.
Response 3:
Thank you for your valuable suggestion that will help us greatly in our future study. Your points are completely right in principle. In the case of this study, we did find quite some variation in a single patient with scars at different locations, or different regions of the same scar.
The lesion-level analysis (n = 36) was retained because some patients presented with relatively large scar areas involving multiple anatomical regions. A single testing window for 3D or dermoscopy may not be able to reflect the actual scar characters of an individual patient , in another word, this type of measurement would possibly create bias of the data. In particular, scar characteristics and treatment responses may vary across different parts of the same scar, and clinical management often requires region-specific treatment approaches. In this context, multiple tests on the same scar would better include the varied scar characters to provide more balanced data. This is the rationale of the study design. Your understanding will be greatly appreciated. Similar approaches have been reported in previous dermatological studies, where multiple lesions from the same patient were analyzed at the lesion level to capture lesion-specific characteristics.
1. Van Nguyen L, Ly H Q, Vo H T, et al. Clinical Features and the Outcome Evaluations of Keloid and Hypertrophic Scar Treatment with Triamcinolone Injection in Mekong Delta, Vietnam – A Cross-Sectional Study[J]. Clinical, Cosmetic and Investigational Dermatology, 2023, 16: 3341-3348.
2. Yu A, Yick K L, Ng S P, et al. A study of using a simple 2D image analysis method to monitor the surface area of hypertrophic scars on hand during pressure therapy[J]. Burns, 2020, 46(7): 1548-1555.
3. Cavalié M, Sillard L, Montaudié H, et al. Treatment of keloids with laser-assisted topical steroid delivery: a retrospective study of 23 cases[J]. Dermatologic Therapy, 2015, 28(2): 74-78.
Comments 4:
A paired Student’s t-test was used for statistical comparison. However, if the VSS is treated as an ordinal scale, a nonparametric test such as the Wilcoxon signed-rank test may be more appropriate.
Response 4:
We thank the reviewer for this important suggestion. We agree that the Vancouver Scar Scale (VSS) represents an ordinal scale and that a nonparametric test may therefore be more appropriate for statistical comparison. In response to this comment, we additionally performed the Wilcoxon signed-rank test for paired comparisons of VSS scores before and after treatment. The results were consistent with those obtained using the paired-sample t tests, and the statistical significance remained unchanged. These additional analyses have been provided in the table S1.
We have revised the manuscript in the following sections:
(line 30)
The original sentence:
“paired student t- tests and standardized response means (SRMs)”
has been revised to:
“paired-sample t tests (one-tailed), Wilcoxon signed-rank test and standardized response means (SRMs)”
(line 177-180)
The original sentence:
“For the longitudinal treatment‒response analysis, paired-sample t tests (one-tailed) were conducted to compare pre- and post-treatment values obtained by the three measurement tools (VSS, dermoscopy, and Antera 3D®).”
has been revised to:
“For the longitudinal treatment–response analysis, paired-sample t tests (one-tailed) were conducted to compare pre- and post-treatment values obtained by the three measurement tools (VSS, dermoscopy, and Antera 3D®), and the Wilcoxon signed-rank test was additionally performed for VSS to provide a nonparametric assessment.”
(line 214-215)
Additional concluding sentence has also been included:
“Similar results were obtained using Wilcoxon signed-rank tests. See Table S1.”
Comments 5:
The manuscript states that “Dermoscopy presented the lowest variability, indicating limited discriminative power.” This interpretation is questionable and requires clarification. Low variability may reflect high measurement consistency or a homogeneous baseline sample rather than limited discriminative ability.
Response 5:
We thank the reviewer for this insightful comment and agree that the interpretation of variability requires careful consideration. We acknowledge that lower variability does not necessarily indicate limited discriminative ability. As the reviewer pointed out, it may also reflect factors such as high measurement consistency of the technique or relatively homogeneous baseline characteristics of the scars included in the study. In our study, the coefficient of variation (CV) was used to explore the ability of different tools to capture inter-individual differences. We observed that Antera 3D® showed relatively higher variability in several parameters, which may reflect its greater sensitivity in detecting subtle differences in scar pigmentation and volume. However, we recognize that CV values may be influenced by multiple factors beyond the intrinsic characteristics of the scars themselves, including measurement stability and sample heterogeneity. At present, it is difficult to fully control these potential sources of variability. To address the reviewer’s concern, we have clarified that the lower variability observed in dermoscopic parameters may reflect several factors, including measurement stability or limited heterogeneity within the sample. We have also expanded the discussion to acknowledge that CV-based comparisons may be influenced by multiple variables and that further studies with larger cohorts and more controlled conditions are needed to better evaluate the discriminative performance of different assessment tools. We indeed appreciate your insightful comments.
We have revised the manuscript in the following section:
(line 302-307)
Additional sentence has been included:
“Besides, using the CV values as an indicator of the sensitivity of scar assessment tools also has certain limitations. CV values may be influenced by multiple factors beyond the intrinsic characteristics of the scars themselves, including measurement stability and sample heterogeneity. Therefore, further studies are required to more rigorously evaluate the performance of different assessment methods.”
Comments 6:
The interpretation of the standardized response mean (SRM) requires further justification. Claims such as “outperforming traditional scales” based solely on SRM values should be reconsidered or supported with stronger comparative statistical evidence.
Response 6:
We thank the reviewer for this insightful comment. We agree that the superiority of an assessment tool should not be inferred solely based on SRM values. Accordingly, we have revised the wording in the manuscript to avoid potential overinterpretation.
In addition, the use of the standardized response mean (SRM) to evaluate responsiveness to treatment has been reported in some clinical outcome measurement studies. Previous scar assessment studies have also used SRM to evaluate the responsiveness of instruments such as POSAS to longitudinal scar changes (see below). Therefore, in the present study, SRM was used as an exploratory metric to compare the responsiveness of different scar assessment methods.
1. Choo A M H, Ong Y S, Issa F. Scar Assessment Tools: How Do They Compare? [J]. Frontiers in Surgery, 2021, 8.
2. Kučinskaitė A, Stundys D, Gervickaitė S, et al. Aesthetic Evaluation of Facial Scars in Patients Undergoing Surgery for Basal Cell Carcinoma: A Prospective Longitudinal Pilot Study and Validation of POSAS 2.0 in the Lithuanian Language[J]. Cancers, 2024, 16(11): 2091.
Nevertheless, we acknowledge that SRM alone cannot fully determine the overall performance of an assessment tool. To address this limitation, we have added a statement in the Discussion section emphasizing that additional metrics and larger studies are needed to more comprehensively evaluate the sensitivity and discriminative ability of different scar assessment tools.
We have revised the manuscript in the following sections:
(line 184-185)
Additional sentence has been included:
“Tools with higher SRM values were interpreted as being more sensitive to treatment-induced changes.”
(line 297-302)
Additional sentence has been included:
“In addition, the comparison among assessment tools in this study was mainly based on p values from paired tests and SRM values. Although these indicators reflect responsiveness to treatment-induced changes, they may not fully capture the overall performance of different assessment methods. Future studies incorporating larger cohorts and more comprehensive validation approaches are needed to further evaluate the sensitivity and reliability of scar assessment tools.”
In addition, we removed the phrase of “outperforming traditional scales” in the text according to your suggestion.
(line 43-46)
The original sentence:
“In conclusion, Antera 3D® provides an objective, sensitive, and spatially precise as-sessment of PPS, outperforming traditional scales in capturing subtle and early changes.”
has been revised to:
“In conclusion, Antera 3D® offers an objective, sensitive, and spatially precise approach for PPS assessment, and may provide additional quantitative information for evaluating subtle and early changes alongside traditional scar assessment scales.”
Comments 7:
For a meaningful comparison of evaluation tools, validation against a reference or gold standard is necessary. The absence of such validation limits the ability to conclude superiority or accuracy.
Response 7:
We thank the reviewer for this valuable comment. We fully agree that validation against a reference or gold standard would be important for a meaningful comparison of different evaluation tools. However, at present, there is no universally accepted gold standard for quantitative scar assessment, particularly in pediatric populations. One of the aims of our study was to explore the characteristics of several commonly used evaluation methods and provide preliminary evidence that may contribute to the identification of a more reliable assessment framework. We acknowledge that establishing a true reference standard requires larger cohorts and more comprehensive validation. Accordingly, we have clarified in the revised Discussion section that future studies with expanded sample sizes will be necessary to further validate these findings and move toward the development of a more robust reference standard for scar evaluation.
We have revised the manuscript in the following sections:
(line 310-315)
Additional sentence has been included:
“Future studies should include larger and more diverse patient cohorts to further validate the performance of different scar assessment tools. In addition, longitudinal studies with standardized imaging protocols and multicenter collaboration may help reduce measurement variability and improve the comparability of results. Integrating objective imaging techniques with clinical evaluation scales may also provide a more comprehensive framework for scar assessment”
(line 317-320)
Additional sentence has been included:
“Ultimately, such efforts may contribute to the establishment of a more reliable refer-ence standard for scar evaluation, particularly in pediatric populations where stand-ardized assessment methods remain limited.”
Comments 8:
Study limitations are not clearly discussed and should be explicitly acknowledged.
Response 8:
We thank the reviewer for this helpful suggestion. We agree that the limitations of the study should be more clearly and explicitly discussed. In the revised manuscript, we have expanded the Discussion section to more clearly acknowledge several limitations of the present study, including the relatively small sample size, the potential lack of independence among scar regions derived from the same patient, the limitations of using statistical indicators such as SRM to compare evaluation tools, and the absence of a universally accepted reference standard for scar assessment. These points have now been explicitly discussed in the revised manuscript to provide a more balanced interpretation of the findings and to outline directions for future research.
(line 297-307)
Additional sentence has been included:
“In addition, the comparison among assessment tools in this study was mainly based on p values from paired tests and SRM values. Although these indicators reflect responsiveness to treatment-induced changes, they may not fully capture the overall performance of different assessment methods. Future studies incorporating larger cohorts and more comprehensive validation approaches are needed to further evaluate the sensitivity and reliability of scar assessment tools. Besides, using the CV values as an indicator of the sensitivity of scar assessment tools also has certain limitations. CV values may be influenced by multiple factors beyond the intrinsic characteristics of the scars themselves, including measurement stability and sample heterogeneity. There-fore, further studies are required to more rigorously evaluate the performance of different assessment methods.”
Comments 9:
The manuscript would benefit from careful grammatical revision and professional language editing.
Response 9:
We thank the reviewer for this helpful suggestion. The manuscript has been carefully revised for grammar, clarity, and overall language quality. In addition, the text has undergone professional language editing to improve readability and ensure that the manuscript meets the linguistic standards of the journal.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors described "Comparison of quantitative evaluation and conventional scar
scale analysis for pediatric pathological scars". As they mentioned, there is no definitive method to assess the severity of scars, both subjective and objective. So, this topic should be informative and attractive for potential readers. I have some suggestions and questions to improve this manuscript.
- More figure legends should be added especially in Figure1. What are L *,a* and Vulumn? Where did you evaluate on the scars? Why did you choose the sites of scars?
- How did you determine the sample size? Thirty five should be small.
- Why did you choose "Children"? Why wasn't it acceptable for adults?
Author Response
Comments 1:
More figure legends should be added especially in Figure1. What are L *,a* and Vulumn? Where did you evaluate on the scars? Why did you choose the sites of scars?
Response 1:
We thank the reviewer for this helpful and constructive comment. We agree that the explanation of the parameters and the measurement locations in Figure 1 should be clearer. In the revised manuscript, we have expanded the legend of Figure 1 to clarify which parameters correspond to each assessment tool. In addition, more detailed explanations of the parameters L*, a*, and volumn have been added to the Methods section to improve clarity and ensure that the measurement indicators are clearly defined.
Furthermore, we have clarified the criteria used for selecting the scar regions. In general, representative areas that reflect the main characteristics of the scar were selected for evaluation, such as regions showing prominent pigmentation, erythema, or elevation. These areas usually correspond to the most clinically relevant parts of the scar and are often the regions requiring treatment. In addition, to provide a more balanced data, the selection of the scar regions are generally the reflection of diverse characters at the different regions, and avoid choosing similar scar characters, so as to avoid induced clustering data of the same patient. We have clarified in the revised Discussion section
We have revised the manuscript in the following sections:
(line 130-137)
Additional sentence has been included:
“For imaging analysis, scars requiring treatment were preferentially selected, particularly those presenting with prominent pigmentation, vascularity, or elevation. Generally, the selected areas also need to reflect the diverse characters at the different regions, so as to avoid induced clustering data of the same patient. Measurements were primarily obtained from representative regions within the central area of the scar, which were considered to best reflect typical scar characteristics while avoiding the boundary between scar tissue and surrounding normal skin.”
(line 156-158)
Additional sentence has been included:
“The a* value and green value reflect the degree of vascularity, whereas the L* value reflects the level of pigmentation within the scar.”
(line 191-193)
Additional sentence has been included:
“( VSS: Pigmentation, Vascularity, Height, Pliability; Dermoscopy: Green value, a* value, L* value; Antera 3D®: Pigmentation, Vascularity, Volume)”
Comments 2:
How did you determine the sample size? Thirty-five should be small.
Response 2:
We thank the reviewer for this important comment. We agree that the sample size remains relatively limited. As responding to another reviewer for the same concern, recruiting pediatric patients for time consuming and tedious evaluation, which need full cooperation of children and patients, would be difficult. Meanwhile, we also feel that proper scar evaluation plays an extremely important role in predicting prognosis and guide effective treatment for pediatric scar patients. Thus, we wish that we can publish this preliminary study, although it may not be perfect in the current sample size, and our future cohort study with a large sample size may offset the current limitations.
In addition, we did find that studies involving imaging-based quantitative evaluation of hypertrophic scars often involve relatively small cohorts in adult patients, also likely due to the difficulty of in recruiting eligible patients and the need for longitudinal follow-up as showed in the following references:
1. Tawfik A A, Ali R A. Evaluation of botulinum toxin type A for treating post burn hypertrophic scars and keloid in children: An intra-patient randomized controlled study[J]. Journal of Cosmetic Dermatology, 2023, 22(4): 1256-1260.
2. Bray R, Forrester K, Leonard C, et al. Laser Doppler imaging of burn scars: a comparison of wavelength and scanning methods[J]. Burns: Journal of the International Society for Burn Injuries, 2003, 29(3): 199-206.
Nevertheless, we acknowledge that the relatively small sample size represents a limitation of the study. This limitation has been clarified in the Discussion section, and future studies with larger cohorts are warranted to further validate these findings.
(line 307-308)
“In terms of study design, limitations include the relatively small sample size.”
(line 310-311)
Additional sentence has been included:
“Future studies should include larger and more diverse patient cohorts to further validate the performance of different scar assessment tools.”
Comments 3:
Why did you choose "Children"? Why wasn't it acceptable for adults?
Response 3:
We thank the reviewer for this insightful comment. We chose to focus on pediatric patients for several reasons. Firstly, although scar assessment has been widely studied in adult populations, studies specifically addressing scar evaluation in children remain relatively limited. Our study was therefore designed in part to address this gap and to contribute evidence on the evaluation of scars in pediatric patients. Secondly, our institution is a specialized pediatric medical center where the majority of patients are children. As a result, pediatric scar management represents a major component of our clinical practice, and focusing on this population allows us to leverage our clinical expertise and patient resources. Finally, although the present study focuses on children, the evaluation tools investigated in this study are not inherently age-specific. Therefore, the main conclusions regarding the comparative performance of these assessment methods may also be informative for scar evaluation in adult populations.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsNIL
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors revised the manuscript precisely. Thank you for this opportunity.