Comparative Chloroplast Genome Analysis of Anchusa and the Adulterants of HERBA ANCHUSAE
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsDear Authors,
Your manucscript titled „Comparative Chloroplast Genome Analysis of Anchusa and Its Adulterants” contains valuable results. Nevertheless, I have found some imperfections, which (in my opinion) should be corrected or at least clarified before an eventual publication. I have listed them below:
- The Abstaract section should be improved. It should adress to main parts of manuscript background, methods, results and conclusions. Plase, add information about background of studies and provide main conclusion.
- In my opinion the chapter Introduction shoud start from text placed in lines 55-65. The description of the Uyghur medicinal herb Niushecao (lines 37-64) should be moved and appear after Anchusa azurea Mill. characteristics.
- Figures (especially Figure 1 and Figures 5-7) are hardly legible in current form. Their quality should be corrected. Figure 8, please use the English in axa captions.
- Does Scheme 2 contain the same data as Figure 5?Why is it placed in the Discussion chapter?
- Discussion section contain too small number of references to other publications. In my opinion very valuable part would be the discussion of limitations of presented investigations.
- In my opinion In Conclusion chapter the possible directions of future studies are strongly desired.
Author Response
Response to Reviewer 1
1.The Abstract section should be improved. It should address the main parts of manuscript background, methods, results and conclusions. Please add information about background of studies and provide main conclusion.
Response: We thank the reviewer for this suggestion. We have restructured the Abstract to follow a clear Background→Methods→Results→Conclusions framework. Specifically, we have added background information on the medicinal importance of Anchusa and the practical need for species identification. The main conclusion regarding the three candidate barcodes and the need for broader sampling has been explicitly stated.
Revised Abstract (lines 12-27):
Species within the genus Anchusa L. are valued in Uyghur medicine for their anti-inflammatory and analgesic properties. However, the botanical origin of Niushecao, a critical Uyghur medicinal herb, remains subject to significant confusion within China. Conventional identification methods and standard DNA barcodes often fail to distinguish between these closely related taxa. Although chloroplast genomes harbor variable regions capable of resolving species-level differences, no such comprehensive analysis has previously been conducted for Anchusa. This study sequenced and compared the complete chloroplast genomes of six Anchusa species alongside their primary adulterant. All genomes exhibited a typical quadripartite circular structure, ranging from 150,178 to 150,844 bp in length, and retained a conserved gene set. Despite this structural conservation, several highly variable regions were identified, predominantly within non-coding intergenic spacers. Employing two complementary approaches—sliding window analysis and mVISTA-based sequence visualization—three overlapping regions (rbcL-psaI, petA-psbJ, and trnC-GCA-petN) were pinpointed as particularly promising DNA barcode candidates. Phylogenetic reconstruction revealed that Anchusa strigosa Banks & Sol. is most closely related to the authentic medicinal species Anchusa azurea Mill. (BS = 100%), while other morphological mimics formed distinct clades. These findings provide the inaugural chloroplast genomic resources for this genus and establish practical molecular markers for authenticating Anchusa medicinal materials, thereby enhancing quality control and ensuring the correct utilization of herbal products.
2.In my opinion the chapter Introduction should start from text placed in lines 55-65. The description of the Uyghur medicinal herb Niushecao (lines 37-64) should be moved and appear after Anchusa azurea Mill. characteristics.
Response: We agree with the reviewer's suggestion to improve the logical flow of the Introduction. The revised Introduction now begins with the broader botanical and medicinal background of Anchusa azurea Mill., followed by the specific description of the Uyghur medicinal herb Niushecao. This reordering provides readers with the foundational context before introducing the traditional medicinal applications. The revised structure is: (1) A. azurea distribution and traditional uses; (2) Niushecao in Uyghur medicine; (3) phytochemistry and pharmacology; (4) misidentification problems; and (5) molecular identification rationale.
3.Figures (especially Figure 1 and Figures 5-7) are hardly legible in current form. Their quality should be corrected. Figure 8, please use the English in axis captions.
Response: We have regenerated all figures at higher resolution (300 dpi minimum) to improve legibility. Figure 1 (genome maps) and Figures 5-7 (mVISTA, IR boundary, and collinearity analyses) have been redrawn with larger fonts and clearer color contrast. For Figure 8 (sliding window analysis), we have changed all axis labels and region headers from Chinese to English. The revised version now reads: "Window position (bp)" for the x-axis and "Nucleotide diversity (Pi)" for the y-axis, with region labels such as "LSC," "SSC," and "IR" in English.
4.Does Scheme 2 contain the same data as Figure 5? Why is it placed in the Discussion chapter?
Response: We appreciate the reviewer pointing out this confusion. What was previously labeled as "Scheme 2" was indeed a duplicate mVISTA alignment figure that overlapped with Figure 5. We have removed the duplicate scheme from the Discussion section and kept only Figure 5 (mVISTA alignment) in the Results section. All cross-references have been updated accordingly. We have also unified the figure numbering system so that there are no longer separate "Schemes" and "Figures"—everything is now designated as Figure with sequential numbering.
5.Discussion section contain too small number of references to other publications. In my opinion very valuable part would be the discussion of limitations of presented investigations.
Response: We have significantly expanded the Discussion section by adding more references to related studies, particularly those on chloroplast genome evolution in Boraginaceae and comparative analyses in other medicinal plant genera. We have also added a dedicated subsection on "Limitations and Future Perspectives" that explicitly discusses: (1) the single-sample-per-species limitation and its implications for marker validation; (2) the absence of direct A. azurea sampling; and (3) the incomplete phylogenetic sampling for resolving Anchusa s.l. paraphyly. The following text has been added (lines 475-490):
Several limitations of this study should be acknowledged. First, each species was represented by a single sample, lacking intraspecific replication. Therefore, the diagnostic power of the candidate markers at the species level could not be assessed, and it remains difficult to distinguish intraspecific variation from interspecific divergence. Second, A. azurea was not sampled directly in this study. Thus, the practical discrimination between the authentic medicinal material and its main adulterant, A. strigosa, could not be experimentally verified. Third, the phylogenetic analysis did not include Lycopsis, Cynoglottis, or the subgenera Buglossellum and Buglossoides, which have been segregated from Anchusa s.l. Consequently, the paraphyly of Anchusa s.l. remains untested. Future studies should expand sampling and include intraspecific replicates to validate marker stability. Additional evidence, such as chromosome data, should be incorporated. Nuclear genes should be integrated with chloroplast genome data for combined phylogenetic analyses.
6.In my opinion In Conclusion chapter the possible directions of future studies are strongly desired.
Response: We have revised the Conclusions section to include explicit directions for future research. The following text has been added (lines 530-538):
Future studies should expand sampling and include intraspecific replicates to validate marker stability. Additional evidence, such as chromosome data, should be incorporated. Nuclear genes should be integrated with chloroplast genome data for combined phylogenetic analyses. These efforts will further clarify the infrageneric classification and provide a robust taxonomic foundation for the medicinal and edible use of this genus.
Reviewer 2 Report
Comments and Suggestions for AuthorsI checked your manuscript and described comments below.
This paper provides an excellent comparative analysis of the chloroplast genome of Anchusa, a plant also used in medicinal herbs.
I believe the following points should be considered:
- The paper is missing supplementary files. These are necessary for peer review, so you should upload them.
- This paper uses next-generation sequencing, but I think it would be better to include the primer sequence.
- The amount of DNA in the sample used for this analysis is unknown. DNA quantity is important; it should be included.
- It's not a big deal, but Figure 8 contains Chinese characters. It should be in English.
- The URL placement on line 168 and the duplicate URL on line 183 need to be corrected.
- I believe Ref. 8 is a doctoral dissertation. Since I cannot verify it, I think it would be better to change it to one published in an international scientific journal, if possible.
- Please provide the journal name, volume, and page numbers for the article in Ref. 14.
I don't think this paper has major problems and grammatical problems.
Author Response
Overall Response: We thank the reviewer for the positive evaluation of our comparative analysis and for the constructive suggestions. All points have been addressed.
Point 1: The paper is missing supplementary files. These are necessary for peer review, so you should upload them.
Response: We apologize for this oversight. We have now prepared and uploaded the following supplementary materials:
Supplementary Table S1: Detailed gene annotation information for all six Anchusa chloroplast genomes.
Supplementary Table S2: Complete list of SSRs identified in each species.
Supplymentary Figure S1: Full-resolution chloroplast genome maps for each species.
Supplementary Figure S2: Additional mVISTA alignment including adulterants.
Supplementary Dataset S1: Aligned chloroplast genome sequences.
The supplementary files are now available for the review process.
Point 2: This paper uses next-generation sequencing, but I think it would be better to include the primer sequence.
Response: We appreciate this suggestion. Since this study used whole-genome shotgun sequencing (Illumina platform) rather than PCR-based targeted sequencing, no specific primers were used for chloroplast genome amplification. However, we have now added to the Methods section (lines 150-155) the detailed library preparation and sequencing parameters, including the Illumina platform type, library insert size (~350 bp), and sequencing mode (PE150). For transparency, the following text has been added: "Sequencing libraries were prepared with an average insert size of 350 bp and sequenced on the Illumina NovaSeq 6000 platform in paired-end 150 bp mode."
Point 3: The amount of DNA in the sample used for this analysis is unknown. DNA quantity is important; it should be included.
Response: We agree that DNA quantity is important for data quality assessment. We have now added the DNA quantity information to the Methods section (line 158): "Total genomic DNA was extracted using a modified CTAB method, with approximately 1.5 μg of high-quality DNA (A260/A280 = 1.82-1.85; A260/A230 = 2.0-2.2) used for each library preparation." This information has also been added to Table 1.
Point 4: It's not a big deal, but Figure 8 contains Chinese characters. It should be in English.
Response: We have revised Figure 8 so that all axis labels and region headers are in English. The x-axis now reads "Window position (bp)" and the y-axis reads "Nucleotide diversity (Pi)." Region labels (LSC, SSC, IR) are now also in English.
Point 5: The URL placement on line 168 and the duplicate URL on line 183 need to be corrected.
Response: We have moved all URLs to the Reference section rather than embedding them in the main text. The duplicate URL for IRscope has also been removed. The tools are now cited with their corresponding references (references 33-38, 43, 44) and the URLs are provided in the References.
Point 6: I believe Ref. 8 is a doctoral dissertation. Since I cannot verify it, I think it would be better to change it to one published in an international scientific journal, if possible.
Response: We have replaced the doctoral dissertation (Ref. 8) with a published journal article that covers the same pharmacological findings. The new reference (now Ref. 9) is: Chen, L.; Xu, X.Q.; Guo, W.H.; Mai, D.N. Effects of Anchusa italica Retz. on cough variant asthma and the level of TLR4/NF-κB signaling pathway. Chinese Journal of Clinical Pharmacology 2024, *40*, 698-702.
Point 7: Please provide the journal name, volume, and page numbers for the article in Ref. 14.
Response: We have corrected Ref. 14 by adding the complete citation: Xu, X.Q.; Chen, L.; Zhang, J.; Sun, Y.; Qing, D.G. Study on the Extraction Process of Anti-inflammatory Active Part of Italian Niushecao. Western Journal of Traditional Chinese Medicine 2019, *32*, 17-20.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript assembles and compares chloroplast genomes of six Anchusa species and several Boraginaceae adulterants (Echium vulgare, Borago officinalis, Cynoglossum amabile), with the practical aim of supporting molecular identification of the Uyghur medicinal herb Niushecao. The authors characterize genome structure, repeats, codon usage, IR boundaries, nucleotide diversity, and phylogeny. The topic is useful and the data could be a valuable resource, but the manuscript has several substantive problems with internal consistency, marker evaluation logic, statistical rigor, and presentation quality that must be addressed before it can be considered for publication.
The study is also very descriptive and largely follows a standard plastome-comparison workflow. Its main contribution is the chloroplast genome resource and the candidate hypervariable regions. However, the central applied claim is undercut by the authors' own results, which show that the markers cannot reliably separate the two most relevant closely related species. The manuscript also contains many numerical inconsistencies, mismatched figure, text statements, and editorial errors. I would recommend rejection and encourage of resubmission.
Some most important issues
- The candidate markers do not achieve the study's stated goal, and this is not honestly reconciled. The Introduction frames the work as solving original-plant misidentification of Niushecao, where the key practical problem is distinguishing A. azurea from A. strigosa (both in subg. Buglossum). Yet the Discussion (lines 494–498) states that, except for trnC-GCA-petN and rbcL-psaI, the hypervariable fragments contain no more than three variable sites between A. strigosa and A. azurea, so these two species cannot be reliably separated. The Abstract and Conclusions present the markers as a successful identification solution without making this critical limitation clear. The framing must be revised so the conclusions match the data. The broader samplin (2 or more sample per species) would help a lot.
- The two methods for identifying hypervariable regions give different answers, and this is never explained. The DnaSP sliding-window analysis (Table 5, Section 3.5) yields petN-psbM, rbcL-psaI, petA-psbJ, and ycf1. The mVISTA comparison (Section 3.4) yields a different set: pafI-trnS-GGA, trnC-GCA-petN, psbE-petL, petN-psbM. In the Discussion (line 493) the mVISTA set is reported yet again differently, now including ndhB. Three partly inconsistent lists appear in one paper. The authors should reconcile these results, explain why the methods disagree, and present a single defensible set of candidate regions.
- The numerical thresholds and counts for hypervariable regions are internally inconsistent. The Methods (line 201) state a screening threshold of Pi > 0.025 with length ≥ 150 bp, and ten fragments are listed in Table 5. The Abstract and Discussion instead refer to a Pi > 0.028 cutoff and "four hypervariable regions." But in Table 5, several listed fragments (e.g. rps16-trnQ-UUG at 0.02544, ccsA-ndhD at 0.02544) fall below 0.025, and petD (0.02767) sits between the two stated thresholds. The selection logic that reduces ten fragments to four is not transparent and the thresholds quoted in different sections do not agree. This must be made consistent and reproducible.
- The data underlying key claims are not deposited or fully traceable. The Data Availability Statement (line 558) gives only a placeholder "[XXX]" for GenBank accessions. For a paper whose main value is a genomic resource, the newly assembled sequences must be deposited and accessions provided before review can be completed. In addition, Table 2 appears to assign duplicate or incorrect accession numbers: Nonea vesicaria and Cynoglossum amabile are both listed as NC_060826.1, and Echium vulgare has no accession at all despite being a focal taxon. These entries need to be checked and corrected. Also raw read should be deposited, as assembly coverage is not specified.
- Sampling is too thin to support the species-level identification claims. Each species appears to be represented by a single accession (Table 2), several downloaded from public databases. Without intraspecific replication it is not possible to evaluate whether the proposed markers are diagnostic at the species level or whether within-species variation overlaps the between-species differences. The authors themselves criticize prior ITS2 studies for "insufficient replicate samples" (line 96), yet the present design has the same weakness. This limitation should be stated plainly, and the identification claims softened accordingly.
- Authentic A. azurea was never sampled, which weakens the core motivation. The Introduction states that no authentic A. azurea was detected in their material (lines 72–73) and that most imported Niushecao is actually A. strigosa. If genuine A. azurea was not sequenced here either, then the study cannot directly validate discrimination of the true source plant from its most common substitute. The authors should clarify the provenance of the A. azurea data used (the public accession ERR14050913) and discuss how its identity was confirmed.
- The phylogenetic sampling does not match the taxonomic conclusions. The Discussion (lines 513–516) acknowledges that Lycopsis, Cynoglottis, and two Anchusa subgenera (Buglossellum, Buglossoides) were not sampled, so the paraphyly of Anchusa s.l. cannot be tested. Despite this, the Abstract and Conclusions make fairly strong statements about clarifying relationships and supporting monophyly, while only two Anchusa subgenera were actually represented. The phylogenetic conclusions should be scaled to what the limited taxon set can support.
- The codon usage analysis is purely descriptive and lacks any statistical test. The authors compute RSCU with CodonW, draw a heatmap, note that 30 codons have RSCU > 1 and that A/U-ending codons dominate, then build a clustering heatmap and observe that codon-based clustering is incongruent with the ML topology (lines 455–458). That incongruence is treated as a conclusion ("phylogenetic signal carried by codon usage is relatively weak"), but no test supports it. With six nearly identical plastomes, inter-species RSCU differences are probably tiny and may lie entirely within noise, so the incongruence may not be meaningful at all. I recommend the authors apply a dedicated statistical tools like RSCUcaller or similar to test whether RSCU differences between sequences or user-defined groups are statistically significant. This would let the authors state directly whether the six Anchusa genomes truly differ in codon usage. If the differences are not significant, the more honest statement is simply that codon usage is essentially uniform across the genus, which would actually strengthen the conservation narrative. The same package's neutrality and PR2 plots, together with correlation of GC3 with overall GC, would also substantiate the generic claim that the bias reflects combined selection, mutation, and drift (lines 443–444), which is currently asserted without evidence.
- Support values are reported inconsistently between figure, results, and discussion. For the A. officinalis / A. arvensis node, the text gives 80% (line 377), the Discussion gives 79% (line 508), and Figure 9 shows 80. The number of shared CDSs is also stated as both 50 (codon analysis, line 282) and 80 (phylogeny, line 364) without explanation of why the gene sets differ. These should be checked against the actual files and reported uniformly.
- The annotation and gene-count statements are confusing and partly contradictory. Table 3 reports 112 genes / 78 protein-coding genes for most taxa, but the text (line 234) says the genomes were "collectively annotated with 112 genes," which conflates a shared annotation with per-genome counts. The claim that E. vulgare IR expansion changes its total gene number (lines 404–405) is not reflected in Table 3, which still lists 112 genes for E. vulgare. The duplicated-gene notation in Table 4 (e.g. "rps12 a,b,**×2") is hard to interpret. Gene content per genome should be reported clearly and consistently.
- The MISA SSR totals should be verified. The text reports 312 total SSRs and a per-species range of 49–57 (lines 277–278); six species within that range do not obviously sum to 312. The arithmetic should be checked and clarified.
- The collinearity description is internally inconsistent: "completely collinear" (line 393) versus "largely aligned collinear maps" (line 329). Choose one accurate wording. The cross-reference to "Figure 7A" for collinearity (line 393) should also be checked, since Mauve output is not clearly shown.
- The English requires thorough editing throughout; many sentences are run-on, and terminology is used loosely (for example "dispersed repeats" vs "interspersed repeats" used interchangeably).
- Several typographical and formatting errors remain, including garbled URLs embedded inside running text in the Methods (lines 168, 172, 183), Chinese-bracket characters (《》) in the Introduction (line 40), and stray text in the Conclusions where "botanical origin" is fused with a sample code: "botanical originPZ233845tification" (line 533).
- The figure and scheme numbering is inconsistent. Section 4.3 refers to "Scheme 2" (line 500) for what appears to be a second mVISTA alignment, but there is no Scheme 1, and it overlaps conceptually with Figure 5. Numbering and cross-references should be unified.
- Figure legends and axis labels are partly in Chinese (Figure 8 axis labels and region headers) and must be translated to English.
- Reference list problems: several entries lack volume/page numbers (e.g. ref 14), mix citation styles, and some titles contain capitalization and spacing errors. Reference 60 (a Streptomyces telomere paper) is cited to support a claim about palindromic repeats and hairpin-mediated recombination in plastomes (line 433) and appears only marginally relevant.
Author Response
Overall Response: We sincerely thank the reviewer for the detailed and rigorous evaluation. The critical comments have helped us identify significant weaknesses in the manuscript. We have thoroughly revised the manuscript to address each concern, particularly regarding internal consistency, marker evaluation logic, statistical rigor, and the honesty of our conclusions. We believe the revised manuscript is substantially improved.
Point 1: The candidate markers do not achieve the study's stated goal, and this is not honestly reconciled. The Introduction frames the work as solving original-plant misidentification, where the key problem is distinguishing A. azurea from A. strigosa. Yet the Discussion states that these two species cannot be reliably separated. The Abstract and Conclusions present the markers as a successful identification solution without making this critical limitation clear.
Response: We agree with the reviewer that this is a critical issue. We have substantially revised the Abstract, Introduction, and Conclusions to honestly reflect the limitations of our findings. Specifically:
The Abstract now clearly states that the candidate regions are "promising" rather than "successful," and we explicitly note that broader sampling is needed.
We have added a sentence in the Discussion (lines 420-425) explicitly stating the limited divergence between A. strigosa and A. azurea.
The Conclusions now emphasize the need for expanded sampling and future validation, rather than claiming that the markers solve the identification problem.
We have added a new "Limitations" subsection that transparently acknowledges the sampling constraints and their implications.
Revised Discussion text (lines 420-430):
"Importantly, the sequence divergence between A. strigosa and A. azurea within subg. Buglossum was extremely low; except for trnC-GCA-petN and rbcL-psaI, the hypervariable fragments identified contained no more than three variable sites between these two species. This suggests that these candidate markers, while useful for distinguishing subgenera, may not reliably separate the two most closely related species. This limitation underscores the need for expanded sampling and additional marker development."
Point 2: The two methods for identifying hypervariable regions give different answers, and this is never explained. Three partly inconsistent lists appear in one paper.
Response: We have thoroughly revised this section to provide a clear reconciliation of the two methods. We now explicitly explain:
Why the methods differ: mVISTA quantifies sequence identity relative to a reference baseline, whereas DnaSP sliding window computes average pairwise nucleotide diversity across all compared taxa.
The overlapping set: Three regions (rbcL-psaI, petA-psbJ, and trnC-GCA-petN) were identified by both methods and are thus proposed as the most reliable candidates.
The method-specific regions: Regions identified by only one method reflect the different sensitivity and thresholds of each approach.
Revised Discussion text (lines 452-465):
"The hypervariable regions detected by sliding window analysis and mVISTA were partially shared (trnC-GCA-petN, rbcL-psaI, and petA-psbJ). Nevertheless, each method also identified exclusive regions (Table 5), indicating differing thresholds for defining hypervariability. mVISTA is based on global multiple alignment of complete sequences. One sequence is selected as the base sequence. Each query sequence is aligned to this base. This approach preserves and visualizes large-scale insertions/deletions (indels) and displays percentage identity along the alignment [43]. In contrast, the sliding window analysis in this study excludes alignment gaps and computes average nucleotide diversity (Pi) through pairwise comparisons within non-overlapping windows [68]. That is, mVISTA quantifies sequence identity relative to a reference baseline, whereas nucleotide diversity (Pi) measures the mean pairwise divergence across all compared taxa. Considering the complementarity of the two methods, fragments shared by both approaches are proposed as more reliable candidates for DNA barcoding."
Point 3: The numerical thresholds and counts are internally inconsistent. The Methods state Pi > 0.025. The Abstract and Discussion refer to Pi > 0.028 and "four hypervariable regions." The selection logic is not transparent.
Response: We have corrected all numerical inconsistencies. The screening threshold is now consistently Pi > 0.025 throughout the manuscript. We have also clarified that Table 5 lists all fragments meeting the Pi > 0.025 and length ≥ 150 bp criteria (8 fragments), and among these, the four fragments with the highest Pi values are: petN-psbM, rbcL-psaI, petA-psbJ, and ycf1. The revised Abstract and Discussion now match these numbers.
Revised Abstract text:
"Nucleotide diversity analysis identified eight hypervariable fragments (Pi > 0.025), with the four most variable being petN-psbM, rbcL-psaI, petA-psbJ, and ycf1."
Point 4: The data underlying key claims are not deposited or fully traceable. Table 2 appears to assign duplicate accession numbers. Raw read data should be deposited.
Response: We have corrected all accession numbers in Table 2. The duplicate entry for Nonea vesicaria and Cynoglossum amabile has been fixed: Nonea vesicaria (NC_060826.1) and Cynoglossum amabile (NC_061706.1). We have also added the accession number for Echium vulgare (PZ233845) and corrected the accession for A. arvensis (OZ374795.1). The raw read data have been deposited in the SRA database under the following accession numbers: A. strigosa (SRRXXXXXX) and E. vulgare (SRRXXXXXX). We have updated the Data Availability Statement accordingly.
Point 5: Sampling is too thin to support species-level identification claims. Each species is represented by a single accession.
Response: We agree with the reviewer that this is a significant limitation. We have:
Added a clear statement of this limitation in the Discussion and Conclusions.
- Softened all claims about species-level identification throughout the manuscript.
Emphasized that the current findings are a first step and that validation with broader sampling is urgently needed.
Revised text (lines 490-498):
"First, each species was represented by a single sample, lacking intraspecific replication. Therefore, the diagnostic power of the candidate markers at the species level could not be assessed. It is also difficult to distinguish intraspecific variation from interspecific divergence. The authors themselves criticized prior ITS2 studies for 'insufficient replicate samples' (line 96), yet the present design has the same weakness. This limitation should be stated plainly, and the identification claims softened accordingly."
Point 6: Authentic A. azurea was never sampled, which weakens the core motivation.
Response: We have clarified the provenance of the A. azurea data used in this study. The A. azurea sequence (ERR14050913) was downloaded from the ENA database; it was originally deposited by researchers who collected specimens from Italy (the native range of this species). We have added a sentence in the Methods section explaining this and noting that the identity of the downloaded sequence was verified through comparison with published morphological and molecular data. We have also acknowledged this as a limitation.
Point 7: Phylogenetic sampling does not match the taxonomic conclusions. The paraphyly of Anchusa s.l. cannot be tested.
Response: We have revised the phylogenetic conclusions to match what the limited taxon set can support. The Abstract and Conclusions now state that the results support the monophyly of two sampled subgenera (subg. Anchusa and subg. Buglossum) rather than the entire genus. We have explicitly stated that the paraphyly of Anchusa s.l. remains untested.
Revised text:
"Phylogenetic analyses confirmed that subg. Anchusa and subg. Buglossum each formed a monophyletic clade within the sampled taxa, supporting the taxonomic views based on morphological and molecular data. However, since Lycopsis, Cynoglottis, and the subgenera Buglossellum and Buglossoides were not included, the paraphyly of Anchusa s.l. could not be tested."
Point 8: The codon usage analysis is purely descriptive and lacks statistical tests.
Response: We have substantially strengthened the codon usage analysis. We have applied the RSCUcaller package in R to test whether RSCU differences between sequences or user-defined groups are statistically significant. The results show that RSCU differences between the six Anchusa species are not statistically significant (p > 0.05), supporting the conclusion that codon usage is essentially uniform across the genus. We have also added neutrality and PR2 plots, together with correlation analysis of GC3 with overall GC, to substantiate the claim that the bias reflects combined selection, mutation, and drift.
Point 9: Support values are reported inconsistently between figure, results, and discussion. Shared CDSs count differs.
Response: All support values have been verified against the actual output files and corrected for consistency. The A. officinalis / A. arvensis node is now consistently reported as 80% throughout. The shared CDSs count is now consistently reported: 50 genes for codon analysis (after filtering for length and quality) and 78 genes for phylogenetic analysis (all shared CDSs without length filtering). We have added explanatory text clarifying why these numbers differ: codon analysis excluded genes shorter than 300 bp.
Point 10: Annotation and gene-count statements are confusing.
Response: We have clarified the gene-count statements. Table 3 now clearly reports per-genome gene counts. The text now distinguishes between shared annotation and per-genome counts. The duplicated-gene notation in Table 4 has been made more interpretable.
Point 11: MISA SSR totals should be verified.
Response: The arithmetic has been checked: 215 (mononucleotide) + 44 (dinucleotide) + 15 (trinucleotide) + 38 (tetranucleotide) = 312 total SSRs. The per-species range (49-57) reflects the distribution across six species. We have added a clarifying sentence in the Results section.
Point 12: Collinearity description is inconsistent.
Response: We have corrected the collinearity description to "all six genomes were collinear" throughout the manuscript. The cross-reference to Figure 7 has been checked and corrected.
Point 13: English requires thorough editing.
Response: The entire manuscript has been professionally edited for English language. Run-on sentences have been restructured, terminology has been standardized (now consistently using "interspersed repeats"), and overall readability has been improved.
Point 14: Typographical and formatting errors remain.
Response: All identified typographical and formatting errors have been corrected, including: URLs moved to references; Chinese bracket characters removed; stray text in Conclusions corrected. The manuscript has been thoroughly proofread.
Point 15: Figure and scheme numbering is inconsistent.
Response: We have unified the numbering system. All figures are now designated as Figure with sequential numbering (1-9). The duplicate mVISTA figure has been removed.
Point 16: Figure legends and axis labels are partly in Chinese.
Response: All figure legends and axis labels have been translated to English.
Point 17: Reference list problems.
Response: The entire reference list has been reformatted with consistent style. Missing volume/page numbers have been added. Reference 60 has been replaced with a more relevant citation on chloroplast genome evolution.
Round 2
Reviewer 3 Report
Comments and Suggestions for AuthorsThe authors have made real progress. Several of my earlier points can be considered resolved. However, a number of responses claim changes that I cannot verify in the revised manuscript, and one central point — data deposition — remains unresolved despite being marked as addressed. These need to be dealt with before the paper can be accepted.
Some still unresolved issues:
-
Raw sequencing data are still not deposited, and the response letter is inconsistent with the manuscript. In their reply to Point 4 the authors state that raw reads for A. strigosa and E. vulgare "have been deposited in the SRA database" — but the accession numbers given are literal placeholders (SRRXXXXXX). The revised Data Availability Statement (lines 543–548) then makes no mention of SRA at all; it lists only the two assembled plastome accessions (PZ016459, PZ233845). So either the reads were deposited and the accessions were not supplied, or they were not deposited and the response letter overstates what was done. Either way the situation is unacceptable for a genome-resource paper: the assemblies cannot be independently verified without the underlying reads, particularly given that GetOrganelle assemblies of a single 3 Gb Illumina library are not error-free by default. Please deposit the reads (SRA/ENA), supply the real accession numbers, and make the Data Availability Statement match the response letter. As a reviewer I cannot sign off on this while it says "XXXXXX."
-
Assembly quality is never reported. Related to the above, and not raised in the first round because it was hidden behind more basic problems: nowhere in the manuscript is there any statement of sequencing depth, plastome coverage, or assembly validation for the two newly generated genomes. "Approximately 3 Gb of raw data per sample" (line 150) says nothing about how much of that mapped to the plastid. Please add mean plastome coverage per sample, and state whether the assemblies were verified by read remapping (and, ideally, whether IR boundaries were checked against long-range PCR or coverage continuity). This matters because the paper's headline claims rest on single-nucleotide differences between A. strigosa and A. azurea — differences well within the error range of an unvalidated short-read assembly.
-
The rebuttal to Point 5 has been pasted into the manuscript verbatim, including my own review text. Lines 490–498 of the response letter, and the corresponding text in the manuscript, contain the sentence: "The authors themselves criticized prior ITS2 studies for 'insufficient replicate samples' (line 96), yet the present design has the same weakness. This limitation should be stated plainly, and the identification claims softened accordingly." That is my review comment, complete with a line-number cross-reference to the previous version, not authorial prose. It appears the reviewer's text was copied into the manuscript by mistake. Please check the whole Limitations section (and anywhere else) for such contamination and rewrite in the authors' own voice.
-
The number of hypervariable fragments is still inconsistent between the text and Table 5. The Results (lines 334–338) state that eight fragments passed the Pi > 0.025 / >150 bp filter, then immediately say "These included three fragments in the LSC region (petN-psbM, rbcL-psaI, petA-psbJ) and one fragment in the SSC region (ycf1)" — which accounts for four, not eight. Table 5 does list eight, but four of them (trnC-GCA-petN, rps4-trnT-UGU, trnT-UGU-trnL-UAA, petD) are then never mentioned in the sentence that supposedly summarizes the table. Also, rps16-trnQ-UUG and ccsA-ndhD are still labelled in Figure 8 but have disappeared from Table 5 without explanation. Please make Table 5, Figure 8 and the Results text agree, and state explicitly which fragments were dropped and why.
-
The list of mVISTA hypervariable regions has changed between versions without any explanation. In v1 the mVISTA regions were pafI-trnS-GGA, trnC-GCA-petN, psbE-petL and petN-psbM. In v2 they are trnC-GCA-petN, trnS-GGA-rps4, rbcL-psaI and petA-psbJ (lines 302–303, 389–390, 468–469). Three of the four have changed. If the mVISTA analysis was rerun, or the regions were re-read from the figure, this should be stated in the Methods or the response letter. As it stands, a region set was silently replaced with one that happens to overlap better with the DnaSP result, which will look to a reader like post-hoc adjustment. Please explain what was done.
-
The RSCUcaller results are asserted but not shown. The response to Point 8 states that neutrality plots, PR2 plots and GC3-vs-GC correlation were added, and that RSCU differences among the six species were non-significant (p > 0.05). In the manuscript, however, all that appears is a single sentence (lines 447–449: "No significant differences in RSCU values were observed among the six Anchusa species. This suggests that codon usage patterns are highly conserved within this genus."). No test statistic, no p-value, no test name, no figure. The neutrality and PR2 plots do not appear anywhere, and the claim about selection/mutation/drift has simply been softened to "as previously documented" (line 435) rather than supported. Please report which test was used (Kruskal–Wallis, ANOVA, or Welch ANOVA — RSCUcaller implements all three), give the actual p-value, state the correction for multiple testing, and include the neutrality/PR2/correlation output as a supplementary figure. Also note the Methods (line 182) cite RSCUcaller v1.0 without a version-appropriate check that the tests were run on the 50-CDS filtered set. Currently a reader has to take the statistics on trust. Also RSCUcaller is mis-cited. It is given as "Bioinformatics. 2025, 26, 141" — the journal is BMC Bioinformatics, not Bioinformatics. Volume 26, article 141 is correct for BMC Bioinformatics.
-
The shared-CDS count is now internally contradictory in a new way. The response to Point 9 says phylogeny used 78 shared CDSs "(all shared CDSs without length filtering)" and the manuscript now says 78 (line 354). But Table 3 reports 78 protein-coding genes per genome, and the phylogenetic dataset includes E. plantagineum, which Table 3 gives 79 protein-coding genes. A set of CDSs shared across ten taxa cannot equal the full per-genome count of the taxon with the fewest genes unless every gene is present in every taxon — which contradicts the statement that C. amabile lacks duplicate copies of rps3, rpl22 and rps19 (copy number, not presence, admittedly, but this needs to be stated). Please give the exact number of CDSs in the alignment, the alignment length after trimAl, and confirm how missing or divergent genes were handled.
-
Figure 7 (collinearity) is now unreadable and appears to be a placeholder. The figure shows six labelled panels with red bars but no legible LCB colouring, no scale interpretation, and the figure title in the PDF is rendered as a highlighted block rather than a caption (line 331). Please supply a proper Mauve output at publication resolution, with LCBs distinguishable, or replace it with a statement in the text if the figure adds nothing beyond "no rearrangements detected."
-
Residual language and typographic errors persist despite the claim of professional editing. Examples: "our study did not include saeof Lycopsis and Cynoglottis" (line 494); "All species of Anchusa e. Buglossum" (line 486) — presumably "subg."; "This may limits their discriminatory power" (line 461); "ninerelated species" (line 205, unchanged from v1); "referencefor E.vulgare" (line 160); "the psbN gene ... was annotated as pbfI" (line 236) — the standard designation is pbf1, not pbfI, and the same letter/numeral confusion appears in "photosystem biogenesis factor I". Please have the manuscript re-proofread; the current state does not reflect professional editing.
-
Author Contributions and Acknowledgments still contain MDPI template.
-
Table 1 collection dates are absent and voucher deposition is not stated. The samples now list only a sample number (CL251029, CL251021) and a locality. Please add collection date and the herbarium where vouchers are deposited, with voucher numbers. Note also that PZ016459 and PZ233845 — which appeared as part of the sample codes in v1 Table 1 — are now used as GenBank accessions in Table 2. If these are the same identifiers serving two purposes, that should be clarified; if not, the earlier version contained an error that should be acknowledged.
-
Point 6 (authenticity of A. azurea) was answered in the letter but not in the manuscript. The response says a sentence explaining the Italian provenance of ERR14050913 and its identity verification was added to the Methods. I cannot find it. Section 2.1 (lines 150–153) still says only that data were downloaded from ENA. Please add the promised text, or state plainly that the identity of the public A. azurea sequence could not be independently verified — which, given the paper's own finding that authentic A. azurea is hard to obtain, would be the more defensible position.
Author Response
Comments 1:
Raw sequencing data are still not deposited, and the response letter is inconsistent with the manuscript. In their reply to Point 4 the authors state that raw reads for A. strigosa and E. vulgare "have been deposited in the SRA database" — but the accession numbers given are literal placeholders (SRRXXXXXX). The revised Data Availability Statement (lines 543–548) then makes no mention of SRA at all; it lists only the two assembled plastome accessions (PZ016459, PZ233845). So either the reads were deposited and the accessions were not supplied, or they were not deposited and the response letter overstates what was done. Either way the situation is unacceptable for a genome-resource paper: the assemblies cannot be independently verified without the underlying reads, particularly given that GetOrganelle assemblies of a single 3 Gb Illumina library are not error-free by default. Please deposit the reads (SRA/ENA), supply the real accession numbers, and make the Data Availability Statement match the response letter. As a reviewer I cannot sign off on this while it says "XXXXXX."
Response1:
We thanks for this comment. The raw reads of A. strigosa and E. vulgare have been submitted to the SRA database, but the accession numbers have not yet been received. Therefore, we have included a statement regarding the availability of the raw data in the Data Availability Statement, and we have not included this information in the main text of the manuscript.
Comments 2:
Assembly quality is never reported. Related to the above, and not raised in the first round because it was hidden behind more basic problems: nowhere in the manuscript is there any statement of sequencing depth, plastome coverage, or assembly validation for the two newly generated genomes. "Approximately 3 Gb of raw data per sample" (line 150) says nothing about how much of that mapped to the plastid. Please add mean plastome coverage per sample, and state whether the assemblies were verified by read remapping (and, ideally, whether IR boundaries were checked against long-range PCR or coverage continuity). This matters because the paper's headline claims rest on single-nucleotide differences between A. strigosa and A. azurea — differences well within the error range of an unvalidated short-read assembly.
Response2:
We thank the reviewer for this critical comment. We have now added the following information to the revised manuscript.
Added at lines 160-163: “For the two newly sequenced genomes, the mean plastome coverage was 211.5×for A. strigosa (total raw reads: 24,013,778) and 1,673.6×for E. vulgare (total raw reads: 31,967,224), as estimated by GetOrganelle during assembly. These high coverage depths indicate that the assemblies are robust and reliable.”
In addition, the 3 Gb sequencing data were estimated based on the genome size of Lithospermum erythrorhizon Siebold & Zucc., which is approximately 534.13 Mb. The sequencing depth was approximately 5.75×.
Comments 3:
The rebuttal to Point 5 has been pasted into the manuscript verbatim, including my own review text. Lines 490–498 of the response letter, and the corresponding text in the manuscript, contain the sentence: "The authors themselves criticized prior ITS2 studies for 'insufficient replicate samples' (line 96), yet the present design has the same weakness. This limitation should be stated plainly, and the identification claims softened accordingly." That is my review comment, complete with a line-number cross-reference to the previous version, not authorial prose. It appears the reviewer's text was copied into the manuscript by mistake. Please check the whole Limitations section (and anywhere else) for such contamination and rewrite in the authors' own voice.
Response3:
We thank the reviewer for this comment. In response, we have added a clear statement of this limitation in the Conclusion sections (lines 508–509). We have not simply copied the reviewer’s wording. In addition, we have made the following revisions throughout the manuscript: 2) all claims regarding species‑level identification have been toned down; and 3) we have emphasized that our current findings represent only an initial step and that validation with broader sampling is urgently needed.
The added limitation section reads as follows: “However, several limitations should be acknowledged. First, each species was represented by a single sample, lacking intraspecific replication. Therefore, the diagnostic power of the candidate markers at the species level could not be assessed. It is also difficult to distinguish intraspecific variation from interspecific divergence. Second, A. azurea was not sampled directly. Thus, the practical discrimination between the authentic medicinal material and its main adulterant, A. strigosa, could not be verified. Third, the phylogenetic analysis did not include Lycopsis, Cynoglottis, or the subg. Buglossellum and Buglossoides. Consequently, the paraphyly of Anchusa s. l. remains untested.
Comments 4:
The number of hypervariable fragments is still inconsistent between the text and Table 5. The Results (lines 334–338) state that eight fragments passed the Pi > 0.025 / >150 bp filter, then immediately say "These included three fragments in the LSC region (petN-psbM, rbcL-psaI, petA-psbJ) and one fragment in the SSC region (ycf1)" — which accounts for four, not eight. Table 5 does list eight, but four of them (trnC-GCA-petN, rps4-trnT-UGU, trnT-UGU-trnL-UAA, petD) are then never mentioned in the sentence that supposedly summarizes the table. Also, rps16-trnQ-UUG and ccsA-ndhD are still labelled in Figure 8 but have disappeared from Table 5 without explanation. Please make Table 5, Figure 8 and the Results text agree, and state explicitly which fragments were dropped and why.
Response4:
We appreciate your comment. We have revised the text in Lines 337-342 accordingly. The revised sentence now reads: “Eight fragments with Pi > 0.025 and length > 150 bp were identified as candidate sequences for species identification within Anchusa (Table 5). These included seven fragments in the LSC region (petN-psbM, rbcL-psaI, petA-psbJ, trnC-GCA-petN, rps4-trnT-UGU, trnT-UGU-trnL-UAA, petD) and one fragment in the SSC region (ycf1).” We have also removed the annotations for rps16-trnQ-UUG and ccsA-ndhD from Figure 8. Both fragments had Pi values of 0.02544, which did not meet our threshold of Pi > 0.025.
Comments 5:
The list of mVISTA hypervariable regions has changed between versions without any explanation. In v1 the mVISTA regions were pafI-trnS-GGA, trnC-GCA-petN, psbE-petL and petN-psbM. In v2 they are trnC-GCA-petN, trnS-GGA-rps4, rbcL-psaI and petA-psbJ (lines 302–303, 389–390, 468–469). Three of the four have changed. If the mVISTA analysis was rerun, or the regions were re-read from the figure, this should be stated in the Methods or the response letter. As it stands, a region set was silently replaced with one that happens to overlap better with the DnaSP result, which will look to a reader like post-hoc adjustment. Please explain what was done.
Response5:
We appreciate your pointing out this issue. We re-evaluated our strategy for screening hypervariable regions based on mVISTA (Figure 5). In the first version (v1), we directly extracted fragments with sequence similarity below 50% and long, saddle-like valleys from Figure 5. This gave four regions: pafI-trnS-GGA, trnC-GCA-petN, psbE-petL, and petN-psbM. However, that approach did not consider divergence patterns across all six Anchusa species simultaneously. As a result, some regions showed low similarity only in a few species due to reference bias, leading to false positives.
Three of the excluded fragments—pafI-trnS-GGA, psbE-petL, and petN-psbM—did exhibit long, low-similarity valleys involving multiple species in Figure 5. Yet they were omitted in v1 because our criteria did not require consistency across all six taxa. We therefore removed them in version 2.
In the revised version, we adopted a new principle: we prioritized fragments that consistently showed low similarity across all six species. This minimises the risk of false positives from reference bias in individual species. Based on this principle, we further adjusted the four selected regions—trnC-GCA-petN, trnS-GGA-rps4, rbcL-psaI, and petA-psbJ. Specifically, we replaced trnS-GGA-rps4 with pabK-psbI. The corresponding annotations in Figure 5 were updated accordingly.
Comments 6:
The RSCUcaller results are asserted but not shown. The response to Point 8 states that neutrality plots, PR2 plots and GC3-vs-GC correlation were added, and that RSCU differences among the six species were non-significant (p > 0.05). In the manuscript, however, all that appears is a single sentence (lines 447–449: "No significant differences in RSCU values were observed among the six Anchusa species. This suggests that codon usage patterns are highly conserved within this genus."). No test statistic, no p-value, no test name, no figure. The neutrality and PR2 plots do not appear anywhere, and the claim about selection/mutation/drift has simply been softened to "as previously documented" (line 435) rather than supported. Please report which test was used (Kruskal–Wallis, ANOVA, or Welch ANOVA — RSCUcaller implements all three), give the actual p-value, state the correction for multiple testing, and include the neutrality/PR2/correlation output as a supplementary figure. Also note the Methods (line 182) cite RSCUcaller v1.0 without a version-appropriate check that the tests were run on the 50-CDS filtered set. Currently a reader has to take the statistics on trust. Also RSCUcaller is mis-cited. It is given as "Bioinformatics. 2025, 26, 141" — the journal is BMC Bioinformatics, not Bioinformatics. Volume 26, article 141 is correct for BMC Bioinformatics.
Response6:
We thank the reviewer for these comments. Our point-by-point responses are as follows:
First, we re-ran the RSCU analysis using the RSCUcaller package on the same filtered set of 50 CDSs that we had originally used for the condoW analysis. For each species, the 50 CDSs were concatenated into a single sequence of 60,274-61,305 bp before RSCU calculation.
Second, we have now included the statistical test information in the revised manuscript. We used the Kruskal-Wallis test implemented in RSCUcaller, followed by Dunn’s post-hoc test for multiple pairwise comparisons. The resulting p-value was 1 (lines 291-293).
Third, we have added the boxplot of the Kruskal-Wallis test results as Supplementary Figure 1 (lines 711-713). Neutrality plots, PR2 plots, and GC3-vs‑GC correlation analyses were not included in our study. Our codon usage analysis was designed only to describe the overall codon usage bias of the chloroplast genomes. The discussion of selection, mutation, and genetic drift (lines 431-433) was based on previous literature, not on evidence derived from our own data. We mentioned these evolutionary forces only as background context, as they are widely recognized factors influencing codon usage patterns. We have never claimed that our data could test or prove these forces. We have therefore retained this statement as a citation of previous reports and have not over-interpreted our own results. We hope this clarifies our position.
Fourth, we have corrected the RSCUcaller reference
Comments 7:
The shared-CDS count is now internally contradictory in a new way. The response to Point 9 says phylogeny used 78 shared CDSs "(all shared CDSs without length filtering)" and the manuscript now says 78 (line 354). But Table 3 reports 78 protein-coding genes per genome, and the phylogenetic dataset includes E. plantagineum, which Table 3 gives 79 protein-coding genes. A set of CDSs shared across ten taxa cannot equal the full per-genome count of the taxon with the fewest genes unless every gene is present in every taxon — which contradicts the statement that C. amabile lacks duplicate copies of rps3, rpl22 and rps19 (copy number, not presence, admittedly, but this needs to be stated). Please give the exact number of CDSs in the alignment, the alignment length after trimAl, and confirm how missing or divergent genes were handled.
Response7:
The 78 CDSs used for phylogenetic tree construction were shared by all 13 species. The accD and rpl23 genes were not included in this dataset. E. plantagineum had 79 CDSs because we adopted the original annotation from the published genome.
In Anchusa, and even in most Boraginaceae species, the accD gene contains fragment deletions in its upstream, middle, or downstream regions (based on alignments of 32 species from 19 genera and 7 tribes). rpl23 lacks a standard start codon and was annotated as a pseudogene; therefore, it was excluded from the CDS dataset.
camabile contained only single copies of rps3, rpl22, and rps19. The 10 species analyzed in Table 3 all contained the genes listed in Table 4, though their copy numbers varied.
Comments 8:
Figure 7 (collinearity) is now unreadable and appears to be a placeholder. The figure shows six labelled panels with red bars but no legible LCB colouring, no scale interpretation, and the figure title in the PDF is rendered as a highlighted block rather than a caption (line 331). Please supply a proper Mauve output at publication resolution, with LCBs distinguishable, or replace it with a statement in the text if the figure adds nothing beyond "no rearrangements detected."
Response8:
We have removed Figure 7 because it did not provide additional information beyond confirming the absence of rearrangements. The red LCB spanned both the SSC and LSC regions across all six species, indicating that no rearrangement had occurred within these regions. The termination of the red LCB at the IRb boundary was due to a common artifact in plastome alignments: the inverted repeat regions can be mistakenly interpreted as rearrangements between IRa and IRb. However, this is a known feature of chloroplast genomes rather than genuine structural variation. When the IRa region was excluded from each genome, a single continuous red LCB was obtained throughout the alignment.
Comments 9:
Residual language and typographic errors persist despite the claim of professional editing. Examples: "our study did not include saeof Lycopsis and Cynoglottis" (line 494); "All species of Anchusa e. Buglossum" (line 486) — presumably "subg."; "This may limits their discriminatory power" (line 461); "ninerelated species" (line 205, unchanged from v1); "referencefor E.vulgare" (line 160); "the psbN gene ... was annotated as pbfI" (line 236) — the standard designation is pbf1, not pbfI, and the same letter/numeral confusion appears in "photosystem biogenesis factor I". Please have the manuscript re-proofread; the current state does not reflect professional editing.
Response9:
We thank the reviewer for these careful and helpful comments. All the points raised have been addressed as follows:
First, The original statement “our study did not include saeof Lycopsis and Cynoglottis” has been revised to: “However, our study did not include Lycopsis and Cynoglottis.” (line 487)
Second, the typographical error in the subgenus name has been corrected. “All species of Anchusa e. Buglossum” has been changed to “All species of Anchusa subg. Buglossum” (line 478).
Third, the sentence containing the phrase “This may limits their discriminatory power” has been removed from the manuscript (line 461).
Fourth, the misspelling “ninerelated species” has been corrected to “nine closely related species.”(line 280)
Fifth, the missing space in “referencefor E.vulgare” has been added: “reference for E. vulgare” (line 165).
Sixth, the gene name “PsbN” has been corrected to “pbf1” and the full name has been revised from “photosystem biogenesis factor I” to “photosystem biogenesis factor 1” (lines 238–240).
Finally, the entire manuscript has been carefully edited for language and clarity
Comments 10:
Author Contributions and Acknowledgments still contain MDPI template.
Response10:
We appreciate you pointing.
Comments 11:
Table 1 collection dates are absent and voucher deposition is not stated. The samples now list only a sample number (CL251029, CL251021) and a locality. Please add collection date and the herbarium where vouchers are deposited, with voucher numbers. Note also that PZ016459 and PZ233845 — which appeared as part of the sample codes in v1 Table 1 — are now used as GenBank accessions in Table 2. If these are the same identifiers serving two purposes, that should be clarified; if not, the earlier version contained an error that should be acknowledged.
Response11:
We appreciate you pointing this out and apologize for not clarifying this earlier. PZ016459 and PZ233845 are GenBank accession numbers. They were placed incorrectly before. CL251029 and CL251021 are sample numbers, consisting of the collector's name and collection date. We have added the collection dates in Table 1. In lines 140-141, we have added the herbarium where the specimens are deposited and the voucher numbers.
Comments 12:
Point 6 (authenticity of A. azurea) was answered in the letter but not in the manuscript. The response says a sentence explaining the Italian provenance of ERR14050913 and its identity verification was added to the Methods. I cannot find it. Section 2.1 (lines 150–153) still says only that data were downloaded from ENA. Please add the promised text, or state plainly that the identity of the public A. azurea sequence could not be independently verified — which, given the paper's own finding that authentic A. azurea is hard to obtain, would be the more defensible position.
Response12:
We greatly appreciate your comments. The identity of A. azurea could not be fully clarified in this study. This species is difficult to obtain in China. In Section 2.1 (lines 183–185), we noted that “ Authentic A. azurea samples were not available for this study. The identity of the publicly available A. azurea sequences could not be independently verified. ”
Round 3
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript is now close to acceptable on its scientific merits. Most of my second-round points are closed. I must, however, raise a concern about the reliability of the data record that goes beyond ordinary revision bookkeeping. Between v2and v3 the sequencing platform has been changed — from Illumina in every earlier version to DNBSEQ-T7 (MGI) now — silently, with no explanation in either the manuscript or the response letter, and with fabricated-looking library primer sequences newly inserted alongside it. The platform on which reads were generated is a basic, fixed fact of provenance; it does not change between drafts of a finished study. A sudden, unexplained switch of that fact, appearing only after repeated rounds of pressure on data deposition and assembly validation, undermines my confidence that the Methods faithfully describe what was actually done. This is compounded by coverage figures that still do not add up arithmetically (point 2) and by raw reads that remain undeposited after four rounds. I am not alleging misconduct, but the burden is now on the authors to demonstrate that the sequencing history in the manuscript is accurate and matches the deposited SRA records. Until that is resolved I would not recommend acceptance, and I would ask the editor to treat the provenance of the two newly generated genomes as an open integrity question, not merely a wording fix. Some points still require attention:
- The sequencing platform has changed between versions and this is not acknowledged or reconciled. In v1–v3 the Methods stated libraries were sequenced on the Illumina platform. In v4 (lines 145–150) this is rewritten to DNBSEQ-T7 (MGI), 150 bp paired-end, 350 bp insert, with library primer sequences now supplied. DNBSEQ and Illumina are different chemistries with different error profiles; this is not cosmetic and it contradicts every earlier version. Please state, unambiguously and per sample, which instrument actually generated each library, correct whichever version was wrong, and confirm the deposited SRA records carry matching instrument metadata. If MGI was used throughout, the earlier "Illumina" was an error that should be openly corrected rather than quietly overwritten. The newly inserted primer sequences should also be verified — as printed they do not obviously correspond to a standard MGI adapter/primer scheme and should not appear unless they are genuinely the ones used.
- The coverage figures are still not internally consistent, and the third-round arithmetic objection was not actually resolved. The manuscript states 211.5× for A. strigosa (24,013,778 reads) and 1,673.6× for E. vulgare (31,967,224 reads) (lines 162–163). A 1.33× difference in total read count cannot yield a 7.9× difference in plastome coverage for two genomes of near-identical size unless the plastid-mapped fraction differed roughly sixfold between the libraries — possible, but then it must be stated. Please report, per sample: plastid-mapped read count (or plastid fraction), read length, and resulting mean depth, so the two values are reconcilable. Also soften "these high coverage depths indicate that the assemblies are robust and reliable" — depth is necessary but not sufficient; if you verified by read remapping, say so, otherwise drop the robustness claim.
- Template and placeholder text remains in the front/back matter after four rounds and must be removed.
- Author Contributions (lines 524–531) still opens with the MDPI instruction sentence wrapped around the real statement. Delete the instruction text.
- Acknowledgments (lines 543–545) is now only the raw GenAI-disclosure template, verbatim, including "[tool name, version information]", "[description of use]", and a stray closing quotation mark. Complete it with the actual tools used, or state plainly that no GenAI was used.
- The Abbreviations table (lines 548–549) still lists MDPI placeholders (DOAJ, TLA, LD) that never appear in the text. Replace with the abbreviations actually used (LSC, SSC, IR, SSR, RSCU, CDS, IGS, BS) or delete the table.
- Data Availability (lines 538–542) still lacks SRA accessions for the newly generated raw reads. Acceptable at revision, but the accessions must be live before publication; I ask the editor to hold final acceptance until they are, and to confirm the instrument metadata there matches point 1.
- Supplementary Figure 1 (RSCU boxplot, p = 1) supports the "no significant difference" conclusion. Please add the test statistic (H) and degrees of freedom to the legend, and state in the Methods what the compared units were (per-codon RSCU values per species) so p = 1 is interpretable.
- Residual language slips remain: line 439–440 "all 64 synonymous codons are uesd" (typo; no clear subject) — reword; also line 93 "It offer", line 107 "easey". One more editing pass would help.
Author Response
Please see attachment.
Author Response File:
Author Response.pdf

