Review Reports
- Toar Jean Maurice Lalisang 1,*,
- Vania Myralda Giamour Marbun 1 and
- Aisyah Fitriannisa Prawiningrum 2,3
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Tomislav Tosti Reviewer 4: Chen Dong Reviewer 5: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe study presented in the manuscript is very interesting; however, it should be clearly emphasized that it is based on a small, exploratory dataset. A total of 15 FFPE samples were included in the analysis, of which only 11 cases remained after quality control. In addition, this is a tumor-only design, without matched normal tissue and without a control group. For this reason, the manuscript should be presented more cautiously. In my opinion, this is a pilot and prioritization study rather than a basis for far-reaching conclusions about disease-associated variants.
In this context, it might be useful to reconsider the use of the terms “GWAS” and “genome-wide SNP study.” The title, abstract, and keywords imply a genome-wide association study, but in reality, this work is more similar to a screening SNP analysis on tumor-only material with functional prioritization of the findings. The current terminology could therefore exaggerate the scope of the study.
The biological interpretation sometimes appears a bit too strong regarding the data presented. While the focus on NASP and GPR78 is an interesting starting point, these should be viewed as preliminary candidates for further validation rather than well-established biological or clinical targets. Likewise, the section on clinical correlations is interesting, but given the small sample size and lack of a true comparative analysis, it remains more observational than analytical.
Additionally, the consistency of the nomenclature for GPR78/GRP78/HSPA5 should be verified. GPR78 is used in the title and results, while in the discussion, the authors switch to GRP78/HSPA5 and cite literature related to that molecule. This causes unnecessary confusion and needs clear clarification.
Finally, the manuscript would definitely benefit from more careful language editing. The text is understandable, but it contains quite a few minor errors, awkward phrases, and inconsistencies in terminology, which reduce its clarity and overall impression.
Comments on the Quality of English LanguageSeveral types and grammatical errors may be detected; the authors should again carefully check their manuscript.
Author Response
[General Comment]
The study presented in the manuscript is very interesting; however, it should be clearly emphasized that it is based on a small, exploratory dataset. A total of 15 FFPE samples were included in the analysis, of which only 11 cases remained after quality control. In addition, this is a tumor-only design, without matched normal tissue and without a control group. For this reason, the manuscript should be presented more cautiously. In my opinion, this is a pilot and prioritization study rather than a basis for far-reaching conclusions about disease-associated variants.
Response:
We sincerely thank the Reviewer for this constructive and important observation. We fully agree that this study should be presented as a small, exploratory, tumor-only analysis. Accordingly, we have revised the manuscript to consistently emphasize its pilot and hypothesis-generating nature throughout the Abstract, Introduction, Discussion, and Conclusions. All interpretations have been tempered to avoid overstatement, and we have explicitly framed the findings as preliminary candidates requiring validation rather than established disease-associated variants.
Specifically, the Abstract now states that the study is an “exploratory SNP-array–based profiling” and the Conclusions emphasize the “feasibility” of the approach and describe NASP and GPR78 as “preliminary candidates.” The Discussion section now includes a comprehensive limitations paragraph explicitly acknowledging the tumor-only design without matched normal tissue, the small sample size (n = 11), the absence of sequencing-based validation, and the descriptive nature of clinical correlations (Discussion, Limitations paragraph).
[Comment on GWAS terminology]
It might be useful to reconsider the use of the terms “GWAS” and “genome-wide SNP study.” The title, abstract, and keywords imply a genome-wide association study, but in reality, this work is more similar to a screening SNP analysis on tumor-only material with functional prioritization of the findings. The current terminology could therefore exaggerate the scope of the study.
Response:
We agree entirely that the term “GWAS” may be misleading in the absence of association testing and a control cohort. Accordingly, we have revised the title to “An Initial Indonesian Genome-Wide SNP-Array Study with Functional Variant Prioritization Reveals NASP and GPR78 Candidate SNVs in FFPE Hepatocellular Carcinoma,” which now includes the word “Candidate” to temper the implication of confirmed findings. In the Abstract and Introduction, we now consistently describe the study as an “exploratory SNP-array–based profiling with functional variant prioritization” rather than a genome-wide association study. We have also revised the Keywords to remove “GWAS” and replaced it with terminology that more accurately reflects the study design (Keywords section, revised manuscript).
[Comment on biological interpretation]
The biological interpretation sometimes appears a bit too strong regarding the data presented. While the focus on NASP and GPR78 is an interesting starting point, these should be viewed as preliminary candidates for further validation rather than well-established biological or clinical targets.
Response:
We appreciate this important feedback. The interpretation of NASP and GPR78 has been revised throughout the manuscript to present them consistently as preliminary, hypothesis-generating candidates for further validation, rather than established biological targets. In the Discussion, NASP is described as “the highest-ranked candidate gene” that provides a “biologically plausible signal that warrants further functional validation,” and GPR78 is framed as “a novel, hypothesis-generating signal, rather than a biologically established contributor to HCC.” Clinical observations are explicitly noted as descriptive only, and the Conclusions state that findings “should not be considered evidence of disease-associated driver events” (Discussion and Conclusions sections).
[Comment on GPR78/GRP78/HSPA5 nomenclature]
The consistency of the nomenclature for GPR78/GRP78/HSPA5 should be verified. GPR78 is used in the title and results, while in the discussion, the authors switch to GRP78/HSPA5 and cite literature related to that molecule. This causes unnecessary confusion and needs clear clarification.
Response:
We thank the Reviewer for identifying this important nomenclature issue. We confirm that the correct gene identified in our analysis is GPR78 (G protein-coupled receptor 78). All incorrect references to GRP78 (glucose-regulated protein 78 kDa, also known as HSPA5/BiP) have been removed from the body text of the revised manuscript to ensure consistency. We acknowledge that GRP78 and GPR78 are entirely distinct molecules; the previous confusion arose from a typographical error. The literature citations referencing GRP78/HSPA5 (References 15–17) have been retained in the reference list as they were included in the original submission, but the Discussion text now clearly distinguishes GPR78 as the candidate gene identified in our study and no longer conflates it with GRP78/HSPA5 biology (Discussion section, revised manuscript).
[Comment on language editing]
The manuscript would definitely benefit from more careful language editing. The text is understandable, but it contains quite a few minor errors, awkward phrases, and inconsistencies in terminology, which reduce its clarity and overall impression.
Response:
We appreciate this feedback. The entire manuscript has been carefully edited for language clarity, grammar, and consistency in terminology. We have also standardized scientific nomenclature and removed awkward phrasing throughout.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis paper is interesting, but why is it called GWAS when there was no association analysis?
How can you distinguish between somatic and germline variants?
What are the advantages of using SNP arrays instead of sequencing?
Why apply Hardy-Weinberg filtering to tumor samples?
What is the purpose of the weighting scheme?
How reliable is gene prioritization with n=11?
Why are known HCC driver genes absent?
What biological process links NASP to liver cancer development?
Are the variants confirmed by sequencing?
Are these variants present in Indonesian population databases?
There is limited comparison with current HCC genetics.
Not discussed: Major HCC driver genes, such as TP53, CTNNB1, TERT, AXIN1, and ARID1A.
Author Response
[Comment 1] Why is it called GWAS when there was no association analysis?
Response:
We agree that the term “GWAS” is misleading in the absence of association testing and a control cohort. We have revised the title, abstract, and keywords to remove the term “GWAS” and now describe the study as an “exploratory SNP-array–based profiling with functional variant prioritization,” which more accurately reflects the design (Title, Abstract, Keywords).
[Comment 2] How can you distinguish between somatic and germline variants?
Response:
Due to the tumor-only design without matched normal tissue, it is not possible to distinguish somatic mutations from rare germline variants. We have now emphasized this limitation more explicitly in the Discussion, stating: “The tumor-only design without matched normal tissue precludes the distinction between somatic mutations and rare germline variants. As a result, the prioritized variants may represent population-specific rare variants rather than tumor-specific alterations.” The interpretation of findings has been revised throughout to reflect their hypothesis-generating nature (Discussion, Limitations paragraph).
[Comment 3] What are the advantages of using SNP arrays instead of sequencing?
Response:
SNP-array platforms provide a cost-effective and practical approach for genome-wide interrogation, particularly in FFPE-derived samples with limited DNA quality. In this study, the SNP-array approach was used as an initial screening and hypothesis-generating strategy. The Discussion now explicitly states: “The use of SNP-array genotyping in FFPE-derived tumor samples represents a pragmatic approach in settings where sequencing resources may be limited and DNA quality is suboptimal. While next-generation sequencing provides higher resolution, SNP-array platforms can still yield informative genome-wide signals when combined with rigorous quality control and systematic functional annotation” (Discussion, Methodological Perspective paragraph).
[Comment 4] Why apply Hardy–Weinberg filtering to tumor samples?
Response:
Hardy–Weinberg equilibrium filtering was applied solely as a technical quality control heuristic to identify potential genotyping artifacts. We have clarified in the Methods that “Hardy–Weinberg equilibrium and minor allele frequency filters were evaluated only as supplementary quality control heuristics to identify potential technical artifacts and were not interpreted as biologically expected features of tumor-derived genotypes” (Methods, SNP-Array Preprocessing and Quality Control subsection).
[Comment 5] What is the purpose of the weighting scheme?
Response:
The weighting scheme was designed to provide a structured and transparent prioritization framework for variants based on predicted functional impact. By integrating multiple in silico prediction tools (SIFT, PolyPhen-2, CADD, REVEL, SpliceAI, and MaxEntScan), the scheme allows aggregation of variant-level information into a gene-level burden score, enabling identification of candidate genes in a tumor-only dataset where conventional association testing is not feasible. The rationale is now clarified in the Methods under the “Weighting Scheme” and “Gene-Level Burden Analysis” subsections.
[Comment 6] How reliable is gene prioritization with n=11?
Response:
We agree that the small sample size limits the robustness of the prioritization. We have emphasized throughout the revised manuscript that the results should be interpreted as exploratory and hypothesis-generating, rather than definitive. The primary purpose of this analysis is to generate a shortlist of candidate genes for further validation in larger cohorts. The Discussion explicitly states: “The small sample size (n = 11 after quality control) limits the robustness and generalizability of the findings and increases the risk of stochastic signals” (Discussion, Limitations paragraph).
[Comment 7] Why are known HCC driver genes absent?
Response:
The absence of canonical HCC driver genes such as TP53, CTNNB1, TERT, AXIN1, and ARID1A likely reflects the limited coverage of SNP-array platforms, which are not optimized for comprehensive mutation discovery and may not capture key somatic alterations outside of predefined probe regions. The small sample size further reduces the likelihood of detecting recurrent driver events. We have now explicitly discussed this limitation and contextualized our findings within the current HCC genomic landscape in the revised Discussion: “Notably, well-established HCC driver genes such as TP53, CTNNB1, TERT, AXIN1, and ARID1A were not identified in the prioritized dataset. This observation likely reflects the inherent limitations of SNP-array platforms. These factors underscore that the present study is not designed to identify canonical drivers, but rather to generate a prioritized shortlist of candidate variants for subsequent investigation” (Discussion).
[Comment 8] What biological process links NASP to liver cancer development?
Response:
NASP encodes a histone-binding protein involved in chromatin assembly, cell cycle progression, and DNA replication. Dysregulation of chromatin organization is a well-recognized hallmark of cancer. The Discussion now states: “Although NASP has not been extensively characterized in hepatocellular carcinoma, prior studies in other malignancies and experimental models suggest a potential role in promoting cellular proliferation and tumor growth. In this context, the prioritization of NASP in our dataset provides a biologically plausible signal that warrants further functional validation.” We have supported this with citations including Li et al. (2019) on NASP in melanoma proliferation, and studies on c-Myc–driven hepatocarcinogenesis (Moon et al., 2021; Sequera et al., 2022) (Discussion, NASP paragraph; References 11–14).
[Comment 9] Are the variants confirmed by sequencing?
Response:
No orthogonal validation using sequencing was performed in this study. We have acknowledged this as a limitation and explicitly stated that sequencing-based confirmation is required in future studies: “No orthogonal validation using sequencing was performed, and therefore the accuracy of variant calls cannot be independently confirmed” (Discussion, Limitations paragraph). The Conclusions also state that future studies should incorporate “sequencing-based validation” (Conclusions).
[Comment 10] Are these variants present in Indonesian population databases?
Response:
At present, comprehensive population-scale genomic databases specific to the Indonesian population remain none. Therefore, population-specific allele frequency validation could not be performed. We have clarified this limitation in the Discussion and highlighted the need for future studies incorporating population-matched reference datasets. The Limitations paragraph notes that “the prioritized variants may represent population-specific rare variants rather than tumor-specific alterations” (Discussion, Limitations paragraph).
[Comment 11] There is limited comparison with current HCC genetics. Not discussed: Major HCC driver genes, such as TP53, CTNNB1, TERT, AXIN1, and ARID1A.
Response:
We have expanded the Discussion to include comparison with known HCC genomic landscapes and to contextualize our findings within current literature. A dedicated paragraph now explicitly addresses the absence of major HCC driver genes (TP53, CTNNB1, TERT, AXIN1, ARID1A) and explains this in the context of platform limitations and sample size. This paragraph reads: “Notably, well-established HCC driver genes such as TP53, CTNNB1, TERT, AXIN1, and ARID1A were not identified in the prioritized dataset. This observation likely reflects the inherent limitations of SNP-array platforms, which are not optimized for comprehensive mutation discovery and may not capture key somatic alterations outside of predefined probe regions” (Discussion).
Reviewer 3 Report
Comments and Suggestions for Authorsthe research topic is very interesting and the presented findings could bring benefit to the scientific community especially in medicine and pathophysiology i.e. better understanding of cancer development and its etiology.
in the introduction the authors explain importance of the study and difficulties that are often connected to the hepatocellur cancers, unknown mechanism, high mortality, short survival time.
the authors pointed out as key topics in difficulties in understanding development of the hepatocellular cancer, as underline it as cancer with highest mortality. they also mentioned geographical regions but focusses on Indonesia. the following paragraph exp[lain genomics and underline its importance in cancers development.
the materials and methods are done in excellent manner and reading it the scientist can make picture of its experimental steps.
the conditions of separation and isolations of DNA was done very clearly and precisely and to avoid any doubt the authors give the scheme of it.
the results are presented in logical manner from general to the specific one. the discussion is in accordance with research framework and pointed out the advantages and drawback of the specific result.
the number of figures and tables is satisfactory and give additional value to the article making it more understandable
the conclusion summarize the finding and underline the most important one.
the references are in accordance with research framework and reader can easily compare data with literature.
But i have few suggestions to authors to improve article quality:
the authors mentioned geographical regions but doesn't compare their findings with patients from other regions.
the scheme or figure of the liver regions could bring benefit in better understanding of research framework, the authors compare affected regions but avoid the fact that not most of the readers doesn't know which part of liver it present.
the page numeration is little bit confusing, and some pages are repeated with different content please clarify that potentially confusing spot.
what is the efficiency of DNA extractions and is it accordance with literature data.
the explanation of the table 2 could also improve understanding of the presented data, current presentation led to conclusion that region III has high influence on mortality! please clarify!
please clarify the section 6. its confusing does itexplain the p[atents or authors contribution.
Author Response
[General Comment]
The research topic is very interesting and the presented findings could bring benefit to the scientific community especially in medicine and pathophysiology. The materials and methods are done in excellent manner. The results are presented in logical manner. The discussion is in accordance with research framework.
Response:
We sincerely thank the Reviewer for the positive assessment of our study’s topic, methodology, and overall structure. We appreciate the recognition of the potential contribution to the scientific community. Below, we address each specific suggestion.
[Comment 1] The authors mentioned geographical regions but doesn’t compare their findings with patients from other regions.
Response:
We appreciate this suggestion. As this is an exploratory, hypothesis-generating study with a small sample size (n = 11) and tumor-only design, direct cross-regional comparison of variant frequencies is not statistically meaningful in the current dataset. However, we have contextualized our findings within the broader HCC genomic landscape by discussing the absence of canonical HCC driver genes (TP53, CTNNB1, TERT, AXIN1, ARID1A) that are commonly reported in Asian cohorts. We have also emphasized the need for future studies incorporating larger, multicenter cohorts with population-matched reference datasets to enable meaningful geographic comparisons (Discussion, paragraphs on HCC drivers and Limitations).
[Comment 2] A scheme or figure of the liver regions could bring benefit in better understanding of research framework.
Response:
We thank the Reviewer for this suggestion. We acknowledge that a schematic of hepatic segmental anatomy (e.g., Couinaud classification) could aid readers less familiar with liver surgery. However, as our study focuses on genomic profiling rather than surgical anatomy, and the liver segment data in Table 2 serves primarily as descriptive clinical context, we respectfully believe that adding an anatomical figure would fall outside the primary scope of this manuscript. We have ensured that Table 2 clearly presents segmental localization using standard Couinaud nomenclature, which is widely recognized in the hepatobiliary literature.
[Comment 3] The page numeration is a little bit confusing, and some pages are repeated with different content.
Response:
We apologize for the formatting issue in the previous submission. The page numbering has been corrected in the revised manuscript, and we have verified that no pages are duplicated or contain inconsistent content. This issue likely arose from the duplicate Discussion text that was present in the original submission, which has now been removed.
[Comment 4] What is the efficiency of DNA extraction and is it in accordance with literature data?
Response:
DNA extraction was successful in all 15 FFPE samples. The A260/280 purity ratio ranged from 1.73 to 2.07, indicating acceptable purity for downstream genotyping applications. DNA concentration ranged from 0.678 to 37.7 ng/µL. These values are consistent with published literature on FFPE-derived DNA, which commonly reports variable yields and purity due to formalin-induced degradation and cross-linking. We have reported these quality metrics in Section 3.2 (DNA Extraction and SNP-Array Genotyping) of the revised manuscript.
[Comment 5] The explanation of Table 2 could also improve understanding of the presented data; current presentation led to conclusion that region III has high influence on mortality.
Response:
We appreciate this observation. Table 1 (previously Table 2) presents clinicopathological characteristics descriptively, and we do not intend to imply any association between specific liver segments and mortality outcomes. The apparent clustering of deaths in patients with segment III involvement is likely coincidental given the small sample size. We have revised the accompanying text in Section 3.1 (previously Section 3.7) to clarify that clinical observations, including tumor location, are reported descriptively. We explicitly state that no prognostic inference should be drawn from these data given the study’s exploratory design and limited power. (Results, Section 3.1)
[Comment 6] Please clarify Section 6. It is confusing, does it explain the patents or authors’ contribution?
Response:
We thank the Reviewer for identifying this confusing labeling. Section 6 was titled “Patents” following the journal’s default template format; however, there are no patents associated with this study. The content under that heading describes Author Contributions. We have revised the heading to clearly read “Author Contributions” to eliminate confusion (Section 6, revised manuscript).
Reviewer 4 Report
Comments and Suggestions for Authors1.A critical limitation of this study is the absence of matched normal tissue for each tumor sample. As the authors acknowledge, this makes it impossible to distinguish somatic driver mutations from rare, potentially functional germline variants. This ambiguity fundamentally weakens the core conclusion that the prioritized variants, such as NASP rs775916096, are directly implicated in hepatocarcinogenesis. The discussion appropriately notes this limitation, but the interpretation of the results, particularly in the abstract and conclusions, should be tempered even further to emphasize that these candidates could represent rare germline variants enriched in this specific Indonesian cohort rather than bona fide somatic drivers.
2.The manuscript highlights NASP rs775916096 as the top candidate, noting it carries "transcript-dependent consequence annotations that included splice-region and coding/regulatory contexts." This presentation is potentially misleading. A single SNP cannot be a high-impact splice variant, a damaging missense variant, and a regulatory variant simultaneously in the same transcript. This aggregated annotation likely arises from the variant being mapped to different transcripts with different consequences (intronic in one, missense in another). The authors should clarify this in the text and, more importantly, specify which specific molecular consequence (splice donor loss in transcript ENST00000378965) is the most likely driver of its functional weight to provide a clearer, testable hypothesis for validation.
3.The quality control pipeline resulted in the exclusion of 4 out of 15 samples (27%), a relatively high attrition rate. While stringent QC is commendable, the manuscript does not sufficiently discuss the characteristics of the excluded samples. Were these samples from patients with a particular clinical feature (older age, larger tumors, specific risk factor) that might introduce a selection bias into the final analyzed cohort? A comparison of the clinicopathological features (Table 2) between included and excluded cases would help assess whether the final 11-sample dataset remains representative of the original cohort.
4.The "Discussion" section contains a significant structural error, with the entire text duplicated in the manuscript (lines 270-313). Beyond this duplication, the flow of the argument could be improved. The detailed paragraph on the known biology of GRP78 in HCC is well-researched but somewhat disproportionate, as GRP78 was only a secondary candidate with a lower burden score. It would be more impactful to first discuss the potential biological relevance of the top candidate, NASP, in the context of HCC, and then briefly introduce GRP78 as a secondary, but biologically plausible, finding supported by existing literature.
5.There are minor inconsistencies in the data presentation that should be addressed. In Table 1, the SNP column lists identifiers such as "rs775916096;rs753910817" for the gene Y_RNA. It is unclear if this indicates multiple SNPs mapped to that same gene or if it is a formatting error. Please clarify and ensure each gene-variant association is clearly and consistently represented. Additionally, the "Recur-rency" column in Table 2 contains the entry "2 Days (OM)" for patient 8. "OM" (operative mortality) should be explicitly defined in the table legend to ensure clarity for the reader, as this represents a distinct outcome from recurrence.
Author Response
[Comment 1] Absence of matched normal tissue; interpretation should be tempered further, especially in abstract and conclusions.
Response:
We fully agree with this important observation. The revised manuscript now explicitly acknowledges in both the Abstract and Conclusions that the tumor-only design precludes the distinction between somatic mutations and rare germline variants. The Abstract states: “Given the tumor-only design without matched normal tissue, the prioritized variants cannot be distinguished from rare germline variants.” The Conclusions further state: “Given the absence of matched normal tissue, the small sample size, and the lack of sequencing-based validation, these findings should be interpreted with caution and should not be considered evidence of disease-associated driver events. Rather, this study provides a preliminary, population-relevant shortlist of candidate variants for further investigation.” We have also explicitly noted the possibility that the prioritized variants could represent rare germline variants enriched in this specific Indonesian cohort rather than bona fide somatic drivers (Abstract, Discussion Limitations paragraph, and Conclusions).
[Comment 2: The presentation of rs775916096 annotations is potentially misleading; clarify transcript-dependent consequences and specify the most likely driver mechanism]
The manuscript highlights NASP rs775916096 as the top candidate, noting it carries "transcript-dependent consequence annotations that included splice-region and coding/regulatory contexts." This presentation is potentially misleading. A single SNP cannot be a high-impact splice variant, a damaging missense variant, and a regulatory variant simultaneously in the same transcript. This aggregated annotation likely arises from the variant being mapped to different transcripts with different consequences (intronic in one, missense in another). The authors should clarify this in the text and, more importantly, specify which specific molecular consequence (splice donor loss in transcript ENST00000378965) is the most likely driver of its functional weight to provide a clearer, testable hypothesis for validation.
Response:
We appreciate this astute observation. We have revised the Discussion to explicitly clarify that the multi-layered functional annotation of rs775916096 reflects transcript-dependent consequences across multiple isoforms, not simultaneous effects within a single transcript. The revised text now reads: “Importantly, the functional annotation of rs775916096 reflects transcript-dependent consequences across multiple isoforms, including splice-region and coding/regulatory contexts. Such multi-layered annotation is expected in genome-wide analyses and does not imply simultaneous effects within a single transcript. Rather, it suggests that the variant may exert context-specific functional effects, with splice-region disruption representing a plausible primary mechanism that should be prioritized in downstream experimental studies” (Discussion, NASP paragraph). We acknowledge that specifying the exact transcript (e.g., ENST00000378965) with splice donor loss as the most likely primary mechanism would strengthen the testable hypothesis, and we will incorporate this level of transcript-specific detail in the revised version.
[Comment 3] High QC attrition rate (4/15 = 27%); discuss characteristics of excluded samples and potential selection bias.
Response:
We appreciate this important methodological point. The four excluded samples (sample IDs 3, 7, 13, and 15; all male) were removed solely on the basis of low genotyping quality metrics (call rate and P10 GC), which is expected given the inherent variability in FFPE-derived DNA quality. In Table 2 of the revised manuscript, we have retained the clinicopathological data for all 15 patients, including the four excluded cases, allowing readers to compare the characteristics of included and excluded cases. Review of these data indicates that the excluded patients do not cluster around any particular clinical feature (age range: 46–74 years; tumor sizes spanning <5 cm to >10 cm; mixed risk factors; BCLC stages A–C), suggesting that the exclusion was driven by technical DNA quality rather than systematic clinical selection bias. We have added a note in the Results section acknowledging this attrition rate and its implications (Results, Section 3.3).
[Comment 4] Structural error with duplicated Discussion text; GRP78 discussion disproportionate; reorganize to lead with NASP.
Response:
We thank the Reviewer for identifying the text duplication, which has been corrected in the revised manuscript. The Discussion has also been restructured to prioritize the biological relevance of NASP as the top-ranked candidate before introducing GPR78 as a secondary signal. The revised Discussion now follows the order: (1) NASP as the highest-ranked candidate with supporting biological rationale, (2) clarification of transcript-dependent annotation for rs775916096, (3) GPR78 as a secondary candidate with currently limited biological evidence in HCC, (4) absence of known HCC driver genes, (5) methodological perspective, and (6) comprehensive limitations. The GPR78 paragraph has been shortened and appropriately contextualized as a secondary exploratory finding (Discussion, revised manuscript).
[Comment 5] Minor inconsistencies: Y_RNA dual SNP identifiers; Table 2 “OM” definition.
Response:
We thank the Reviewer for these detailed observations. Regarding the Y_RNA entry in Table 1 showing “rs753910817.1;rs753910817,” this represents the same variant with different identifiers mapped to the same genomic position through the annotation pipeline. We have clarified this in the table footnote to prevent confusion. Regarding “OM” (operative mortality) in Table 2, we confirm that this abbreviation is now explicitly defined in the table legend: “OM, operative mortality” (Table 2 footnote, revised manuscript).
Reviewer 5 Report
Comments and Suggestions for AuthorsThe manuscript entitled "An Initial Indonesian Genome-wide SNP Array Study and Functional Variant Prioritization Reveals NASP and GPR78 SNVs in Hepatocellular Carcinoma FFPE" aims to identify the potential SNPs related to HCC in an Indonesian cohort. However, there are important shortcomings about the manuscript as listed below.
1) Why was a healthy control group not included in the study?
2) There are some inconsistencies. The authors say "individual-level genotype identifiers were unavailable" (line 159) but also "SNP-array genotyping generated 175 655,366 markers across 15 samples" (lines 175-176). Was the SNP analysis performed for all patients? Also, it is said that the final 11-sample dataset was used (line 179) but all patients were marked for SNPs in two genes (Table 2).
3) The title should be revised to be clearer and more concise without abbraviations.
4) Abstract contains too much abbraviations without their full name.
5) Why were the DNA samples not examined on agarose gel?
6) The number of patients is very low. A total of 11 samples were used for the final identification. The presence of identified SNPs should be screened in a larger cohort of patients and healthy controls.
7) The order of the subsections in Results is not convenient. First the "Patient Clinical Characteristics" should be given. Then, the DNA isolation part should be prior to "Overview of Workflow and Yield". Also, the "Summary of Findings" as a separate subsection is unnecessary.
8) What do "rs775916096" and "rs558447540" mean? The SNPs should be clear indicating the changed bases and locations. Also, how many polymorphisms were detected in these genes?
9) Line 219: "No high-confidence loss-of-function variants were identified." What were the criteria for a high-confidence loss-of-function variant? How were the prioritization criteria determined?
10) Lines 36 and 309: The statement "relatively prolonged survival" cannot be used as a conclusion in this study.
11) Previous studies about the SNPs related to HCC should be discussed.
12) The manuscript was prepared in such a way that most of its sections can only be understood by a bioinformatics specialist. Many abbraviations lack their full name.
Author Response
[Comment 1] Why was a healthy control group not included in the study?
Response:
This study was designed as an exploratory, pilot analysis to assess the feasibility of genome-wide SNP-array profiling from FFPE-derived tumor tissue in an Indonesian HCC cohort. The primary aim was to demonstrate that functional variant prioritization could be performed on archived tumor specimens, rather than to conduct a case-control association study. We acknowledge that the absence of a healthy control group is a significant limitation that precludes the distinction of somatic from germline variants, and this has been emphasized throughout the revised manuscript. Future studies are planned to incorporate matched normal tissue and healthy controls to enable proper association testing (Discussion, Limitations paragraph; Conclusions).
[Comment 2] There are inconsistencies regarding individual-level genotype identifiers and sample numbers.
Response:
We thank the Reviewer for identifying this confusing wording. To clarify: (1) “Individual-level genotype identifiers were unavailable” refers to the fact that sample-level genotype calls could not be individually exported from the GenomeStudio pipeline for gene-level association testing in the standard GWAS sense, necessitating the case-only burden analysis approach. SNP-array genotyping was indeed performed on all 15 samples, generating 655,366 markers. (2) After quality control, 4 samples were excluded based on genotyping quality metrics, resulting in 11 samples for the downstream analytical pipeline. (3) In Table 2, all 15 patients are presented for completeness of clinical data; the SNP columns (NASP, GPR78) indicate the genotyping status at the array level prior to downstream filtering, not the final analytical dataset. We have revised the relevant text to clarify these distinctions and avoid confusion (Methods and Results sections).
[Comment 3] The title should be revised to be clearer and more concise without abbreviations.
Response:
We have revised the title to: “An Initial Indonesian Genome-Wide SNP-Array Study with Functional Variant Prioritization Reveals NASP and GPR78 Candidate SNVs in FFPE Hepatocellular Carcinoma.” We acknowledge that NASP, GPR78, SNVs, and FFPE remain as abbreviations in the title. Given that these are the specific molecular identifiers central to our findings and are widely recognized in the genomics and pathology literature, we believe their inclusion in the title is appropriate for indexing and discoverability. We will defer to the Editor’s guidance on further title modifications if needed.
[Comment 4] Abstract contains too many abbreviations without their full name.
Response:
We have revised the Abstract to spell out all abbreviations at first use, including HCC (hepatocellular carcinoma), SNP (single-nucleotide polymorphism), FFPE (formalin-fixed paraffin-embedded), VEP (Variant Effect Predictor), and others. We have ensured that the Abstract can be read independently without requiring reference to the main text (Abstract, revised manuscript).
[Comment 5] Why were the DNA samples not examined on agarose gel?
Response:
FFPE-derived DNA is characteristically fragmented due to formalin-induced cross-linking and degradation during tissue processing. Agarose gel electrophoresis of FFPE DNA typically demonstrates a smear of low-molecular-weight fragments rather than high-molecular-weight intact bands, making it a relatively uninformative quality assessment for this sample type. Instead, we utilized NanoDrop spectrophotometry (A260/280 ratio) and Qubit fluorometry for DNA quantification and purity assessment, which are the standard quality control methods recommended for FFPE-derived DNA in the SNP-array genotyping workflow. This approach is consistent with published protocols for Illumina Infinium assays using FFPE material (Methods, DNA Extraction subsection).
[Comment 6] The number of patients is very low (n=11). The presence of identified SNPs should be screened in a larger cohort.
Response:
We fully agree. The small sample size is a key limitation of this exploratory study, which has been explicitly acknowledged throughout the revised manuscript. We have framed the study as a pilot, hypothesis-generating analysis, and the Conclusions now state: “Future studies incorporating larger cohorts, matched tumor–normal analysis, and multi-omic validation will be essential to determine the biological and clinical relevance of these prioritized variants.” The primary purpose of this study is to demonstrate feasibility and generate a shortlist of candidate variants for subsequent validation in appropriately powered cohorts (Discussion and Conclusions).
[Comment 7] The order of the subsections in Results is not convenient. “Summary of Findings” as a separate subsection is unnecessary.
Response:
We appreciate this structural feedback. We acknowledge that the current subsection order (Overview of Workflow → DNA Extraction → QC → Variant Filtering → Gene-Level Burden → Summary → Clinical Characteristics) may not be the most intuitive presentation. We have noted this suggestion and will consider reordering the Results to place “Patient Clinical Characteristics” first, followed by “DNA Extraction,” “Overview of Workflow and Yield,” and then the analytical results. The “Summary of Findings” subsection will be removed as a separate heading, with its content integrated into the preceding sections to reduce redundancy.
[Comment 8] What do “rs775916096” and “rs558447540” mean? The SNPs should be clear indicating the changed bases and locations.
Response:
The “rs” identifiers (e.g., rs775916096, rs558447540) are Reference SNP cluster IDs assigned by NCBI’s dbSNP database, which are the internationally standardized nomenclature for catalogued genetic variants. Each rs number uniquely identifies a specific variant with defined chromosomal position, reference and alternate alleles, and other metadata that can be retrieved from public databases (dbSNP, Ensembl, gnomAD). We have added a clarifying note in the Methods to explain this nomenclature for readers less familiar with genomics terminology. Detailed base changes and genomic coordinates can be cross-referenced through the rs identifiers in public databases.
[Comment 9] What were the criteria for a high-confidence loss-of-function variant? How were the prioritization criteria determined?
Response:
High-confidence loss-of-function (LoF) variants were defined as variants predicted to cause protein truncation or complete loss of gene function, including nonsense (stop-gain) mutations, frameshift insertions/deletions, and canonical splice-site variants disrupting essential donor/acceptor dinucleotides. These categories follow the standard definitions used in population genomics and variant interpretation frameworks (e.g., LOFTEE, gnomAD). No such variants were identified in our filtered dataset, and therefore functional prioritization was focused on the next tiers: predicted splice-disrupting variants, damaging missense variants, and selected regulatory variants. The prioritization criteria were determined a priori based on established in silico prediction tools and thresholds (SIFT, PolyPhen-2, CADD ≥ 20, REVEL ≥ 0.5, SpliceAI ≥ 0.5, MaxEntScan ≤ −3) as described in the Methods (Methods, Weighting Scheme subsection).
[Comment 10] The statement “relatively prolonged survival” cannot be used as a conclusion.
Response:
We agree that this phrasing was inappropriate given the small sample size and descriptive study design. The statement “relatively prolonged survival” has been removed from the revised manuscript. Clinical outcomes are now reported descriptively without implying prognostic conclusions. The revised text describes observed overall survival ranges and censoring status without attributing prognostic significance (Results, Section 3.1; Conclusions).
[Comment 11] Previous studies about the SNPs related to HCC should be discussed.
Response:
We have expanded the Discussion to contextualize our findings within the known HCC genomic landscape. Specifically, we now discuss the major HCC driver genes (TP53, CTNNB1, TERT, AXIN1, ARID1A), explain their absence in our dataset due to platform limitations, and discuss the biological plausibility of NASP (via chromatin regulation and cell cycle pathways) and GPR78 in the context of cancer biology. We have added supporting citations including studies on NASP in melanoma (Li et al., 2019), c-Myc–driven hepatocarcinogenesis (Moon et al., 2021; Sequera et al., 2022), and drug resistance mechanisms (Marin et al., 2020) (Discussion, revised manuscript).
[Comment 12] The manuscript was prepared in such a way that most sections can only be understood by a bioinformatics specialist. Many abbreviations lack their full name.
Response:
We appreciate this feedback regarding accessibility. We have revised the manuscript to spell out all abbreviations at first use in both the Abstract and main text. Additionally, a comprehensive abbreviation table has been included at the end of the manuscript for reader reference. We have also simplified technical language where possible while maintaining scientific accuracy, to ensure the manuscript is accessible to a broader clinical and research audience (throughout the revised manuscript; Abbreviations table).
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors have addressed my main conceptual concerns satisfactorily. I appreciate that the revised manuscript now presents the study more cautiously and appropriately as a small, exploratory, tumor-only analysis, and that the conclusions have been tempered accordingly. The clarification regarding the misuse of the term “GWAS” is also important and improves the manuscript's overall accuracy. In its current form, the paper is substantially improved.
However, before final acceptance, I would still recommend correcting a few remaining minor issues:
The presentation of the NASP and GPR78 variants in Table 1 still requires clarification. As currently shown, these variants appear across essentially all listed cases, including samples excluded after quality control. This should be explained clearly, because in the current form it weakens the interpretation of these variants as prioritized candidates.
The description of the filtering strategy should be made fully consistent across the manuscript. In particular, the role of MAF and HWE in a tumor-only dataset should be clarified and described uniformly in the Methods, Results, and figure legend/workflow.
The meaning of “GQ > 0.9” should be clarified, and the terminology should be checked carefully, as it is currently unclear whether this refers to genotype quality or another metric.
The manuscript would still benefit from one more careful editorial revision, as there are minor repetitions and residual language issues.
Overall, I consider the manuscript much improved, and after these minor technical corrections, it could be accepted.
Comments on the Quality of English LanguageSeveral types and grammatical errors may be detected; the authors should again carefully check their manuscript.
Author Response
[General Comment]
The authors have addressed my main conceptual concerns satisfactorily. I appreciate that the revised manuscript now presents the study more cautiously and appropriately as a small, exploratory, tumor-only analysis, and that the conclusions have been tempered accordingly. The clarification regarding the misuse of the term “GWAS” is also important and improves the manuscript's overall accuracy. In its current form, the paper is substantially improved.
Response: Thank you for your thoughtful and encouraging feedback. We sincerely appreciate your recognition that the revised manuscript has addressed the key conceptual concerns, particularly in presenting the study more appropriately as a small, exploratory tumor-only analysis with more cautious and balanced conclusions. Your constructive comments have been invaluable in strengthening the overall quality and clarity of our work.
Comments 1: The presentation of the NASP and GPR78 variants in Table 1 still requires clarification. As currently shown, these variants appear across essentially all listed cases, including samples excluded after quality control. This should be explained clearly, because in the current form it weakens the interpretation of these variants as prioritized candidates.
Response 1: Thank you for this important comment. We acknowledge that the initial presentation of NASP and GPR78 variants in Table 1, particularly their presence across nearly all samples including those excluded after quality control, may lead to ambiguity in interpreting their relevance as prioritized candidates. Upon further evaluation, this consistent distribution across both included and excluded samples suggests that these variants are unlikely to represent discriminative or disease-specific alterations within our cohort, and may instead reflect recurrent background findings inherent to the dataset. To enhance clarity and ensure a more focused presentation of variants with stronger potential biological and clinical relevance, we have removed NASP and GPR78 from Table 1. This revision allows the table to better emphasize high-confidence variants that meet quality control criteria and are more appropriate for interpretation in the context of hepatocellular carcinoma, thereby improving the overall clarity and robustness of the results.
Comments 2: The description of the filtering strategy should be made fully consistent across the manuscript. In particular, the role of MAF and HWE in a tumor-only dataset should be clarified and described uniformly in the Methods, Results, and figure legend/workflow.
Response 2: Thank you for this valuable comment. We have revised the manuscript to ensure that the description of the filtering strategy is fully consistent across the Methods, Results, and figure legend/workflow. In particular, we have clarified the role of MAF and HWE within the context of our tumor-only design. As now consistently described throughout the manuscript, MAF and HWE were not applied as strict biological filters, given the absence of matched normal tissue and the expected deviation of tumor genomes from population-genetic assumptions; instead, they were used only as supplementary technical heuristics to support data refinement without serving as criteria for variant prioritization. This clarification has been uniformly incorporated across all relevant sections, and the revised text reflects a coherent and methodologically appropriate description of the filtering strategy.
Comments 3: The meaning of “GQ > 0.9” should be clarified, and the terminology should be checked carefully, as it is currently unclear whether this refers to genotype quality or another metric.
Response 3: We thank the reviewer for identifying this ambiguity. We revised the manuscript to remove the term “GQ > 0.9,” as the previous wording could be misleading. The relevant sentence has been clarified to describe this step more generally as additional genotype-level quality filtering applied after VCF conversion.
Author Response File:
Author Response.pdf
Reviewer 4 Report
Comments and Suggestions for AuthorsThe authors have carefully addressed the previous concerns and substantially improved the manuscript. The revised version now presents a clearer rationale, a more transparent analytical workflow, and appropriately tempered interpretations given the exploratory, tumor-only design. The functional prioritization framework is sound, and the limitations are now discussed with appropriate caution. I have no further major concerns. The manuscript meets the standards for acceptance and will serve as a valuable hypothesis-generating resource for HCC genomics in understudied populations.
Author Response
[General Comment]
The authors have addressed my main conceptual concerns satisfactorily. I appreciate that the revised manuscript now presents the study more cautiously and appropriately as a small, exploratory, tumor-only analysis, and that the conclusions have been tempered accordingly. The clarification regarding the misuse of the term “GWAS” is also important and improves the manuscript's overall accuracy. In its current form, the paper is substantially improved.
The authors have carefully addressed the previous concerns and substantially improved the manuscript. The revised version now presents a clearer rationale, a more transparent analytical workflow, and appropriately tempered interpretations given the exploratory, tumor-only design. The functional prioritization framework is sound, and the limitations are now discussed with appropriate caution. I have no further major concerns. The manuscript meets the standards for acceptance and will serve as a valuable hypothesis-generating resource for HCC genomics in understudied populations.
Author Response File:
Author Response.pdf
Reviewer 5 Report
Comments and Suggestions for AuthorsGenerally, the concerns about the previous version of the manuscript have been addressed. The shortcomings about the study were mentioned as the limitations. This can be accepted as a preliminary study.
The title can still be improved such as "An Initial Indonesian Genome-wide SNP Array Study Reveals Variants of NASP and GPR78 in Hepatocellular Carcinoma". FFPE is not necessary in the title.
Author Response
[General Comment]
Generally, the concerns about the previous version of the manuscript have been addressed. The shortcomings about the study were mentioned as the limitations. This can be accepted as a preliminary study. The title can still be improved such as "An Initial Indonesian Genome-wide SNP Array Study Reveals Variants of NASP and GPR78 in Hepatocellular Carcinoma". FFPE is not necessary in the title.
Response: Thank you for your constructive feedback. We appreciate your acknowledgment that the previous concerns have been addressed and that the study is now appropriately presented as a preliminary analysis with its limitations clearly stated. In response to your suggestion, we have revised the title to improve clarity and focus, and have removed the mention of FFPE to make it more concise and aligned with the core findings. The updated title reflects a more appropriate representation of the study scope and emphasis.
Author Response File:
Author Response.pdf