Skip to Content
BiomedicinesBiomedicines
  • Article
  • Open Access

20 April 2026

An Initial Indonesian Genome-Wide SNP-Array Study with Functional Variant Prioritization Reveals NASP and GPR78 Candidate SNVs in Hepatocellular Carcinoma

,
,
,
,
,
and
1
Digestive Surgery Division, Faculty of Medicine, Cipto Mangunkusumo Hospital, Universitas Indonesia, Jakarta 10430, Indonesia
2
Department of Medical Chemistry, Faculty of Medicine, Universitas Indonesia, Jakarta 10430, Indonesia
3
Bioinformatics Core Facilities, Indonesian Medical Education and Medical Research Institute (IMERI), Faculty of Medicine, Universitas Indonesia, Jakarta 10430, Indonesia
*
Author to whom correspondence should be addressed.

Abstract

Background/Objectives: Population-specific genomic data are essential for understanding hepatocellular carcinoma (HCC) biology, particularly in underrepresented regions. This study aimed to perform exploratory single-nucleotide polymorphism (SNP)-array-based profiling of HCC tumor samples from Indonesian patients and to prioritize candidate functional variants using a systematic in silico framework. Methods: This retrospective cross-sectional study included 15 resected HCC cases with available formalin-fixed paraffin-embedded (FFPE) tumor tissue. Genome-wide SNP genotyping was performed using the Illumina Asian Screening Array. Following quality control and filtering, variants were annotated using the Ensembl Variant Effect Predictor. A case-only functional prioritization approach incorporating multiple in silico prediction tools was applied, followed by gene-level burden aggregation. Results: After multistep filtering, 11 samples and 104 prioritized variants were retained for analysis. Variants consisted predominantly of splice-region, missense, and regulatory changes. Gene-level burden analysis identified Nuclear Autoantigenic Sperm Protein (NASP, rs775916096) as the highest-ranked candidate gene, while G protein-coupled receptor 78 (GPR78, rs558447540) emerged as a secondary candidate with predicted functional annotations but currently limited biological evidence in HCC. Given the tumor-only design without matched normal tissue, the prioritized variants cannot be distinguished from rare germline variants. Conclusions: This exploratory SNP-array study provides a hypothesis-generating framework for functional variant prioritization in Indonesian HCC. NASP and GPR78 represent preliminary candidates that require validation in larger cohorts with matched normal tissue and sequencing-based confirmation.

1. Introduction

Hepatocellular carcinoma (HCC) is a biologically heterogeneous malignancy and remains one of the leading causes of cancer-related death worldwide. Treatment outcomes are determined not only by intrinsic tumor biology but also by the extent of preserved hepatic function. In Indonesia, the burden of HCC remains particularly substantial. In 2022, an estimated 23,805 new cases and 23,383 deaths were reported, underscoring the exceptionally high fatality of this malignancy, with mortality nearly approaching incidence [1]. Advances in genomic technologies have substantially improved our current understanding of HCC biology by elucidating key molecular pathways involved in hepatocarcinogenesis, thereby facilitating the identification of candidate biomarkers, therapeutic targets, and potential mechanisms of treatment resistance [2,3,4].
In settings characterized by distinct geographic backgrounds, population structure, and risk factor distributions, locally generated genomic data are essential for hypothesis generation and for informing the design of appropriately powered downstream validation studies [3,5]. Although high-depth sequencing remains the preferred approach for comprehensive mutation discovery, alternative genomic platforms may still yield informative genome-wide data at a more feasible scale. In this regard, genome-wide SNP array genotyping is a practical strategy for capturing tumor-derived variant signals from archived tissue specimens, particularly when integrated with stringent quality control and systematic functional annotation.
In this study, we report genome-wide SNP-array profiles derived from formalin-fixed paraffin-embedded tumor specimens obtained from Indonesian patients who underwent surgical resection for HCC at Dr. Cipto Mangunkusumo Hospital. We further implemented a transparent functional prioritization framework that incorporates established in silico prediction tools to nominate candidate functional variants, followed by case-only gene-level aggregation to rank genes by cumulative predicted deleterious burden. This study is intended as a hypothesis-generating resource rather than a case–control association analysis, with the aim of generating a rational shortlist of candidate genes and variants for subsequent validation using matched normal tissue, sequencing-based confirmation, and correlation with clinicopathologic characteristics and clinical outcomes.

2. Materials and Methods

Study Design
In this retrospective cross-sectional study, we included patients with hepatocellular carcinoma (HCC) who underwent hepatectomy at Dr. Cipto Mangunkusumo Hospital between 2007 and 2023 and for whom archived formalin-fixed paraffin-embedded (FFPE) tumor tissue was available for genomic analysis. Eligibility criteria comprised histopathologic confirmation of HCC in the resection specimen and adequate FFPE material for DNA extraction and SNP array genotyping. Clinical and pathological variables, including suspected etiologic risk factors, Child–Pugh class, and Barcelona Clinic Liver Cancer (BCLC) stage, were systematically abstracted from the medical records. Tumor characteristics were ascertained from operative reports, pathology records, and imaging findings, including anatomical location, maximal tumor diameter, tumor multiplicity, and other relevant morphologic features. Data on recurrence and surveillance were obtained from institutional follow-up records and, where necessary, reconfirmed through structured contact with patients or family members during the data collection period. Clinical data extraction and genome-wide SNP-array genotyping were performed in November 2024. Prior to analysis, all datasets were de-identified, and only anonymized data were accessible to the investigators.
FFPE DNA Extraction & Purification
Genomic DNA was isolated from FFPE tumor tissue using the QIAamp DNA FFPE Tissue Kit (Qiagen, Hilden, Germany) according to the manufacturer’s protocol. For each sample, six 10 µm-thick sections were processed. Following xylene-based deparaffinization, tissue was lysed under denaturing conditions with proteinase K. To reverse formalin-induced cross-linking, lysates were incubated at 90 °C prior to column-based purification. DNA was subsequently captured on a silica membrane, washed to remove residual contaminants, and eluted in 20 µL of the supplied elution buffer. DNA concentration and purity were evaluated before downstream analysis using NanoDrop spectrophotometry and Qubit 4.0 fluorometry.
Because FFPE-derived DNA is frequently limited in quantity and may be contaminated with residual impurities, an additional cleanup and concentration step was performed using the DNA Clean & Concentrator-5 kit (Zymo Research, Irvine, CA, USA). DNA Binding Buffer was added at 2–7 times the sample volume, and the mixture was loaded onto a Zymo-Spin column for centrifugation-based DNA capture. After two washes with DNA Wash Buffer, DNA was eluted in 8 µL of DNA Elution Buffer or nuclease-free water following a brief incubation at room temperature. For SNP-array genotyping, DNA input was normalized to 50 ng per sample whenever possible; when DNA yield was insufficient, the maximum recoverable amount was used. Purified DNA was stored at −20 °C until genotyping.
SNP Array Genotyping
Genome-wide SNP genotyping was conducted using the Infinium Asian Screening Array (ASA) BeadChip (Illumina, San Diego, CA, USA) following the manufacturer’s standard Infinium assay protocol. In brief, genomic DNA was subjected to whole-genome amplification with an incubation at 37 °C for 20–24 h, followed by 1 h of enzymatic fragmentation. The amplified DNA was then precipitated at 4 °C for 30 min and resuspended before hybridization. Samples were hybridized to the BeadChip at 48 °C for 16–24 h. After post-hybridization washing, single-base extension and staining were carried out at 44 °C in accordance with the Infinium workflow. Arrays were scanned using the Illumina iScan platform to generate raw intensity data (IDAT files), which were subsequently used for genotype calling and downstream quality control.
SNP-Array Preprocessing and Quality Control (GenomeStudio and PLINK)
Genome-wide SNP-array data generated from FFPE tumor-derived DNA were processed using a standardized genotyping quality control workflow. Raw intensity (IDAT) files were imported into GenomeStudio v2.0 (Illumina) for genotype calling and initial sample-level assessment. Sample quality was evaluated using the call rate and the P10 GC metric, and the call rate versus P10 GC plot was visually inspected to identify outlier samples. Samples falling outside the main cluster or exhibiting a call rate below 0.80 were excluded. SNP statistics were subsequently recalculated in GenomeStudio, and low-quality markers flagged by the software were removed prior to data export.
GenomeStudio output files were then converted to PLINK binary format (BED/BIM/FAM) using PLINK v1.9.0-b.7.7 with the --make-bed option. Individual- and variant-level missingness were assessed using individual missingness (IMISS) and locus/SNP missingness (LMISS) metrics, and samples or SNPs with missingness greater than 0.20 were excluded (--mind 0.2, --geno 0.2). Analyses were restricted to autosomal SNPs (chromosomes 1–22). To reduce redundancy in downstream gene-level aggregation analyses, linkage disequilibrium pruning was performed using --indep-pairwise 50 5 0.2.
Given the tumor-only design and the absence of matched normal tissue, downstream analyses were structured to minimize reliance on population-genetic assumptions. Accordingly, the core quality control framework was based primarily on genotyping performance metrics, including call rate, missingness, and linkage disequilibrium pruning, whereas population allele frequency filtering was applied principally during the annotation stage using external reference databases. Hardy–Weinberg equilibrium and minor allele frequency filters were evaluated only as supplementary quality control heuristics to identify potential technical artifacts and were not interpreted as biologically expected features of tumor-derived genotypes. Following quality control, genotype data were converted to Variant Call Format (VCF) using GTC2VCF for downstream annotation.
Overall, the dataset comprised 655,366 SNPs before quality control, of which 104 remained after the application of the GenomeStudio- and PLINK-based filtering pipeline.
Annotation and Functional Filtering
Variant annotation was performed using the Ensembl Variant Effect Predictor (VEP; GRCh38, web interface) to assign each SNP to its corresponding gene context and to predict its functional consequence. In this study, specific variants are identified using their Reference SNP cluster ID (rs number), which is the standardized nomenclature assigned by the NCBI dbSNP database indicating the specific chromosomal position and allele changes. The potential deleterious impact of missense variants on protein function was assessed using Sorting Intolerant From Tolerant (SIFT), PolyPhen-2, Combined Annotation Dependent Depletion (CADD), and Rare Exome Variant Ensemble Learner (REVEL), whereas the putative effects of splice-region variants on pre-mRNA splicing were evaluated using SpliceAI and MaxEntScan [6,7,8,9,10]. Population allele frequency data were obtained from the Genome Aggregation Database (gnomAD), and variants were defined as rare when the global allele frequency was <0.01. No putative loss-of-function variants were identified in the analyzed dataset; therefore, functional prioritization was focused on predicted splice-disrupting variants, deleterious missense variants, and selected regulatory variants retained for exploratory hypothesis generation.
Weighting Scheme
Each variant was assigned a functional weight based on predicted molecular impact:
  • Variants with strong evidence of splice disruption, defined as a maximum SpliceAI delta score ≥ 0.5 or a MaxEntScan splice site score difference ≤ −3, were assigned a weight of 3.
  • Damaging missense variants were assigned a weight of 2 and were defined as missense variants with REVEL ≥ 0.5 or CADD PHRED ≥ 20, supported by a deleterious SIFT prediction or a damaging PolyPhen-2 prediction.
  • Regulatory or untranslated region variants, including upstream, promoter-annotated, or untranslated region (UTR) variants, were assigned a weight of 1.
  • Variants not meeting these criteria were assigned a weight of 0.
Gene-Level Burden Analysis
In the absence of individual-level genotype identifiers and an external control cohort, we applied a case-only gene-level burden ranking framework to capture the cumulative predicted functional impact of variants within each gene. For each gene, a raw burden score was calculated by summing the weights of all mapped variants. To stabilize the distribution and reduce the influence of extreme values, the raw burden score was log-transformed as follows:
Log Burden = ln (1 + raw burden score)
Genes were then ranked by their log-transformed burden scores, and candidate genes were defined as those in the upper 5th percentile of the ranked distribution for downstream interpretation and exploratory hypothesis generation.

3. Results

3.1. Patient Clinical Characteristics

A total of 15 patients with HCC who underwent hepatectomy were included in the study. The cohort was predominantly male and exhibited generally preserved hepatic function, with the majority of patients classified as Child–Pugh class A. Tumors involved multiple hepatic segments, most frequently segments II, III, and VIII. Hepatitis B virus infection was the most common underlying risk factor, followed by hepatitis C virus infection, nonalcoholic steatohepatitis, and no clearly documented etiology. Most patients presented with early- to intermediate-stage disease, corresponding to BCLC stage A or B. Tumor size most commonly ranged from 5 to 10 cm, and histopathologic evaluation most often showed moderate tumor differentiation (Table 1).
Table 1. Clinicopathological Characteristics, Demographic Features (Age and Sex), and Selected SNPs in Patients with Hepatocellular Carcinoma.
During follow-up, HCC recurrence was documented in three patients. Overall survival remained consistent throughout the first year after surgery, followed by stepwise declines at approximately 19, 22, 26, and 33 months, corresponding to subsequent deaths. Several patients were censored during follow-up, including long-term survivors with follow-up extending to 41 months. Observed overall survival ranged from 16 to 41 months, and median overall survival was not reached within the study period, indicating that more than half of the cohort remained alive at the last follow-up.

3.2. DNA Extraction and SNP-Array Genotyping

DNA extraction was successful in all 15 FFPE tumor samples. The A260/280 purity ratio ranged from 1.73 to 2.07, and the DNA concentration ranged from 0.678 to 37.7 ng/µL. All samples subsequently underwent SNP-array processing using the Infinium Asian Screening Array (ASA) BeadChip (Illumina), and the arrays were scanned using the Illumina iScan platform. Initial genotype calling in GenomeStudio v2.0 identified 655,366 SNP markers across 15 patients, comprising 12 male and 3 female individuals.

3.3. Overview of Workflow and Yield

Genome-wide SNP-array profiling was performed on 15 FFPE tumor specimens from patients with HCC, yielding 104 prioritized variants after sequential quality control, annotation, and functional filtering. From an initial total of 655,366 SNP markers, a multistep analytical pipeline progressively reduced the dataset to a final subset of variants retained for downstream functional prioritization and case-only gene-level aggregation. A schematic summary of the analytical workflow is provided in Figure 1.
Figure 1. Workflow of SNP genotyping, quality control, and functional prioritization. Sequential filtering and annotation steps reduced the initial variant set to identify prioritized candidate genes based on functional impact and gene-level burden analysis. FFPE, formalin-fixed paraffin-embedded; SNP, single nucleotide polymorphism; MAF, minor allele frequency; LD, linkage disequilibrium; HWE, Hardy–Weinberg Equilibrium; VCF, Variant Call Format; REVEL, Rare Exome Variant Ensemble Learner; CADD, Combined Annotation Dependent Depletion; SIFT, Sorting Intolerant From Tolerant; UTR, untranslated region; NASP, Nuclear Autoantigenic Sperm Protein.
In brief, SNP-array genotyping generated 655,366 markers across 15 samples, corresponding to an overall genotyping rate of 0.84365. After GenomeStudio- and PLINK-based preprocessing and marker-level refinement, 223,222 SNPs were retained. Four low-quality samples (sample IDs 3, 7, 13, and 15) were excluded because of poor technical quality, resulting in a final analytic dataset of 11 samples with an overall call rate of 0.917923. Subsequent VCF conversion and additional genotype-level quality filtering further reduced the dataset to 5991 SNPs, of which 104 variants remained after downstream annotation and functional filtering.

3.4. SNP-Array Quality Control and Sample/Marker Filtering

Marker- and sample-level quality control was performed using GenomeStudio and PLINK, with the principal filtering steps summarized in Figure 2 and Figure 3. Initial quality assessment focused on genotyping performance metrics, including call rate, P10 GC, and sample- and variant-level missingness. Samples falling outside the main call rate versus P10 GC cluster or showing a call rate below 0.80 were excluded in GenomeStudio. In PLINK, samples and SNPs with missingness greater than 0.20 were excluded using the --mind 0.2 and --geno 0.2 thresholds, and analyses were restricted to autosomal SNPs. Linkage disequilibrium pruning (--indep-pairwise 50 5 0.2) was then applied to reduce redundancy in downstream gene-level aggregation analyses.
Figure 2. Distribution of missing genotype data across samples and variants. (a) Histogram showing the proportion of missing genotypes per individual, reflecting sample-level data completeness. (b) Histogram showing the proportion of missing genotypes per single-nucleotide polymorphism (SNP), illustrating variant-level missingness across the dataset. These distributions were used to assess data quality prior to downstream genetic analyses and to inform the selection of appropriate quality control thresholds.
Figure 3. Distribution of variant-level allele frequency and Hardy–Weinberg equilibrium statistics. (a) Histogram of minor allele frequency (MAF) across all single-nucleotide polymorphisms (SNPs), illustrating the allele frequency spectrum of the dataset. (b) Histogram of Hardy–Weinberg equilibrium (HWE) test p-values for all SNPs, used to assess genotype quality and identify variants potentially affected by genotyping error, population stratification, or selection prior to quality control filtering.
Because this study used a tumor-only design without matched normal tissue, population-genetic filters, such as minor allele frequency and Hardy–Weinberg equilibrium, were treated as supplementary technical heuristics and were not interpreted as biologically expected properties of the dataset. After these preprocessing and refinement steps, 223,222 SNPs were retained. Four samples (all male; sample IDs 3, 7, 13, and 15) were excluded because of low genotyping quality. The remaining 11 samples (8 male and 3 female) showed an improved overall call rate of 0.917923, and 5991 SNPs were retained for downstream analysis after final quality filtering.

3.5. Variant-Level Filtering and Functional Prioritization

Following the conversion of genotype data to Variant Call Format (VCF), variants were filtered using a genotype quality threshold of GQ > 0.9, resulting in 104 variants eligible for downstream annotation and functional prioritization. After annotation and functional filtering, most variants were excluded because they were not predicted to have meaningful functional consequences. No high-confidence loss-of-function variants were detected. The retained variants were predominantly splice-region, missense, and regulatory or untranslated-region variants (Table 2), which were subsequently prioritized using the prespecified weighting framework for downstream gene-level aggregation.
Table 2. Functional Annotation of Candidate Genes Based on Missense, Splicing, and Regulatory Variants.

3.6. Gene-Level Burden Ranking and Key Candidates

Gene-level aggregation of weighted variants demonstrated a sparse distribution of predicted functional burden across the analyzed genes. Nuclear Autoantigenic Sperm Protein (NASP) was the highest-ranked gene and surpassed the prespecified upper 5th percentile threshold, predominantly driven by rs775916096, which satisfied the criteria for predicted functional significance within the splice/missense/regulatory prioritization framework. G protein-coupled receptor 78 (GPR78) (rs558447540) emerged as a secondary candidate gene, with predicted functional annotations in the missense and regulatory categories, although its cumulative burden score was lower than that observed for NASP. Taken together, the case-only gene-level burden analysis prioritized NASP (rs775916096) as the principal candidate gene and GPR78 (rs558447540) as a secondary candidate for downstream validation and biological interpretation.
Overall, the case-only gene-level burden analysis prioritized NASP (rs775916096) as the leading candidate gene, demonstrating the highest cumulative predicted functional burden across the analyzed dataset. This signal was primarily driven by variants classified as splice-related, missense, and regulatory. GPR78 (rs558447540) was identified as a secondary candidate gene, with predicted functional variants in the missense and regulatory categories, although its aggregate burden was lower than that observed for NASP.

4. Discussion

In this exploratory study, we performed genome-wide SNP-array profiling of FFPE-derived tumor DNA from an Indonesian HCC cohort and applied a structured functional prioritization framework to identify candidate variants [8,9]. From an initial set of 655,366 SNP markers, a multistep pipeline integrating genotyping quality control, variant-level filtering, and in silico annotation yielded 104 prioritized variants across 11 analyzable cases. Gene-level aggregation revealed a sparse distribution of predicted functional burden, highlighting a small number of candidate genes for downstream interpretation.
Among these, NASP emerged as the highest-ranked candidate gene, driven primarily by rs775916096. NASP encodes a histone-binding protein involved in chromatin assembly, cell cycle progression, and DNA replication. Dysregulation of chromatin organization is a well-recognized hallmark of cancer, and accumulating evidence suggests that aberrant histone chaperone activity may contribute to oncogenic transcriptional programs and genomic instability. Although NASP has not been extensively characterized in hepatocellular carcinoma, prior studies in other malignancies and experimental models suggest a potential role in promoting cellular proliferation and tumor growth. In this context, the prioritization of NASP in our dataset provides a biologically plausible signal that warrants further functional validation [11,12,13,14].
Importantly, the functional annotation of rs775916096 reflects transcript-dependent consequences across multiple isoforms, including splice-region and coding/regulatory contexts. Such multi-layered annotation is expected in genome-wide analyses and does not imply simultaneous effects within a single transcript. Rather, it suggests that the variant may exert context-specific functional effects, with splice-region disruption representing a plausible primary mechanism that should be prioritized in downstream experimental studies.
A secondary candidate signal was identified in GPR78 (rs558447540), supported by predicted missense and regulatory annotations. GPR78 is an orphan G protein-coupled receptor with limited functional characterization, and its role in hepatocellular carcinoma remains largely unexplored. Unlike canonical HCC-associated genes, there is currently no substantial evidence linking GPR78 to hepatocarcinogenesis, tumor progression, or therapeutic response [15]. Therefore, this finding should be interpreted cautiously as a novel, hypothesis-generating signal, rather than as a biologically established contributor to HCC. Future studies integrating transcriptomic, proteomic, and functional assays will be necessary to determine whether GPR78 plays a meaningful role in liver tumor biology.
Notably, well-established HCC driver genes commonly reported in Asian cohorts, such as Tumor Protein 53 (TP53), Catenin Beta 1 (CTNNB1), Telomerase Reverse Transcriptase (TERT), Axis Inhibition Protein 1 (AXIN1), and AT-rich interacting domain-containing protein 1A (ARID1A) were not identified in the prioritized dataset [16]. This observation likely reflects the inherent limitations of SNP-array platforms, which are not optimized for comprehensive mutation discovery and may not capture key somatic alterations outside of predefined probe regions. In addition, the relatively small sample size further reduces the likelihood of detecting recurrent driver events. These factors underscore that the present study is not designed to identify canonical drivers, but rather to generate a prioritized shortlist of candidate variants for subsequent investigation.
From a methodological perspective, the use of SNP-array genotyping in FFPE-derived tumor samples represents a pragmatic approach in settings where sequencing resources may be limited and DNA quality is suboptimal. While next-generation sequencing provides higher resolution, SNP-array platforms can still yield informative genome-wide signals when combined with rigorous quality control and systematic functional annotation [17]. In this study, the integration of multiple prediction tools (SIFT, PolyPhen-2, CADD, REVEL, SpliceAI, and MaxEntScan) within a transparent weighting framework allowed for structured prioritization of variants based on predicted functional impact [6,7,8,9,10]. However, it should be emphasized that such in silico predictions are probabilistic and require experimental validation.
The clinicopathological profile of patients in this series generally reflected typical hepatocellular carcinoma characteristics, including a predominance of hepatitis B infection, preserved liver function, and intermediate to large tumor size at presentation. Nevertheless, the small sample size limits the interpretability of these findings, and the results should be considered descriptive only without statistical or population-level significance.
Several important limitations should be acknowledged. First, the tumor-only design without matched normal tissue precludes the distinction between somatic mutations and rare germline variants. As a result, the prioritized variants may represent population-specific rare variants rather than tumor-specific alterations. Second, the small sample size (n = 11 after quality control) limits the robustness and generalizability of the findings and increases the risk of stochastic signals. Third, SNP-array platforms provide limited genomic coverage compared with sequencing-based approaches and may fail to detect key driver mutations. Fourth, no orthogonal validation using sequencing was performed, and therefore the accuracy of variant calls cannot be independently confirmed. Finally, clinical correlation analyses in this cohort remain descriptive and should not be interpreted as evidence of prognostic association.
Taken together, this study should be viewed as a pilot, hypothesis-generating analysis that demonstrates the feasibility of SNP-array-based functional variant prioritization in an Indonesian HCC cohort. The identification of NASP as a top-ranked candidate and GPR78 as a secondary exploratory signal provides a focused starting point for future studies. Subsequent work should incorporate larger, multicenter cohorts, matched normal tissue for somatic variant calling, sequencing-based validation, and integration with transcriptomic and protein-level data to establish biological relevance.

5. Conclusions

This study demonstrates the feasibility of applying a SNP-array–based functional variant prioritization framework to FFPE-derived tumor samples from an Indonesian hepatocellular carcinoma cohort. Within this exploratory, tumor-only design, NASP (rs775916096) emerged as the top-ranked candidate gene, while GPR78 (rs558447540) was identified as a secondary, hypothesis-generating signal with currently limited biological evidence in HCC.
Given the absence of matched normal tissue, the small sample size, and the lack of sequencing-based validation, these findings should be interpreted with caution and should not be considered evidence of disease-associated driver events. Rather, this study provides a preliminary, population-relevant shortlist of candidate variants for further investigation.
Future studies incorporating larger cohorts, matched tumor-normal analysis, and multi-omic validation will be essential to determine the biological and clinical relevance of these prioritized variants. Overall, this work represents an initial step toward generating population-specific genomic insights into Indonesian HCC and provides a foundation for future precision oncology research in this setting.

Author Contributions

T.J.M.L. and V.M.G.M. jointly conceptualized and designed the study, with T.J.M.L. providing overall project leadership and supervision. V.M.G.M. contributed to clinical data collection and curation, interpretation of findings, and manuscript preparation and revision. L.E. performed formal analysis and software-based data processing. F.F. and A.F.P. conducted laboratory procedures and contributed to the investigation and methodology implementation. N.J.Z. and K.N.L.A. supported visualization and contributed to drafting and editing the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

This study was supported by a research grant (PUTI) for the year 2022 from Universitas Indonesia, with contract number NKB-441/UN2.RST/HKP.05.00/2022. The funders had no role in the study design, data collection, data analysis, decision to publish, or manuscript preparation.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board (or Ethics Committee) of the Cipto Mangunkusumo Hospital, Faculty of Medicine, Universitas Indonesia, under approval number KET-399/UN2.F1/ETIK/PPM.00.02/2023.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Acknowledgments

The authors would like to express their sincere gratitude to all individuals and organizations that contributed to the surgical procedures and data acquisition. We also acknowledge the assistance of GPT-5.2 for optimizing the grammar and refining the manuscript. The authors take full responsibility for the content and the final version of this manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ARID1AAT-rich interacting domain−containing protein 1A
ASAAsian Screening Array
AXIN1Axis Inhibition Protein 1
BCLCBarcelona Clinic Liver Cancer
CADDCombined Annotation Dependent Depletion
CTNNB1Catenin Beta 1
FFPEFormalin-Fixed Paraffin-Embedded
gnomADGenome Aggregation Database
GPR78G protein-coupled receptor 78
GQGenotype-Quality
HCCHepatocellular Carcinoma
HWEHardy–Weinberg Equilibrium
IMISSIndividual Missingness
LDLinkage Disequilibrium
LMISSLocus/SNP Missingness
MAFMinor Allele Frequency
NASPNuclear Autoantigenic Sperm Protein
NF-κBNuclear Factor Kappa-Light-Chain-Enhancer of Activated B Cells
P10 GC10% GenCall
REVELRare Exome Variant Ensemble Learner
SIFTSorting Intolerant From Tolerant
SNPSingle Nucleotide Polymorphism
SNVSingle Nucleotide Variant
TERTTelomerase Reverse Transcriptase
TP53Tumor Protein 53
UTRUntranslated Region
VEPVariant Effect Predictor
VCFVariant Call Format

References

  1. Bray, F.; Laversanne, M.; Sung, H.; Ferlay, J.; Siegel, R.L.; Soerjomataram, I.; Jemal, A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin. 2024, 74, 229–263. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Teufel, A.; Staib, F.; Kanzler, S.; Weinmann, A.; Schulze-Bergkamen, H.; Galle, P.R. Genetics of hepatocellular carcinoma. World J. Gastroenterol. 2007, 13, 2271–2282. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Galun, D.; Mijac, D.; Filipovic, A.; Bogdanovic, A.; Zivanovic, M.; Masulovic, D. Precision Medicine for Hepatocellular Carcinoma: Clinical Perspective. J. Pers. Med. 2022, 12, 149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Miki, D.; Ochi, H.; Hayes, C.N.; Aikata, H.; Chayama, K. Hepatocellular carcinoma: Towards personalized medicine. Cancer Sci. 2012, 103, 846–850. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Tümen, D.; Heumann, P.; Gülow, K.; Demirci, C.-N.; Cosma, L.-S.; Müller, M.; Kandulski, A. Pathogenesis and Current Treatment Strategies of Hepatocellular Carcinoma. Biomedicines 2022, 10, 3202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Ioannidis, N.M.; Rothstein, J.H.; Pejaver, V.; Middha, S.; McDonnell, S.K.; Baheti, S.; Musolf, A.; Li, Q.; Holzinger, E.; Karyadi, D.; et al. REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants. Am. J. Hum. Genet. 2016, 99, 877–885. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Kircher, M.; Witten, D.M.; Jain, P.; O’Roak, B.J.; Cooper, G.M.; Shendure, J. A general framework for estimating the relative pathogenicity of human genetic variants. Nat. Genet. 2014, 46, 310–315. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Jaganathan, K.; Kyriazopoulou Panagiotopoulou, S.; McRae, J.F.; Darbandi, S.F.; Knowles, D.; Li, Y.I.; Kosmicki, J.A.; Arbelaez, J.; Cui, W.; Schwartz, G.B.; et al. Predicting Splicing from Primary Sequence with Deep Learning. Cell 2019, 176, 535–548.e24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Naito, T. Predicting the impact of single nucleotide variants on splicing via sequence-based deep neural networks and genomic features. Hum. Mutat. 2019, 40, 1261–1269. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Yeo, G.; Burge, C.B. Maximum Entropy Modeling of Short Sequence Motifs with Applications to RNA Splicing Signals. J. Comput. Biol. 2004, 11, 377–394. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Li, J.-X.; Wei, C.-Y.; Cao, S.-G.; Xia, M.-W. Elevated nuclear auto-antigenic sperm protein promotes melanoma progression by inducing cell proliferation. OncoTargets Ther. 2019, 12, 2105–2113. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Moon, H.; Park, H.; Ro, S.W. c-Myc-driven Hepatocarcinogenesis. Anticancer. Res. 2021, 41, 4937–4946. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Sequera, C.; Grattarola, M.; Holczbauer, A.; Dono, R.; Pizzimenti, S.; Barrera, G.; Wangensteen, K.J.; Maina, F. MYC and MET cooperatively drive hepatocellular carcinoma with distinct molecular traits and vulnerabilities. Cell Death Dis. 2022, 13, 994. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Marin, J.J.G.; Macias, R.I.R.; Monte, M.J.; Romero, M.R.; Asensio, M.; Sanchez-Martin, A.; Cives-Losada, C.; Temprano, A.G.; Espinosa-Escudero, R.; Reviejo, M.; et al. Molecular Bases of Drug Resistance in Hepatocellular Carcinoma. Cancers 2020, 12, 1663. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. GPR78. Expression of GPR78 in Liver Cancer—The Human Protein Atlas [Internet]. Available online: https://v18.proteinatlas.org/ENSG00000155269-GPR78/pathology/tissue/liver%2Bcancer (accessed on 23 March 2026).
  16. Bold, N.; Buyanbat, K.; Enkhtuya, A.; Myagmar, N.; Batbayar, G.; Sandag, Z.; Damdinbazar, D.; Oyunbat, N.; Boldbaatar, T.; Enkhbaatar, A.; et al. High-frequency mutations in TP53, AXIN1, CTNNB1, and KRAS, and polymorphisms in JAK1 genes among Mongolian HCC patients. Cancer Rep. 2025, 8, e70227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. McDonough, S.J.; Bhagwate, A.; Sun, Z.; Wang, C.; Zschunke, M.; A Gorman, J.; Kopp, K.J.; Cunningham, J.M. Use of FFPE-derived DNA in next generation sequencing: DNA extraction methods. PLoS ONE 2019, 14, e0211400. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.