Next Article in Journal
Impact of Bevacizumab Relative Dose Intensity and Treatment-Induced Proteinuria on the Efficacy of Trifluridine/Tipiracil Plus Bevacizumab in Metastatic Colorectal Cancer
Previous Article in Journal
Preclinical Optimization of Magnetotactic Bacteria Therapy for the Treatment of Pancreatic and Rectal Cancer
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Rare Germline Variants in Key Pathways Contribute to Hepatocellular Carcinoma Risk in Hispanic Individuals

1
Section of Epidemiology and Population Sciences, Department of Medicine, Baylor College of Medicine, Houston, TX 77030, USA
2
Dan L Duncan Comprehensive Cancer Center, Department of Medicine, Baylor College of Medicine, Houston, TX 77030, USA
3
Division of Digestive and Liver Diseases, Department of Medicine, UT Southwestern Medical Center, Dallas, TX 75390, USA
4
Section of Gastroenterology and Hepatology, Department of Medicine, Baylor College of Medicine, Houston, TX 77030, USA
5
Center for Innovations in Quality, Effectiveness and Safety (IQuESt), Michael E. DeBakey Veterans Affairs Medical Center, Houston, TX 77030, USA
6
Division of Hepatology, Baylor University Medical Center, Dallas, TX 75246, USA
*
Author to whom correspondence should be addressed.
Cancers 2026, 18(15), 2428; https://doi.org/10.3390/cancers18152428
Submission received: 17 June 2026 / Revised: 17 July 2026 / Accepted: 24 July 2026 / Published: 28 July 2026
(This article belongs to the Section Cancer Epidemiology and Prevention)

Simple Summary

In this study, we performed whole-exome sequencing on 719 Hispanic individuals (455 HCC cases) and conducted a series of rare variant analyses. We identified 29 rare deleterious variants in known HCC genes that are involved in lipid metabolism, immune regulation, and Wnt signaling. These variants are enriched in Hispanics, highlighting population-specific genetic susceptibility to HCC.

Abstract

Background: Hepatocellular carcinoma (HCC) is one of the leading causes of cancer-related death worldwide. In the U.S., HCC incidence is highest among Hispanic populations. With the exception of well-known PNPLA3 variants, the genetic risk factors underlying HCC in Hispanic individuals, particularly rare germline variants, are largely unexplored. Methods: We performed whole-exome sequencing in 719 Hispanic Americans, including 455 HCC cases and 264 controls. We analyzed rare to low frequency, putatively deleterious variants in 86 genes previously implicated in HCC susceptibility. Genetic associations were assessed at single-variant, gene, and cumulative multi-variant burden levels using Fisher’s exact tests and weighted burden tests. Results: We identified 29 rare deleterious variants across 16 genes enriched in HCC cases compared with controls. Single-variant tests identified statistically robust associations in pathways related to lipids metabolism, TM6SF2 R138W, immune regulation, HLA-DRB1 W38X, and Wnt signaling pathway, WNT9A R375H. Exploratory analyses showed that TM6SF2 R138W was further enriched among cases with alcohol and metabolic dysfunction-associated steatotic liver disease. Gene-based burden tests revealed five associated genes, including TM6SF2 and HLA-DRB1. A strong dose-effect relationship was observed, with HCC risk increasing ~4-fold per candidate variant carried (odds ratio, 3.84; 95% confidence interval, 2.40–6.53). Conclusions: Rare deleterious germline variants across lipid metabolism and immune pathways contribute to increased HCC risk in Hispanic individuals. These findings provide a genetic framework for understanding HCC disparities and support future mechanistic and epidemiologic studies in this high-risk population.

1. Introduction

Liver cancer is the sixth most common cancer and the third leading cause of cancer-related deaths worldwide [1]. Hepatocellular carcinoma (HCC) accounts for approximately 90% of primary liver cancers. Major risk factors for HCC include hepatitis B (HBV) and hepatitis C (HCV) virus infections, heavy alcohol consumption, and metabolic disorders [2]. Over the past decade, metabolic dysfunction-associated steatotic liver disease (MASLD), with and without alcohol use, has become an increasingly dominant driver of new HCC cases [3,4]. However, these risk factors explain a variable and low fraction (~10–~45%) of the overall HCC burden, and most individuals with these risk factors do not develop HCC [5]. Evidence of family clustering [6] and synergistic effects among alcohol consumption, metabolic disorders, and genetic predisposition suggest that inherited susceptibility plays a critical role in HCC risk [7].
Hispanic Americans, the largest and fastest-growing minoritized population in the United States, have the highest HCC age-adjusted incidence rates among all racial/ethnic groups [8]. Although this excess burden is partly attributable to a higher burden of metabolic disorders [9], alcoholic liver disease (ALD) and MASLD [10], HCC risk remains elevated even after adjusting for known risk factors, highlighting the potentially important contribution of genetic factors for HCC among Hispanic individuals [11].
Previous genome-wide association studies (GWAS) have identified multiple susceptibility loci and genes for HCC. These include common variants in genes such as Myelin Associated Oligodendrocyte Basic Protein (MOBP), MAU2 Sister Chromatid Cohesion Factor (MAU2) [12], Patatin-like Domain 3, 1-acylglycerol-3-phosphate O-Acyltransferase (PNPLA3) [13], Transmembrane 6 Superfamily Member 2 (TM6SF2) [14], Human Leukocyte Antigen, class II DQ beta 1 (HLA-DQB1) [15], class I polypeptide-related sequence A (MICA) [16], Telomerase Reverse Transcriptase (TERT) [17], Wnt Family Member 3A (WNT3A), Wnt Family Member 9A (WNT9A) [18], Complement C2 (C2) [19] and Sorting and Assembly Machinery component (SAMM50) [20]. However, these GWAS were conducted in mostly European and East Asian populations. As a result, the genetic architecture of HCC in Hispanic populations remains poorly characterized. Moreover, these GWAS studies focused on common variants (minor allele frequency [MAF] > 5%), which are predominantly located in non-coding regions and generally confer a small effect on HCC risk [13,14]. Given that these common variants explain only a limited proportion of HCC heritability (<30% [17,21]), particularly in ancestrally diverse populations, there is a critical need for studies focused on Hispanic groups.
Rare (MAF < 1%) and low-frequency (1% < MAF < 5%), coding variants have the great potential to explain the missing heritability [22], and often have a stronger effect size [23], as demonstrated for BRCA1/2 in breast cancer [24]. Yet, rare variants have not been systematically studied in HCC, and few studies involved Hispanic or Latin American populations. One prior effort, conducted in a Mexican cohort, evaluated six previously reported common variants in PNPLA3, Glucokinase Regulator (GCKR) and Membrane Bound Acylglycerophosphatidylinositol O-acyltransferase MBOAT7 (MBOAT7) with HCC risk [25]. However, this study did not assess rare coding variation and was limited to only three known loci. Thus, despite this initial work, systematic investigation of rare germline variants contributing to HCC risk in Hispanic individuals is still lacking, leaving a major gap in understanding population-specific genetic susceptibility.
In high-throughput sequencing-based rare variant association analyses, the statistical power is often limited due to the low-to-rare allele frequencies of individual variants [26]. However, genes implicated by common causal variants may also harbor rare variants that contribute to disease susceptibility [23,27], a pattern that demonstrated in liver cirrhosis [28]. Thus, focusing on HCC-associated genes may increase statistical power in rare variant analysis. In this study, we performed whole-exome sequencing for 719 Hispanic individuals and examined rare deleterious variants within genes previously implicated in HCC. We hypothesized that rare deleterious variants in key hepatic pathways contribute to HCC susceptibility in Hispanic individuals and that their effects are larger in this high-risk population.

2. Materials and Methods

2.1. Study Participants and Measurements

We conducted a case–control study of HCC among Hispanic individuals in Texas. HCC cases were recruited from clinical sites in Texas as part of the Texas Hispanic HCC collaborative. HCC diagnoses were confirmed by the tumor board at each facility using the American Association for the Study of Liver Diseases (AASLD) criteria, which includes histological and/or radiological diagnosis using characteristic appearance (arterial enhancement and delayed washout) on triple-phase computed tomography (CT) or magnetic resonance imaging (MRI). Internal controls were selected from the Starr County cohort study [29], a population-based cohort of individuals from Starr County, Texas, and had no cancer history. Demographic and clinical variables, such as age (at diagnosis for HCC cases or enrollment for controls), sex, and self-reported race/ethnicity, were obtained from medical records or the study questionnaire. Based on the clinical assessment of the treating physician, HCC etiology was classified as: (i) HCV, if the participant had a positive HCV RNA test or achieved sustained virological response at diagnosis; (ii) HBV, if hepatitis B surface antigen was positive; (iii) ALD based on a clinician-recorded diagnosis of ALD and self-reported current or former heavy alcohol use (≥8 drinks/week for women or ≥15 drinks/week for men); or (iv) MASLD requiring documented hepatic steatosis on liver histology or imaging in the absence of viral hepatitis, ALD, or other clinician-documented etiologies. Viral hepatitis was defined as either HCV or HBV.

2.2. Candidate Gene Selection

We selected genes that had been previously associated with HCC (Experimental Factor Ontology [EFO_0000182]) or MASLD (formerly NAFLD, EFO_0003095) through the GWAS Catalog [30] (accessed 4 August 2025). Additional candidate genes were identified from published GWAS or germline mutation studies identified through a literature review in PubMed. The final curated 86 candidate genes and supporting references are provided in Supplementary Table S1.

2.3. Whole-Exome Sequencing

Whole-exome sequencing of HCC cases and internal controls was performed at the Human Genome Sequencing Center (HGSC) at Baylor College of Medicine. Genomic DNA was extracted from peripheral blood samples that passed intake quality control using the Quant-iT™ PicoGreen® dsDNA Assay Kit (Cat. #P7589, Thermo Fisher Scientific, Waltham, MA, USA). Libraries were prepared using NEB End Repair/dA-tailing and Ligation modules (Cat. # E6050L & E6053L, New England Biolabs, Ipswich, MA, USA) with 500 ng of genomic DNA input. Briefly, DNA samples were fragmented by sonication on a Covaris E220 instrument (Covaris, Woburn, MA, USA) in a 96-well format. Paired-end pre-capture libraries were generated using a fully automated workflow on Beckman Coulter Biomek FXp dual-arm robotic systems (Beckman Coulter, Inc., Stockholm, Sweden). Library preparation included DNA end repair, 3′-adenylation, and ligation to Illumina dual-indexed multiplexing adapters (Cat. #20022370, Illumina, Inc., San Diego, CA, USA). Adapter-ligated DNA fragments were subsequently PCR-amplified using KAPA HiFi DNA polymerase (Cat. #KK2612, Roche, Indianapolis, IN, USA) with primers complementary to the adapter sequences. For exome enrichment, pre-capture libraries were pooled and hybridized to HGSC-ClinExD probes (Twist Bioscience, South San Francisco, CA, USA, custom design), which target the GENCODE and RefSeq gene models and cover 35.9 Mb across approximately 20,000 genes. Post-capture libraries were subsequently combined into a single 90-plex pool for sequencing on one Illumina NovaSeq X lane (Illumina, Inc., San Diego, CA, USA) to generate 150 bp paired-end sequencing reads. Real-Time Analysis software (version 4.6.7) was used to monitor run performance, including cluster density, signal intensity, and phasing/pre-phasing metrics. Whole-exome sequencing was performed once for each individual, and no technical replicate sequencing was conducted. The sequencing coverage and quality statistics for each sample are summarized in Supplementary Table S2.

2.4. Alignment, Variant Calling, and Quality Control

Raw sequencing reads were aligned to the human reference genome (GRCh38) using BWA-MEM (version 0.7.15) [31]. Post-alignment processing, including base quality score recalibration and indel realignment, was performed using the Genome Analysis Toolkit (GATK, version 3.6.0) [32]. Single Nucleotide Variants (SNVs), short insertions or deletions (indels), were called using GATK HaplotypeCaller followed by joint genotyping with GenotypeGVCFs. Variant-level quality control steps included: (1) retaining variants with call rates ≥ 85%; (2) excluding variants with quality score < 30, depth < 10 or allelic balance < 2.0 for heterozygous calls; (3) excluding long indels (≥22 bp) and singleton variants; and (4) removing low-complexity repeats, segmental duplications and low-confidence variants that result from mapping errors, strand bias, weak exon conservation, and noisy background. All retained variants were visualized in Integrative Genome Viewer [33] to verify read depth and the pile-up pattern supporting each variant. Sample-level quality control steps included excluding samples with principal component analysis outliers, abnormal heterozygosity rate, sex discordance, <95% sample completion rates, or unexpected relatedness (defined as identity-by-state > 10%).

2.5. Variant Annotation and Filtering Process

We focused on the variants mapped on the 86 candidate genes. Variants were annotated using Ensembl Variant Effect Predictor [34] (version 114) and filtered based on MAF and functionality as follows (Figure 1): (1) variants with MAF ≥ 1% in UCSC dbSNP tracks (build 155) were excluded; (2) common variants with MAF ≥ 5% in the Admixed American (AMR, n = 30,019) population of Genome Aggregation Database [35] (gnomAD, v4.1.0) were excluded to avoid variants that are rare in the general population but with low or common MAF in Hispanic ancestry groups. This genetic ancestry-defined AMR population includes a substantial proportion (41.76%) of individuals who self-identified as Mexican, Costa Rican, or Latino, making it an appropriate reference population for variant filtering and downstream analyses; (3) variants with MAF ≥ 1% among internal controls were further excluded to eliminate cohort-specific common variants; and (4) variants with a scaled Combined Annotation Dependent Depletion (CADD) [36] score ≥ 15 were retained.

2.6. Variant Prioritization

To streamline downstream analyses, we implemented an integrated variant prioritization framework combining functional consequences, in silico deleteriousness, and CADD score. Variants were first grouped into three tiers: (i) high—protein-truncating LoF variants (such as frameshift or stop-/gain-loss) and missense variants predicted damaging by both Polymorphism Phenotyping v2 (PolyPhen-2) and Sorting Intolerant From Tolerant (SIFT) [37,38]; (ii) medium—missense variants predicted damaging by either tool; and (iii) low—missense variants not predicted damaging but retained based on rarity and CADD. Across tiers, variants with CADD ≥ 15 (top ~3% deleterious in the human genome) were retained, except protein-truncating variants, which were included regardless. Variants were then ranked hierarchically by functional consequence (truncating > missense), in silico predicted deleteriousness tier (high > medium > low), and CADD score (higher prioritized). This integrated approach ensured that variants most likely to disrupt gene function were consistently prioritized for downstream analyses, while integrating functional and computational evidence.

2.7. Single-Variant Tests

MAFs were calculated separately for HCC cases and internal controls. We also included 30,019 individuals from the gnomAD AMR population as an external reference control. Single variant allelic-association tests were performed using Fisher’s exact test for two comparisons: (1) HCC cases versus internal controls (from the Starr County cohort study), and (2) HCC cases versus external gnomAD AMR population controls. Odds ratios (ORs) and 95% confidence intervals (CIs) were computed, and the false discovery rate (FDR) method was applied for multiple testing correction [39]. Stratified analysis was performed to compare the distribution of liver disease etiologies among HCC cases for each individual variant using Fisher’s exact tests. Statistical power for each single-variant comparison was calculated for Fisher’s exact test using the observed allele counts and MAFs in HCC cases, internal controls, and external gnomAD AMR controls. The estimated power for each comparison is summarized in Supplementary Table S3. As a sensitivity analysis, Firth’s penalized logistic regression was additionally performed for comparisons between HCC cases and internal controls, with adjustment for sex and the top three principal components generated by PLINK 2.0 [40] to adjust for population stratification.

2.8. Burden Test of Multiple Rare Deleterious Variants

Gene-based burden tests were performed for genes containing ≥ 2 rare variants using the Sequence Kernel Association Test [41] (SKAT version 2.2.5) R package to evaluate the cumulative effects of rare variants on HCC risk. The weighted burden test first collapses all included variants within a gene with their respective pre-assigned weights into a single burden score and then tests its association with HCC in analyses comparing HCC cases versus internal controls and adjusted for sex. The default Beta (1,25) probability distribution function was used for the weight in the models, which places greater weight on rarer variants. To ensure precise p-value estimation, 10,000 resampling iterations were performed. Sex was included as a covariate in all gene-based burden tests. Multiple testing correction was performed using the FDR method.

2.9. Dose-Effect Analysis

Dose–response effects of rare candidate variants on HCC risk were evaluated by comparing variant carriers and non-carriers in HCC cases and our internal controls using logistic regression models. We fitted logistic regression models for males and females combined and separately, treating the number of prioritized variants as a continuous predictor to estimate the change in HCC risk per additional rare variant carried. The Cochran–Armitage test [42] was used to test for trend.

3. Results

3.1. Overview of Study Cohort and Baseline Characteristics

After quality control, a total of 719 self-reported Hispanic participants were included in the analysis, comprising 455 HCC cases and 264 internal controls (Table 1). Among the HCC cases, 66.4% were male. The leading HCC etiologies were viral hepatitis (29.9%), ALD (23.1%), and MASLD (22.4%).

3.2. Rare Deleterious Variants Enriched in HCC Cases

From 1,995,435 variants detected in 719 Hispanic individuals, we selected 128 rare, potentially functional variants mapped to the 40 candidate genes, including missense (n = 112), frameshift (n = 7), splicing (n = 5), start-loss (n = 3) and stop-gain (n = 1) variants. Among these, 126 variants had CADD ≥ 15, except one frameshift and one start-loss variant.
We further evaluated 29 prioritized variants (located in 16 genes) that showed enrichment in HCC cases compared to external gnomAD AMR controls (nominal p-value < 0.05) by single variant allelic association test (Table 2). Nine of these variants remained statistically significant after FDR correction. Among these 29 variants, two were frameshift variants and the rest were missense variants. Based on deleteriousness levels, five variants were classified as high, six as medium, and 18 as low. Among the five highly deleterious variants, two were protein-truncating variants located in HLA-DRB1, and three were missense variants located in Oncostatin M Receptor (OSMR, n = 2) and TM6SF2 (n = 1). The high deleteriousness variant TM6SF2 R138W also has the highest CADD score, followed by WNT9A R357H. TM6SF2 R138W also has a relatively higher allele frequency compared to other variants in HCC cases and the gnomAD AMR population. Out of the 29 variants, 23 showed higher allele frequencies in the gnomAD AMR compared with East Asian (EAS) and Non-Finnish European (NFE) populations.
In analyses comparing HCC cases with internal controls, we were unable to estimate ORs and 95% CIs for most variants as all internal controls were non-carriers. Despite this, we observed that TM6SF2 R138W was significantly enriched and carriers had approximately 10-fold increased risk of HCC (carriers vs. non-carriers, OR 10.62, 95% CI: 1.67–443.08). When compared with the gnomAD AMR population, nine variants remained significantly associated with HCC risk after multiple testing corrections (Table 2). Among these nine, C2 I484V exhibited the largest effect size (OR 26.43, 95% CI: 2.51–161.78) and TM6SF2 R138W showed a moderate effect size (OR 2.66, 95% CI: 1.55–4.28). In addition, the allele frequency of TM6SF2 R138W was more than two-fold higher in cases (allele frequency: 1.98%) than in the gnomAD AMR controls (allele frequency: 0.75%). As a sensitivity analysis, Firth’s penalized logistic regression yielded largely consistent effect estimates with those obtained from Fisher’s exact test, although confidence intervals remained wide because of the rarity of the variants (Supplementary Table S4).
We performed exploratory analyses to examine the distribution of prioritized variants across different HCC etiologies. The frequency of the 29 prioritized variants appeared to differ among HCC cases by etiology (Figure 2). Variants in HLA-DRB1, OSMR and C2 genes were observed among cases with viral hepatitis, ALD, and MASLD etiologies, while variants in TM6SF2 and APOB were more frequently detected in ALD and MASLD-related HCC. In particular, TM6SF2 was enriched among cases with ALD- or MASLD-related HCC, compared to other etiologies (OR 3.30, 95% CI: 0.98–14.27, nominal p = 0.040), supporting a potential etiology-specific role.

3.3. Gene-Level Burden Analysis Supported Cumulative Rare-Variant Effects

In the gene burden tests, a total of 23 genes containing at least two rare variants were analyzed. Of these, five genes were significantly associated with HCC (p < 0.05, Table 3), comprising 43 ultra-rare and nine rare variants. TM6SF2 remained significantly associated with HCC risk after multiple testing corrections. Among these five genes, TM6SF2 and HLA-DRB1 overlapped with strong single-variant association signals (FDR-adjusted p-value < 0.05) while Obscurin, Cytoskeletal Calmodulin and Titin-interacting RhoGEF (OBSCN) and OSMR overlapped with moderate single-variant association signals (nominal p-value < 0.05, Table 2). Notably, rare variants in TM6SF2 showed a large dose effect (OR 15.29, 95% CI: 2.06–113.51), and rare variants in HLA-DRB1 showed a moderate dose effect (OR 2.50, 95% CI: 1.14–5.50).

3.4. HCC Risk Increased with the Number of Rare Deleterious Variants Carried

To evaluate whether the cumulative burden of rare variants contributed to increased HCC risk, we examined the association between the number of variants and HCC risk. When all 29 prioritized variants were considered, we observed a strong and statistically significant cumulative trend (OR 3.84, 95% CI: 2.40–6.53, Cochran–Armitage test, p-value = 1.56 × 10−8, Table 4). This trend was also observed when analyzed stratifying by sex (Supplementary Figure S2). Among all participants, 592 individuals were non-carriers of these 29 variants, whereas 127 individuals carried at least one variant (108 HCC cases and 19 internal controls, Supplementary Figure S1A). Among these 108 HCC carriers, most harbored a single variant, while 10 cases carried two variants and one HCC carried three variants (Supplementary Figure S1B). Among carriers of TM6SF2 and APOB, co-occurring variants were observed in C2, HLA-DRB1, OBSCN, OSMR and TERT. When the analysis was restricted to the five variants classified as highly deleterious, a significant dose-effect trend was also observed, with a larger estimated effect size compared to the full set of 29 variants (OR 10.00, 95% CI: 3.03–61.85 versus OR 3.84, 95% CI: 2.40–6.53, Supplementary Figure S2).

4. Discussion

In this study, we demonstrate that rare deleterious germline variants are enriched in Hispanic HCC cases and exhibit effect sizes that far exceed those of common variants reported from HCC GWAS among European and Asian populations. We highlighted four variants, TM6SF2 R138W, APOB I3721T, HLA-DRB1 W38X, and C2 I484V which represent major biological pathways that were markedly enriched in HCC cases, suggesting that rare coding variants may contribute to HCC susceptibility in Hispanic individuals (Supplementary Figure S3). These findings indicate population-specific genetic architecture of HCC, with certain rare variants exerting stronger effects in Hispanic individuals than previously observed in European or East Asian cohorts. These findings align with the disproportionate HCC burden and its etiologic features reported in Hispanic individuals, including a higher prevalence of metabolic syndrome [43], as well as different underlying etiologic patterns for HCC [44].
Our study extended prior findings on TM6SF2. Our group previously identified the common variants E167K as a susceptibility locus for HCC in populations of European descent in North American [12]. In the present study, we identified a rare variant, R138W, with a larger observed effect size (OR 2.66) in Hispanic individuals than that previously reported for the common E167K variant in European populations (OR 1.49) [12]. However, this comparison should be interpreted with caution because rare and common variants have different statistical properties, and effect estimates for rare variants may be inflated due to the winner’s curse. In exploratory analysis, R138W appeared to be enriched among ALD- or MASLD-related HCC cases. This observation suggests that this variant may modify risk in the context of alcohol-associated liver injury or metabolic dysregulation, consistent with the established role of TM6SF2 in hepatic lipid metabolism. With an MAF of ~0.94% in the gnomAD AMR population, this variant may affect a non-trivial population-level risk of Hispanic individuals. Given growing evidence that rare variants improve risk stratification and refining disease diagnosis, incorporating R138W into prediction models for HCC arising from ALD- and MASLD-related cirrhosis may enhance risk discrimination in high-risk Hispanic populations.
Although the candidate genes included in this study were selected based on prior evidence of association with HCC, only a subset showed clear enrichment signals in our Hispanic HCC cohort, whereas many previously implicated genes did not. The enrichment of rare deleterious variants in lipid metabolism genes (i.e., TM6SF2 and APOB) aligns with the high burden of MASLD and ALD in Hispanic populations and suggests that metabolic pathways may play a disproportionately large role in inherited HCC susceptibility in this group [43]. Rare variants in immune response genes, including HLA-DRB1 and C2, also showed enrichment. HLA-DRB1 has been implicated in susceptibility to several malignancies [45], supporting a hypothesis of immune-related genetic variation having a substantial impact on carcinogenesis. Given the high polymorphism and functional diversity of HLA loci, these findings warrant cautious interpretation but may reflect ancestry-specific immune modulation relevant to chronic liver injury and carcinogenesis.
Across analytical approaches, lipid metabolism and immune pathways were consistently identified in single-variant, gene-based burden, and dose–response analyses, reinforcing their robustness. In dose–response analysis, the five variants classified as highly deleterious were all mapped to these two pathways, further supporting their contribution to HCC susceptibility. Although OBSCN showed a strong association in the burden test, a few rare variants in this gene were identified in the single-variant analysis. Given that OBSCN is among the largest genes in the genome, it is more likely to accumulate rare variants by chance. Therefore, its association in the burden analysis may partly reflect gene size rather than a true biological contribution to HCC and should be interpreted with caution.
The main strength of our study is that it provides the first exome-sequencing data on HCC of Hispanic individuals in the US, addressing a population at high HCC risk, with high genetic diversity but limited prior genomic characterization. Previous work by Alejandro et al. examined only six known common variants in lipid metabolism genes and reported modest effect sizes (OR ~2.2) for three missense and one intronic variant [25]. In contrast, our study systematically evaluated rare coding variants across 86 genes, offering an initial basis for examining Hispanic-specific variation in HCC. The second major strength is our comprehensive etiological profiling of HCC cases (HBV, HCV, ALD, MASLD and other etiologies), which enhances interpretation of genetic findings. Also, because reliance on external controls alone can introduce confounding from platform differences, batch effects, and ancestry mismatch, we included internal controls drawn from the same geographic regions, and processed and sequenced alongside cases. This design improves geographic and ancestral comparability, reduces technical artifacts, and strengthens confidence in our rare variant filtering and discoveries.
Several limitations should be acknowledged. First, the small sample size limited our statistical power for whole-exome-wide discovery. To mitigate this, we focused on genes previously identified as HCC susceptibility genes from European and East Asian populations to improve statistical power. In addition, residual population stratification may also have influenced the observed associations because rare variants often exhibit different population stratification patterns compared with common variants [46]. Therefore, given the genetic heterogeneity among Hispanics across the US [47], validation in larger and independent cohorts representing diverse Hispanic ancestral backgrounds beyond our primarily Texas-based cohort will be necessary. Second, functional validation of TM6SF2 R138W and other prioritized variants is needed to elucidate their underlying biological mechanisms. Future functional studies could use human hepatocyte cell lines and knock-in mouse models to validate the functional consequences of TM6SF2 R138W on hepatic lipid metabolism and its downstream biological effects [48,49]. Third, the inclusion of affected or at-risk individuals among controls may have attenuated effect estimates. Precursor and related conditions, such as MASLD, ALD and cirrhosis, may also be present among the gnomAD controls, potentially leading to underestimation of true associations.

5. Conclusions

Our findings provide evidence that rare deleterious variants, especially in lipid metabolism and immune response, may contribute to increased HCC risk in Hispanic individuals, and suggest population-specific genetic architecture of HCC. The enrichment of variants such as TM6SF2 R138W and the observed cumulative burden pattern highlighted the interplay between HCC risk and metabolic or environmental factors in this high-risk population. Future integration of coding, regulatory, environmental, and mechanistic data will be required to fully delineate the characteristics of high-risk subgroups and translate these findings into precision prevention strategies for Hispanic populations.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/cancers18152428/s1, Figure S1A: Landscape of the 29 prioritized rare deleterious variants across 127 carriers, including 108 HCC cases (black) and 19 controls (blank white). Each column represents one participant. Upper 4 rows represent HCC status, demographic characteristics (age, sex) and etiology, and lower 29 rows represent variants in genes. Colored cells in gene rows indicate the presence of specific variants. For genes with more than 1 variants, colors distinguish the specific variants. HCC, Hepatocellular carcinoma; ALD, Alcoholic liver dis-ease; MASLD, Metabolic dysfunction-associated steatotic liver disease; HCV, Hepatitis C virus; Figure S1B: Landscape of variants of 11 individuals carrying at least two prioritized variants. Among these individuals, 9 have ALD or MASLD, and 1 has HCV hepatitis. Each column represents one participant. Upper 3 rows represent age, sex and etiology, and lower 19 rows represent variants in genes. HCC, Hepatocellular carcinoma; ALD, Alcoholic liver disease; MASLD, Metabolic dysfunc-tion-associated steatotic liver disease; HCV, Hepatitis C virus; Figure S2: Association between rare variant burden and HCC risk. Odds ratios (ORs) and 95% confidence intervals (CIs) for HCC risk per additional variant carried are shown for all 29 prioritized rare variants (top panel) and for the subset of five variants classified as high deleteriousness (bottom panel). Estimates were obtained from logistic regression models with the number of variants carried as a continuous predictor. Analyses were performed for all participants combined, and stratified by sex. The dashed vertical line indicates OR = 1; Figure S3: Flow chart of variant prioritization and representative biological pathways; Table S1: Curated list of candidate genes associated with HCC and supporting references; Table S2: Sequencing coverage and quality statistics; Table S3: Estimated statistical power of single-variant association comparison; Table S4: Comparison of odds ratios for the 29 prioritized variants estimated using Fisher’s exact test and Firth’s penalized logistic regression.

Author Contributions

Conceptualization: Y.L. and A.P.T.; Methodology: X.L., Y.L., S.T. and P.B.S.; Formal Analysis: X.L., Y.L., S.T. and P.B.S.; Data Curation: X.L., Y.L., S.T. and A.P.T.; Investigation: X.L., Y.L. and A.P.T.; Visualization: X.L. and Y.L.; Resources: A.G.S., R.H., S.K.A., S.K., H.B.E.-S. and A.P.T.; Funding Acquisition: A.G.S., H.B.E.-S. and A.P.T.; Project Administration: A.P.T.; Supervision: A.P.T.; Writing—original draft: X.L. and Y.L.; Writing—review and editing: X.L., Y.L., A.G.S., R.H., S.K.A., S.K., H.B.E.-S. and A.P.T. All authors have read and agreed to the published version of the manuscript.

Funding

We would like to thank all individuals who participated in this study. This work was supported by grants from the Cancer Prevention and Research Institute of Texas (CPRIT; RP150587, RP200537, RP200554, RR210027, RP220119 and RP250355), the National Cancer Institute of the National Institutes of Health (NCI P01 CA263025, NCI P50 CA295495, NCI U01 CA230997, NCI U01 CA271887 and NCI U01 CA283935), and in part by the Center for Gastrointestinal Development, Infection and Injury funded by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK P30 DK56338).

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and the Declaration of Istanbul, and approved by the Institutional Review Board of Baylor College of Medicine (approval number: H-48548; approval date: 7 September 2021). All participants provided informed consent.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

Data supporting the findings of this study are available within the article and its Supplementary Materials. The whole-exome sequencing data generated in this study contain individual-level human genomic information and are subject to institutional, ethical, legal, and participant consent restrictions designed to protect participant privacy. Data and supporting documentation will be made available through Baylor College of Medicine’s section of Digital Commons@TMC. Access to genomic data will be provided in accordance with applicable institutional review board, data use, and privacy requirements. Additional information supporting the findings of this study is available from the corresponding author upon reasonable request.

Conflicts of Interest

Singal has served as a consultant or on advisory boards for Genentech, AstraZeneca, Eisai, Exelixis, Bayer, Merck, Elevar, Boston Scientific, Sirtex, FujiFilm Medical Sciences, Exact Sciences, Helio Genomics, Roche, Glycotest, ImCare, Curve Bio, DELFI, Mursla, Universal Dx, and Abbott. All other authors report no conflicts.

Abbreviations

AASLD, American Association for the Study of Liver Diseases; ALD, alcoholic liver disease; AMR, Admixed American; C2, Complement C2; CADD, Combined Annotation Dependent Depletion; CI, confidence interval; CT, Computed Tomography; EAS, East Asian; FDR, false discovery rate; GATK, Genome Analysis Toolkit; GCKR, Glucokinase Regulator; gnomAD, Genome Aggregation Database; GWAS, genome-wide association studies; HBV, hepatitis B virus; HCC, hepatocellular carcinoma; HCV, hepatitis C virus; HGSC, Human Genome Sequencing Center; HLA-DQB1, human leukocyte antigen, class II DQ beta 1; MAF, minor allele frequency; MASLD, metabolic dysfunction-associated steatotic liver disease; MAU2, MAU2 sister chromatid cohesion factor; MBOAT7, membrane bound acylglycerophosphatidylinositol O-acyltransferase MBOAT7; MICA, human leukocyte antigen, class I polypeptide-related sequence A; MOBP, myelin associated oligodendrocyte basic protein; MRI, Magnetic Resonance Imaging; NAFLD, nonalcoholic fatty liver disease; NFE, Non-Finnish European; OR, odds ratio; OBSCN, obscurin, cytoskeletal calmodulin and titin-interacting RhoGEF; OSMR, oncostatin M receptor; PNPLA3, patatin-like domain 3, 1-acylglycerol-3-phosphate O-acyltransferase; PolyPhen-2, Polymorphism Phenotyping v2; SAMM50, sorting and assembly machinery component; SIFT, Sorting Intolerant From Tolerant; SKAT, Sequence Kernel Association Test; SNV, Single Nucleotide Variant; TERT, telomerase reverse transcriptase; TM6SF2, transmembrane 6 superfamily member 2; WNT3A, Wnt family member 3A; WNT9A, Wnt family member 9A.

References

  1. Sung, H.; Ferlay, J.; Siegel, R.L.; Laversanne, M.; Soerjomataram, I.; Jemal, A.; Bray, F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J. Clin. 2021, 71, 209–249. [Google Scholar] [CrossRef] [PubMed]
  2. Toh, M.R.; Wong, E.Y.T.; Wong, S.H.; Ng, A.W.T.; Loo, L.H.; Chow, P.K.; Ngeow, J. Global Epidemiology and Genetics of Hepatocellular Carcinoma. Gastroenterology 2023, 164, 766–782. [Google Scholar] [CrossRef] [PubMed]
  3. Cholankeril, G.; Thrift, A.P.; Khaderi, S.; Kanwal, F.; El-Serag, H.B.; Duong, H.; Singal, A.G.; Asrani, S.K.; Kourkompetis, T.; Zamir, A.; et al. Temporal Changes in Risk Factors of Hepatocellular Carcinoma in a Prospective Cirrhosis Cohort. Clin. Gastroenterol. Hepatol. 2025, 23, 1268–1270 e1262. [Google Scholar] [CrossRef] [PubMed]
  4. Kanwal, F.; Khaderi, S.; Singal, A.G.; Marrero, J.A.; Loo, N.; Asrani, S.K.; Amos, C.I.; Thrift, A.P.; Gu, X.; Luster, M.; et al. Risk factors for HCC in contemporary cohorts of patients with cirrhosis. Hepatology 2023, 77, 997–1005. [Google Scholar] [CrossRef] [PubMed]
  5. Makarova-Rusher, O.V.; Altekruse, S.F.; McNeel, T.S.; Ulahannan, S.; Duffy, A.G.; Graubard, B.I.; Greten, T.F.; McGlynn, K.A. Population attributable fractions of risk factors for hepatocellular carcinoma in the United States. Cancer 2016, 122, 1757–1765. [Google Scholar] [CrossRef] [PubMed]
  6. Yu, M.-W.; Chang, H.-C.; Liaw, Y.-F.; Lin, S.-M.; Lee, S.-D.; Liu, C.-J.; Chen, P.-J.; Hsiao, T.-J.; Lee, P.-H.; Chen, C.-J. Familial Risk of Hepatocellular Carcinoma Among Chronic Hepatitis B Carriers and Their Relatives. J. Natl. Cancer Inst. 2000, 92, 1159–1164. [Google Scholar] [CrossRef] [PubMed]
  7. Thrift, A.P.; Kanwal, F.; Lim, H.; Duong, H.; Liu, Y.; Singal, A.G.; Khaderi, S.; Asrani, S.K.; Amos, C.I.; El-Serag, H.B. PNPLA3, Obesity, and Heavy Alcohol Use in Cirrhosis Patients May Exert a Synergistic Increase Hepatocellular Carcinoma Risk. Clin. Gastroenterol. Hepatol. 2024, 22, 1858–1866.e4. [Google Scholar] [CrossRef] [PubMed]
  8. Thrift, A.P.; Liu, K.S.; Raza, S.A.; El-Serag, H.B. Recent Decline in the Incidence of Hepatocellular Carcinoma in the United States. Clin. Gastroenterol. Hepatol. 2023, 21, 2418–2420.e3. [Google Scholar] [CrossRef] [PubMed]
  9. Heiss, G.; Snyder, M.L.; Teng, Y.; Schneiderman, N.; Llabre, M.M.; Cowie, C.; Carnethon, M.; Kaplan, R.; Giachello, A.; Gallo, L.; et al. Prevalence of Metabolic Syndrome Among Hispanics/Latinos of Diverse Background: The Hispanic Community Health Study/Study of Latinos. Diabetes Care 2014, 37, 2391–2399. [Google Scholar] [CrossRef] [PubMed]
  10. Browning, J.D.; Szczepaniak, L.S.; Dobbins, R.; Nuremberg, P.; Horton, J.D.; Cohen, J.C.; Grundy, S.M.; Hobbs, H.H. Prevalence of hepatic steatosis in an urban population in the United States: Impact of ethnicity. Hepatology 2004, 40, 1387–1395. [Google Scholar] [CrossRef] [PubMed]
  11. El-Serag, H.B.; Kramer, J.; Duan, Z.; Kanwal, F. Racial differences in the progression to cirrhosis and hepatocellular carcinoma in HCV-infected veterans. Am. J. Gastroenterol. 2014, 109, 1427–1435. [Google Scholar] [CrossRef] [PubMed]
  12. Hassan, M.M.; Li, D.; Han, Y.; Byun, J.; Hatia, R.I.; Long, E.; Choi, J.; Kelley, R.K.; Cleary, S.P.; Lok, A.S.; et al. Genome-wide association study identifies high-impact susceptibility loci for HCC in North America. Hepatology 2024, 80, 87–101. [Google Scholar] [CrossRef] [PubMed]
  13. Trepo, E.; Nahon, P.; Bontempi, G.; Valenti, L.; Falleti, E.; Nischalke, H.D.; Hamza, S.; Corradini, S.G.; Burza, M.A.; Guyot, E.; et al. Association between the PNPLA3 (rs738409 C>G) variant and hepatocellular carcinoma: Evidence from a meta-analysis of individual participant data. Hepatology 2014, 59, 2170–2177. [Google Scholar] [CrossRef] [PubMed]
  14. Stickel, F.; Buch, S.; Nischalke, H.D.; Weiss, K.H.; Gotthardt, D.; Fischer, J.; Rosendahl, J.; Marot, A.; Elamly, M.; Casper, M.; et al. Genetic variants in PNPLA3 and TM6SF2 predispose to the development of hepatocellular carcinoma in individuals with alcohol-related cirrhosis. Am. J. Gastroenterol. 2018, 113, 1475–1483. [Google Scholar] [CrossRef] [PubMed]
  15. Jiang, D.K.; Sun, J.; Cao, G.; Liu, Y.; Lin, D.; Gao, Y.Z.; Ren, W.H.; Long, X.D.; Zhang, H.; Ma, X.P.; et al. Genetic variants in STAT4 and HLA-DQ genes confer risk of hepatitis B virus-related hepatocellular carcinoma. Nat. Genet. 2013, 45, 72–75. [Google Scholar] [CrossRef] [PubMed]
  16. Kumar, V.; Kato, N.; Urabe, Y.; Takahashi, A.; Muroyama, R.; Hosono, N.; Otsuka, M.; Tateishi, R.; Omata, M.; Nakagawa, H.; et al. Genome-wide association study identifies a susceptibility locus for HCV-induced hepatocellular carcinoma. Nat. Genet. 2011, 43, 455–458. [Google Scholar] [CrossRef] [PubMed]
  17. Buch, S.; Innes, H.; Lutz, P.L.; Nischalke, H.D.; Marquardt, J.U.; Fischer, J.; Weiss, K.H.; Rosendahl, J.; Marot, A.; Krawczyk, M.; et al. Genetic variation in TERT modifies the risk of hepatocellular carcinoma in alcohol-related cirrhosis: Results from a genome-wide case-control study. Gut 2023, 72, 381–391. [Google Scholar] [CrossRef] [PubMed]
  18. Trepo, E.; Caruso, S.; Yang, J.; Imbeaud, S.; Couchy, G.; Bayard, Q.; Letouze, E.; Ganne-Carrie, N.; Moreno, C.; Oussalah, A.; et al. Common genetic variation in alcohol-related hepatocellular carcinoma: A case-control genome-wide association study. Lancet Oncol. 2022, 23, 161–171. [Google Scholar] [CrossRef] [PubMed]
  19. Clifford, R.J.; Zhang, J.; Meerzaman, D.M.; Lyu, M.S.; Hu, Y.; Cultraro, C.M.; Finney, R.P.; Kelley, J.M.; Efroni, S.; Greenblum, S.I.; et al. Genetic variations at loci involved in the immune response are risk factors for hepatocellular carcinoma. Hepatology 2010, 52, 2034–2043. [Google Scholar] [CrossRef] [PubMed]
  20. Kawaguchi, T.; Shima, T.; Mizuno, M.; Mitsumoto, Y.; Umemura, A.; Kanbara, Y.; Tanaka, S.; Sumida, Y.; Yasui, K.; Takahashi, M.; et al. Risk estimation model for nonalcoholic fatty liver disease in the Japanese using multiple genetic markers. PLoS ONE 2018, 13, e0185490. [Google Scholar] [CrossRef] [PubMed]
  21. Han, Y.; Shaw, V.R.; Byun, J.; Thrift, A.P.; Zhu, C.; Li, D.; Hatia, R.I.; Kelley, R.K.; Cleary, S.P.; Lok, A.S.; et al. Linkage disequilibrium score regression identifies genetic correlations between hepatocellular carcinoma and clinically relevant traits. Int. J. Cancer 2026, 158, 1193–1203. [Google Scholar] [CrossRef] [PubMed]
  22. Wainschtein, P.; Jain, D.; Zheng, Z.; TOPMed Anthropometry Working Group; NHLBI Trans-Omics for Precision Medicine (TOPMed) Consortium; Cupples, L.A.; Shadyab, A.H.; McKnight, B.; Shoemaker, B.M.; Mitchell, B.D.; et al. Assessing the contribution of rare variants to complex trait heritability from whole-genome sequence data. Nat. Genet. 2022, 54, 263–273. [Google Scholar] [CrossRef] [PubMed]
  23. Cirulli, E.T.; Goldstein, D.B. Uncovering the roles of rare variants in common disease through whole-genome sequencing. Nat. Rev. Genet. 2010, 11, 415–425. [Google Scholar] [CrossRef] [PubMed]
  24. Kuchenbaecker, K.B.; Hopper, J.L.; Barnes, D.R.; Phillips, K.A.; Mooij, T.M.; Roos-Blom, M.J.; Jervis, S.; van Leeuwen, F.E.; Milne, R.L.; Andrieu, N.; et al. Risks of Breast, Ovarian, and Contralateral Breast Cancer for BRCA1 and BRCA2 Mutation Carriers. JAMA 2017, 317, 2402–2416. [Google Scholar] [CrossRef] [PubMed]
  25. Cruz, A.A.; Hernández, J.C.N.; Garza, L.E.C.; Duarte, A.M.; Tijerina, V.L.M.; Garcia, M.E.H.; Peñuelas-Urquides, K.; González-Escalante, L.A.; de León, M.B.; Ramirez, B.S. Contribution of PNPLA3, GCKR, MBOAT7, NCAN, and TM6SF2 Genetic Variants to Hepatocellular Carcinoma Development in Mexican Patients. Int. J. Mol. Sci. 2025, 26, 7409. [Google Scholar] [CrossRef] [PubMed]
  26. Lee, S.; Abecasis, G.R.; Boehnke, M.; Lin, X. Rare-variant association analysis: Study designs and statistical tests. Am. J. Hum. Genet. 2014, 95, 5–23. [Google Scholar] [CrossRef] [PubMed]
  27. Manolio, T.A.; Collins, F.S.; Cox, N.J.; Goldstein, D.B.; Hindorff, L.A.; Hunter, D.J.; McCarthy, M.I.; Ramos, E.M.; Cardon, L.R.; Chakravarti, A.; et al. Finding the missing heritability of complex diseases. Nature 2009, 461, 747–753. [Google Scholar] [CrossRef] [PubMed]
  28. Ghouse, J.; Sveinbjornsson, G.; Vujkovic, M.; Seidelin, A.S.; Gellert-Kristensen, H.; Ahlberg, G.; Tragante, V.; Rand, S.A.; Brancale, J.; Vilarinho, S.; et al. Integrative common and rare variant analyses provide insights into the genetic architecture of liver cirrhosis. Nat. Genet. 2024, 56, 827–837. [Google Scholar] [CrossRef] [PubMed]
  29. Hanis, C.L.; Ferrell, R.E.; Barton, S.A.; Aguilar, L.; Garzaibarra, A.N.A.; Tulloch, B.R.; Garcia, C.A.; Schull, W.J. Diabetes among mexican americans in starr county, texas. Am. J. Epidemiol. 1983, 118, 659–672. [Google Scholar] [CrossRef] [PubMed]
  30. Cerezo, M.; Sollis, E.; Ji, Y.; Lewis, E.; Abid, A.; Bircan, K.O.; Hall, P.; Hayhurst, J.; John, S.; Mosaku, A.; et al. The NHGRI-EBI GWAS Catalog: Standards for reusability, sustainability and diversity. Nucleic Acids Res. 2024, 53, D998–D1005. [Google Scholar] [CrossRef] [PubMed]
  31. Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv 2013, arXiv:1303.3997. [Google Scholar] [CrossRef]
  32. Poplin, R.; Ruano-Rubio, V.; DePristo, M.A.; Fennell, T.J.; Carneiro, M.O.; Van der Auwera, G.A.; Kling, D.E.; Gauthier, L.D.; Levy-Moonshine, A.; Roazen, D. Scaling accurate genetic variant discovery to tens of thousands of samples. bioRxiv 2017. preprint. [Google Scholar] [CrossRef]
  33. Robinson, J.T.; Thorvaldsdóttir, H.; Winckler, W.; Guttman, M.; Lander, E.S.; Getz, G.; Mesirov, J.P. Integrative genomics viewer. Nat. Biotechnol. 2011, 29, 24–26. [Google Scholar] [CrossRef] [PubMed]
  34. McLaren, W.; Gil, L.; Hunt, S.E.; Riat, H.S.; Ritchie, G.R.S.; Thormann, A.; Flicek, P.; Cunningham, F. The Ensembl Variant Effect Predictor. Genome Biol. 2016, 17, 122. [Google Scholar] [CrossRef] [PubMed]
  35. Chen, S.; Francioli, L.C.; Goodrich, J.K.; Collins, R.L.; Kanai, M.; Wang, Q.; Alföldi, J.; Watts, N.A.; Vittal, C.; Gauthier, L.D.; et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 2024, 625, 92–100. [Google Scholar] [CrossRef] [PubMed]
  36. Schubach, M.; Maass, T.; Nazaretyan, L.; Röner, S.; Kircher, M. CADD v1.7: Using protein language models, regulatory CNNs and other nucleotide-level scores to improve genome-wide variant predictions. Nucleic Acids Res. 2024, 52, D1143–D1154. [Google Scholar] [CrossRef] [PubMed]
  37. Adzhubei, I.A.; Schmidt, S.; Peshkin, L.; Ramensky, V.E.; Gerasimova, A.; Bork, P.; Kondrashov, A.S.; Sunyaev, S.R. A method and server for predicting damaging missense mutations. Nat. Methods 2010, 7, 248–249. [Google Scholar] [CrossRef] [PubMed]
  38. Sim, N.-L.; Kumar, P.; Hu, J.; Henikoff, S.; Schneider, G.; Ng, P.C. SIFT web server: Predicting effects of amino acid substitutions on proteins. Nucleic Acids Res. 2012, 40, W452–W457. [Google Scholar] [CrossRef] [PubMed]
  39. Benjamini, Y.; Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J. R. Stat. Soc. Ser. B Methodol. 1995, 57, 289–300. [Google Scholar] [CrossRef]
  40. Chang, C.C.; Chow, C.C.; Tellier, L.C.; Vattikuti, S.; Purcell, S.M.; Lee, J.J. Second-generation PLINK: Rising to the challenge of larger and richer datasets. GigaScience 2015, 4, 7. [Google Scholar] [CrossRef] [PubMed]
  41. Wu, M.C.; Lee, S.; Cai, T.; Li, Y.; Boehnke, M.; Lin, X. Rare-variant association testing for sequencing data with the sequence kernel association test. Am. J. Hum. Genet. 2011, 89, 82–93. [Google Scholar] [CrossRef] [PubMed]
  42. Armitage, P. Tests for Linear Trends in Proportions and Frequencies. Biometrics 1955, 11, 375–386. [Google Scholar] [CrossRef]
  43. Zhang, X.; Heredia, N.I.; Balakrishnan, M.; Thrift, A.P. Prevalence and factors associated with NAFLD detected by vibration controlled transient elastography among US adults: Results from NHANES 2017–2018. PLoS ONE 2021, 16, e0252164. [Google Scholar] [CrossRef] [PubMed]
  44. Kuftinec, G.N.; Levy, R.; Kieffer, D.A.; Medici, V. Hepatocellular Carcinoma and Associated Clinical Features in Latino and Caucasian Patients from a Single Center. Ann. Hepatol. 2019, 18, 177–186. [Google Scholar] [CrossRef] [PubMed]
  45. Kamiza, A.B.; Kamiza, S.; Mathew, C.G. HLA-DRB1 alleles and cervical cancer: A meta-analysis of 36 case-control studies. Cancer Epidemiol. 2020, 67, 101748. [Google Scholar] [CrossRef] [PubMed]
  46. Mathieson, I.; McVean, G. Differential confounding of rare and common variants in spatially structured populations. Nat. Genet. 2012, 44, 243–246. [Google Scholar] [CrossRef] [PubMed]
  47. Conomos, M.P.; Laurie, C.A.; Stilp, A.M.; Gogarten, S.M.; McHugh, C.P.; Nelson, S.C.; Sofer, T.; Fernández-Rhodes, L.; Justice, A.E.; Graff, M.; et al. Genetic Diversity and Association Studies in US Hispanic/Latino Populations: Applications in the Hispanic Community Health Study/Study of Latinos. Am. J. Hum. Genet. 2016, 98, 165–184. [Google Scholar] [CrossRef] [PubMed]
  48. Mahdessian, H.; Taxiarchis, A.; Popov, S.; Silveira, A.; Franco-Cereceda, A.; Hamsten, A.; Eriksson, P.; van’t Hooft, F. TM6SF2 is a regulator of liver fat metabolism influencing triglyceride secretion and hepatic lipid droplet content. Proc. Natl. Acad. Sci. USA 2014, 111, 8913–8918. [Google Scholar] [CrossRef] [PubMed]
  49. Newberry, E.P.; Hall, Z.; Xie, Y.; Molitor, E.A.; Bayguinov, P.O.; Strout, G.W.; Fitzpatrick, J.A.; Brunt, E.M.; Griffin, J.L.; Davidson, N.O. Liver-Specific Deletion of Mouse Tm6sf2 Promotes Steatosis, Fibrosis, and Hepatocellular Cancer. Hepatology 2021, 74, 1203–1219. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Study design and analytical workflow. WES, whole-exome sequencing; HCC, hepatocellular carcinoma; MAF, minor allele frequency; AMR, Admixed American from gnomAD; CADD, Combined Annotation Dependent Depletion; PolyPhen-2: Polymorphism Phenotyping v2; SIFT: Sorting Intolerant From Tolerant.
Figure 1. Study design and analytical workflow. WES, whole-exome sequencing; HCC, hepatocellular carcinoma; MAF, minor allele frequency; AMR, Admixed American from gnomAD; CADD, Combined Annotation Dependent Depletion; PolyPhen-2: Polymorphism Phenotyping v2; SIFT: Sorting Intolerant From Tolerant.
Cancers 18 02428 g001
Figure 2. Distribution of prioritized rare variants by gene across etiologies. Each row represents a gene. The number in parentheses after each gene name indicates the number of prioritized variants identified in that gene. Each column represents an HCC etiology. The number in parentheses after each etiology indicates the total number of HCC cases in that etiology group. Values in each cell represent the number of variant occurrences observed for a given gene within each etiology. The last column shows the total number of variant occurrences across all HCC etiologies. Darker blue shading indicates a higher number of variant occurrences within each cell.
Figure 2. Distribution of prioritized rare variants by gene across etiologies. Each row represents a gene. The number in parentheses after each gene name indicates the number of prioritized variants identified in that gene. Each column represents an HCC etiology. The number in parentheses after each etiology indicates the total number of HCC cases in that etiology group. Values in each cell represent the number of variant occurrences observed for a given gene within each etiology. The last column shows the total number of variant occurrences across all HCC etiologies. Darker blue shading indicates a higher number of variant occurrences within each cell.
Cancers 18 02428 g002
Table 1. Baseline characteristics of study participants and gnomAD AMR reference population.
Table 1. Baseline characteristics of study participants and gnomAD AMR reference population.
Characteristic aCases (n = 455)Controls (n = 264)gnomAD AMR
(n = 30,019)
Sex (Male, %)302 (66.4)89 (33.7)13,775 (45.9)
Age (mean, SD)60.6 (9.50)48.3 (7.67)N.A.
HCC etiology (%)
   Viral hepatitis B or C136 (29.9)--
   Hepatitis C128 (28.1)
   Hepatitis B8 (1.8)
   ALD105 (23.1)--
   MASLD102 (22.4)--
   Other83 (18.2)--
Abbreviations: gnomAD, Genome Aggregation Database; AMR, Admixed American; N.A., not available; ALD, alcoholic liver disease; MASLD, metabolic dysfunction-associated steatotic liver disease. a Missing values (cases/controls): age (16/0); cirrhosis (29/-); HCC etiology (29/-).
Table 2. Prioritized 29 rare deleterious variant candidates *.
Table 2. Prioritized 29 rare deleterious variant candidates *.
GeneVariant aDeleterious Level bCADD cMAF % dMinor Allele CountHCC Cases vs. gnomAD AMR
Case/ControlAMR/EAS/NFE/Overall eCase/Control/AMROR (95% CI)FDR p Value f
HLA-DRB1c.101-1G>AHigh250.39/00.071/0.0065/0.14/0.143/0/225.45 (1.04–18.21)0.1070
HLA-DRB1W38XHigh21.60.73/0.210.049/0.019/0.02/0.0246/1/1815.13 (4.91–39.90)0.0003
TM6SF2R138WHigh29.81.98/0.190.75/0.065/0.00068/0.03118/1/4522.66 (1.55–4.28)0.0061
OSMRV436DHigh23.50.55/00.16/0/0.43/0.345/0/983.38 (1.07–8.18)0.1004
OSMRD262AHigh22.10.22/00.032/0/0.04/0.0312/0/196.95 (0.78–28.91)0.1489
WNT9AR357HMedium28.60.44/00.093/0.02/0.2/0.164/0/564.72 (1.24–12.82)0.0974
AXIN2P245SMedium23.50.33/00.053/0/0.063/0.0523/0/326.20 (1.21–19.87)0.0976
OBSCNT6187MMedium23.30.44/0.760.15/0/0.031/0.0314/4/883.01 (0.80–8.00)0.1782
OBSCNV5337MMedium22.50.22/00.017/0.0045/0.001/0.00172/0/1012.88 (1.37–60.53)0.0976
OBSCNG3945RMedium20.60.33/0.570.058/0/0.0064/0.00793/3/355.67 (1.11–18.02)0.1004
APOBI3721TMedium20.10.33/00.033/0/0/0.00123/0/209.92 (1.88–33.53)0.0467
GRIK1E821KLow24.70.22/00.0083/0/0.0025/0.00252/0/526.41 (2.51–161.69)0.0467
DYSFD1477GLow24.10.44/00.13/0/0.0025/0.144/0/773.43 (0.91–9.18)0.1371
C2T184MLow23.10.22/00.018/0/0.004/0.00382/0/1112.02 (1.29–55.16)0.0976
TLL1N84ILow22.80.22/0.190.03/0.0022/0/0.00122/1/187.32 (0.82–30.64)0.1412
TERTA730TLow20.70.22/00.022/0/0.00093/0.00162/0/1310.17 (1.11–44.99)0.1020
HLA-DRB1Q260PLow19.980.78/00.18/0.11/0.037/0.16/0/774.34 (1.54–9.93)0.0467
OBSCNR6892QLow19.980.22/00.02/0/0.0019/0.00312/0/1211.01 (1.20–49.61)0.1004
PCDH9S1209RLow19.240.77/0.190.25/0.0067/0.27/0.317/1/1473.15 (1.24–6.68)0.0808
CMTR2S331PLow17.881.21/0.380.52/0.0022/0.00017/0.0211/2/3142.33 (1.15–4.24)0.0886
KLHL8S178RLow17.870.33/00.022/0/0/0.000813/0/1315.26 (2.78–55.60)0.0274
APOBR778HLow17.750.22/00.025/0/0.00093/0.00172/0/158.81 (0.98–37.96)0.1152
TLL1A815DLow17.740.22/00.035/0.0022/0/0.00142/0/216.27 (0.71–25.70)0.1716
APOBD3907GLow17.610.22/0.190.025/0/0.0011/0.00772/1/158.81 (0.98–37.95)0.1152
OSMRK154TLow17.50.33/00.07/0/0/0.00263/0/424.72 (0.93–14.82)0.1260
C2I484VLow17.490.22/0.190.0083/0/0.0059/0.0222/1/526.43 (2.51–161.78)0.0467
HLA-DRB1M20TLow16.821.26/0.660.086/0.084/0.055/0.06710/3/3414.93 (6.55–31.04)<0.0001
FZD4V184LLow15.250.33/00.023/0/0/0.000993/0/1414.17 (2.61–50.95)0.0288
OBSCNS4397LLow15.20.22/0.190.02/0.011/0.00059/0.00172/1/1210.96 (1.19–49.37)0.1004
Abbreviations: CADD, Combined Annotation Dependent Depletion; MAF, minor allele frequency; AMR, Admixed American; EAS, East Asian; NFE, European (non-Finnish); OR, odds ratio; CI, confidence interval; FDR, false discovery rate. a The variant c.101-1G>A represents a splice acceptor variant and W38X denotes a stop-gain variant. Variants in bold indicate those significantly enriched in HCC cases compared with the gnomAD AMR population (FDR p-value < 0.05). b Deleteriousness levels were defined based on functional consequences. Protein-truncating variants and missense variants predicted as “probably damaging” by PolyPhen-2 and “deleterious” by SIFT were classified as “High”. Variants predicted as such by either tool were classified as “Medium”, and those predicted as neither were classified as “Low”. c CADD score reflects the predicted deleteriousness of variants in the human genome. Phred-scaled CADD scores of ≥10, ≥20, ≥30 correspond to the top 10%, 1%, and 0.1% most deleterious variants, respectively. d MAF reported in %. e AMR, EAS, NFE and overall MAF were retrieved from Genome Aggregation Database allele frequency. f p-values were adjusted for multiple testing using the FDR method. * Variants were selected using a nominal p < 0.05 in the comparison with the external gnomAD AMR controls. Variants shown in bold remained significant after FDR correction.
Table 3. Burden test results for 23 HCC-related genes with at least 2 rare variants.
Table 3. Burden test results for 23 HCC-related genes with at least 2 rare variants.
GeneN. Variant in MAF Bin aN. Variant Carriers
Case/Control
OR (95% CI) bp-Value cFDR p-Value d
Ultra Rare
0.001–0.005
Low Frequency
0.005–0.0132
TM6SF23125/115.29 (2.06–113.51)0.00040.0092
OBSCN30488/341.62 (1.06–2.49)0.01150.1322
MICA1112/23.53 (0.78–15.90)0.02520.1504
HLA-DRB15233/82.50 (1.14–5.50)0.02620.1504
OSMR4126/62.61 (1.06–6.42)0.04530.2084
MAU2204/05.27 (0.28–98.31)0.07910.3032
TERT205/06.46 (0.36–117.26)0.09270.3044
MBOAT7208/14.71 (0.59–37.84)0.12600.3624
DYSF5229/101.73 (0.83–3.61)0.16940.3903
EVA1C208/14.71 (0.59–37.84)0.16970.3903
MTARC1209/22.64 (0.57–12.33)0.23470.4898
C2408/22.34 (0.49–11.12)0.27620.4898
MRPS110219/71.60 (0.66–3.86)0.27680.4898
KIF1B206/21.75 (0.35–8.74)0.31390.5158
FZD4119/22.64 (0.57–12.33)0.37280.5587
TLL1204/12.33 (0.26–20.98)0.40750.5587
SQSTM1308/31.56 (0.41–5.92)0.41300.5587
TRIM31127/70.57 (0.2–1.65)0.54200.6611
AXIN2308/60.77 (0.26–2.24)0.54610.6611
PCDH94114/71.17 (0.46–2.93)0.73910.8300
HLA-DPA1204/12.33 (0.26–20.98)0.75780.8300
VEPH1608/50.93 (0.3–2.86)0.91230.9270
APOB7123/91.51 (0.69–3.31)0.92700.9270
Abbreviations: MAF, minor allele frequency; OR, odds ratio; CI, confidence interval; FDR, false discovery rate. a MAFs were calculated within the study population (combined cases and controls), and therefore, some variants may have MAF > 0.01 in this dataset despite being classified as rare. b ORs and 95% CIs were estimated by comparing carriers (≥1 included variant in a given gene) with non-carriers. c p-values were calculated from gene-based weighted burden tests. d p-values were adjusted for multiple testing using the FDR method across 23 tests.
Table 4. Distribution of individuals by the number of deleterious variants carried.
Table 4. Distribution of individuals by the number of deleterious variants carried.
GroupN. VariantsCaseControlOR (95% CI) ap-Value
Non-Carriers0347245Ref 1.0-
Carriers197193.84 (2.40–6.53)1.19 × 10−7
2100
310
Abbreviations: OR, odds ratio. a OR was estimated using logistic regression and represents the effect per additional variant, Cochran–Armitage trend test, p = 1.56 × 10−8.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, X.; Liu, Y.; Tsavachidis, S.; Shetty, P.B.; Singal, A.G.; Hernaez, R.; Asrani, S.K.; Khaderi, S.; El-Serag, H.B.; Thrift, A.P. Rare Germline Variants in Key Pathways Contribute to Hepatocellular Carcinoma Risk in Hispanic Individuals. Cancers 2026, 18, 2428. https://doi.org/10.3390/cancers18152428

AMA Style

Li X, Liu Y, Tsavachidis S, Shetty PB, Singal AG, Hernaez R, Asrani SK, Khaderi S, El-Serag HB, Thrift AP. Rare Germline Variants in Key Pathways Contribute to Hepatocellular Carcinoma Risk in Hispanic Individuals. Cancers. 2026; 18(15):2428. https://doi.org/10.3390/cancers18152428

Chicago/Turabian Style

Li, Xiangnan, Yanhong Liu, Spiridon Tsavachidis, Priya B. Shetty, Amit G. Singal, Ruben Hernaez, Sumeet K. Asrani, Saira Khaderi, Hashem B. El-Serag, and Aaron P. Thrift. 2026. "Rare Germline Variants in Key Pathways Contribute to Hepatocellular Carcinoma Risk in Hispanic Individuals" Cancers 18, no. 15: 2428. https://doi.org/10.3390/cancers18152428

APA Style

Li, X., Liu, Y., Tsavachidis, S., Shetty, P. B., Singal, A. G., Hernaez, R., Asrani, S. K., Khaderi, S., El-Serag, H. B., & Thrift, A. P. (2026). Rare Germline Variants in Key Pathways Contribute to Hepatocellular Carcinoma Risk in Hispanic Individuals. Cancers, 18(15), 2428. https://doi.org/10.3390/cancers18152428

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop