1. Introduction
Liver cancer causes over 800,000 new cases and more than 750,000 deaths each year worldwide, accounting for 4.3% of incident cancers and 7.3% of cancer deaths [
1]. Hepatocellular carcinoma (HCC) dominates among all liver cancer cases. It is well established that early-stage disease can be controlled through surgery, ablation, or systemic treatment, and the survival rate is also good. But when it comes to advanced stages, the vast majority of patients, even liver cancer specialists, feel helpless. Even with modern regimens that combine chemotherapy, targeted agents, and checkpoint inhibitors, survival stays stubbornly short in advanced disease [
2,
3,
4]. Better risk tools are needed, not just for predicting outcomes but for matching patients to the right intensity of treatment.
In clinical practice, most patients with portal vein tumor thrombus (PVTT) or extrahepatic spread fall into Barcelona Clinic Liver Cancer (BCLC) stage C [
5,
6,
7]. The thrombus itself drives intrahepatic dissemination, decompensation, portal hypertension, ascites, and variceal bleeding [
8,
9]. Cancer progression is rapid, but treatment options are limited. Luckily, doctors have found that some patients survive considerably longer than the stage labels imply, but we lack simple bedside tools to distinguish them. If those who may benefit from more aggressive treatment measures can be identified at the initial stage of treatment, it will help improve the overall prognosis of the entire patient population. Of course, there are many factors that contribute to the differences in prognosis between different BCLC stage C patients.
We believe metabolism is part of the explanation, though not the whole story. The liver handles glucose, lipids, and proteins. HCC disrupts multiple important metabolic pathways. Insulin resistance, hyperglycemia, dyslipidemia, and hypoalbuminemia appear frequently, reflecting both tumor demand and hepatic reserve [
10,
11]. There is some evidence showing that metabolic syndrome has been tied to HCC risk, treatment response, and long-term outcome [
12,
13,
14]. In advanced stages, chronic inflammation, malnutrition, and therapy-related stress also add further disruption [
15,
16,
17]. We can consider that circulating metabolic markers carry dual information: they signal tumor burden and liver capacity at the same time. That makes them plausible prognostic candidates. Fasting glucose, triglycerides, albumin, and composite indices like triglyceride-glucose (TyG) have all been tested for prognostic value in HCC [
18,
19,
20]. Some are independently associated with survival [
21,
22,
23]. But any single marker captures only a slice of the metabolic picture. Combining several clinical variables is one way forward.
Another valuable direction for exploration is to look beyond bloodwork: artificial intelligence (AI) applied to routine computed tomography (CT) can extract quantitative features that the human eye misses. The core essence of tumors is still the pathological nature of the tumor itself, and imaging can partially display it in a macroscopic dimension. Radiomics extracts high-dimensional numerical descriptors of tumor shape, texture, and heterogeneity from standard CT or magnetic resonance imaging (MRI) [
24,
25,
26]. These can be distilled into a radiomics score. In HCC, such scores have shown promise for predicting microvascular invasion, early recurrence, and treatment response [
27,
28,
29]. What is less settled is whether adding a CT-based radiomics score to routine clinical and metabolic variables actually sharpens survival prediction in patients who already have PVTT. Answering that question would fit squarely into the push for AI-driven biomarkers in cancer care.
A model that combines metabolic variables, radiomics features, and standard clinical indicators has not been tested in HCC patients with PVTT. In this study, we attempted to build this model. Our goal was practical from the start: a tool that uses what is already collected—a blood test and a portal-venous phase CT—and returns individualized survival estimates. A more precise prognostic tool could eventually help guide treatment selection, including for aggressive therapies such as immunotherapy.
2. Materials and Methods
2.1. Patients Selection
Between January 2015 and December 2019, we pulled records of primary HCC patients who had undergone contrast-enhanced abdominal CT at West China Hospital. To be eligible, patients had to have pathologically confirmed HCC, CT evidence of PVTT, complete 5-year survival follow-up data, and routine laboratory tests within three months before CT. We excluded cases with indeterminate PVTT pathology, mixed tumor and bland portal vein thrombus, and any anti-tumor therapy—including transarterial chemoembolization (TACE), radiotherapy, targeted therapy, immunotherapy, or local ablation—within three months before imaging. We excluded patients whose last documented follow-up was shorter than six months from treatment initiation, as their survival status could not be reliably ascertained beyond the last known contact. Patients with complete event data—including those who died within six months—were retained in the cohort. The remaining patients were randomly split into a training set and a validation set. Hereafter in this manuscript, the term ‘validation set’ refers to this internal validation set derived from the same single-center cohort by random split, rather than an external validation set.
2.2. Data Collection
Demographic, clinical, imaging, and laboratory data were pulled from the hospital’s electronic medical record system, all obtained within three months before treatment. Lab tests were run on the hospital’s automated platform. Overall survival was measured from the date of first treatment until death from any cause or last follow-up.
2.3. CT Acquisition and Radiomics Workflow
Portal-venous phase CT images were used for radiomics analysis. All scans were performed using 64-detector row scanners of the same model following the center’s standard protocol. Detailed CT acquisition and reconstruction parameters, including contrast administration, scan timing, and imaging settings, are provided in
Supplementary Material 1. A radiologist with experience in liver imaging manually segmented the primary tumor volume, avoiding vessels and necrosis. A second radiologist reviewed the contours, and disagreements were settled through discussion. Using LIFEx software (version 3.74; CEA-SHFJ, Orsay, France), we extracted a battery of shape, first-order, and texture features, as well as wavelet-transformed variants. All features were normalized to zero mean and unit variance using the training set’s statistics.
2.4. Radiomics Score Construction
In the training set, a LASSO Cox regression model, a supervised machine learning algorithm, was applied to the normalized radiomic features [
30]. Ten-fold cross-validation selected the penalty parameter λ, and features with non-zero coefficients at that λ were retained. These features were linearly combined with their LASSO coefficients to form a radiomics score, or Rad-score. The same normalization parameters and linear formula were then applied to the validation set without modification.
2.5. Clinical Predictor Selection
We first screened clinical variables for collinearity using the variance inflation factor (VIF). Variables with VIF exceeding 5 were excluded from further analysis to address multicollinearity, which can destabilize coefficient estimates in Cox regression. The remaining variables then underwent univariate Cox regression, and those that met a relaxed significance level were carried forward into a multivariate Cox model. Backward elimination was then applied, retaining only predictors that remained significant at a conventional threshold. This procedure produced the clinical model.
2.6. Model Building and Evaluation
Three Cox proportional-hazards models were constructed in the training set. The clinical model contained the predictors retained after backward elimination. The Rad-score model contained only the continuous Rad-score. The combined model included all variables from the clinical model together with the Rad-score, entered simultaneously without further selection. The same model formulas were used in the validation set without refitting.
Discrimination was assessed with C-index and time-dependent AUC at two landmark time points, both in the training set and in the validation set. Calibration was assessed using the calibrate() function from the rms package with 200 bootstrap resamples. Calibration curves were generated by plotting observed survival probabilities against predicted probabilities, stratified by risk deciles. The calibration intercept and slope were derived from a linear regression of observed outcomes on predicted probabilities. Under this parameterization, an ideal intercept of 0 and an ideal slope of 1 indicate perfect calibration. An intercept > 0 indicates systematic underestimation of survival probabilities, while an intercept < 0 indicates systematic overestimation. A slope > 1 indicates overdispersion—predicted probabilities are too extreme—whereas a slope < 1 indicates underdispersion. Decision curve analysis was used to evaluate net benefit across a range of threshold probabilities. Risk scores from each model were dichotomized by the median of a time-dependent ROC curve, and Kaplan–Meier curves with log-rank tests were plotted for both groupings.
All statistical analyses were performed in R (version 4.6.0; R Foundation for Statistical Computing, Vienna, Austria). A two-sided p < 0.05 was considered statistically significant.
4. Discussion
We compared three survival models in HCC patients with PVTT, and the combined model came out ahead on discrimination, calibration, and net benefit. Four clinical variables and a radiomics score built from two CT texture features were included in the final model. Routine lab values and AI-pulled imaging features thus appear to carry complementary prognostic weight in this setting.
PVTT is a turning point. It speeds up intrahepatic spread, worsens portal hypertension, and cuts short the window for effective therapy [
31,
32]. Even with checkpoint inhibitors added to the arsenal, survival above 20 months is uncommon. Still, the “advanced” label hides considerable heterogeneity. Some patients live far longer than the median. Tools that can separate them out, especially tools built from data that are already available, could help clinicians decide how aggressively to treat without piling on extra cost. Before discussing the model’s performance, it is important to acknowledge several constraints related to patient selection. The requirement for pathological confirmation of PVTT excluded a substantial proportion of patients (1089/2157, 50.5%) and may have selected a healthier, more procedure-tolerant subgroup, limiting generalizability. While this choice ensured diagnostic accuracy—distinguishing tumor thrombus from bland thrombus—it does not reflect routine clinical practice where imaging-based diagnosis is standard. Similarly, we excluded 189 patients with follow-up shorter than six months and 254 who had received recent anti-tumor therapy (including TACE, radiotherapy, targeted therapy, immunotherapy, or local ablation) within three months prior to CT. These exclusions were intended to ensure that baseline data reflected the untreated tumor state and to maintain completeness of follow-up data. However, combined with the pathological confirmation requirement, these criteria narrowed the cohort to a highly selected population (134/2157, 6.2%), which may limit the generalizability of our findings. We also acknowledge that a formal comparison of baseline characteristics between included and excluded patients could not be performed. A substantial proportion of excluded patients had incomplete baseline data at the time of their initial visit—many were lost to follow-up or transferred to other institutions before completing the recommended evaluations—which was one of the reasons for their exclusion. Thus, we cannot formally assess whether the analytic cohort is fully representative of the entire screened PVTT population. We also acknowledge that only 35.1% of patients had documented evidence of cirrhosis, whereas all patients were Child–Pugh B or C. Cirrhosis was recorded based on any available clinical evidence (history, imaging, or clinical assessment). However, all patients underwent contrast-enhanced CT as part of the study inclusion criteria; the 35.1% figure reflects those with clearly documented cirrhosis in the medical records, not the prevalence of cirrhosis. Hepatic dysfunction in this cohort may therefore reflect a combination of chronic liver disease, tumor burden, and PVTT-related portal hypertension, rather than cirrhosis alone. These factors should be considered when interpreting the model’s performance.
The clinical predictors that survived backward elimination all have a plausible biological footing. AFP is the oldest biomarker in HCC, and its independent link to survival held in our cohort [
33,
34,
35]. BMI and HDL-C were both protective. A higher BMI in advanced HCC is unlikely to reflect classic obesity-related risk; it more likely signals preserved lean mass and nutritional buffer that help patients withstand the metabolic drain of a large tumor. This aligns with the “obesity paradox” noted by several groups [
36,
37,
38]. HDL-C may point in a similar direction. Low HDL-C has been linked to systemic inflammation and cancer cachexia. In our data, each increment in HDL-C was tied to a sizable drop in mortality, consistent with reports that lipid handling is deeply altered in advanced liver cancer [
39,
40]. Alkaline phosphatase, the fourth predictor, could flag biliary obstruction, heavy hepatic infiltration, or bone involvement, all of which worsen the outlook [
41].
The radiomics part, though pared down to two features, still picked up a prognostic signal that the blood variables missed. GLRLM_SRHGE and GLZLM_SZHGE belong to the high-gray-level run-length and zone-size families. Their positive coefficients mean that tumors with coarser texture and more frequent high-intensity runs on CT foreshadow higher mortality. Other groups have reported that similar texture features carry prognostic value across several cancers, HCC included [
42,
43,
44]. Biologically, these features may encode necrosis, hemorrhage, or architectural disarray that routine radiology does not formally grade [
45,
46,
47]. A practical point worth underlining: the radiomics score came from a single portal-venous phase CT, a scan already embedded in the standard workup for PVTT. No extra imaging, no additional contrast, no biopsy. That makes the approach compatible with everyday clinical workflow rather than a research-only exercise. Several technical limitations of the radiomic analysis should be noted. Formal reproducibility metrics for ROI delineation (such as ICC or Dice coefficients) were not calculated because independent duplicate segmentations were not preserved, although contours were reviewed by two experienced radiologists with consensus resolution. Detailed CT acquisition and reconstruction parameters are provided in
Supplementary Material 1; all scans were performed using the same scanner model following the center’s standard protocol, so no scanner-specific harmonization (e.g., ComBat) was required. Furthermore, the exact identity of individual features selected by LASSO may be sensitive to sampling variability—using 1000 bootstrap resamples with λ.min, the two retained features showed re-selection frequencies of 61.7% and 26.8%—though the composite Rad-score remained robust (optimism-corrected C-index = 0.830). Finally, radiomics feature interpretation, while tied to texture theory, remains inferential without histopathological correlation.
Our combined model yielded a C-index of 0.843 in the training set and 0.815 in the validation set, with 1-year AUCs above 0.94. These numbers sit in the same range as—and sometimes above—published models that depend on genomic panels, circulating tumor DNA, or advanced MRI techniques. Those tools are powerful but are also expensive and geographically concentrated. Our predictors come from a basic blood draw and a single contrast-enhanced CT. The cost difference matters: HCC is most prevalent in regions where healthcare budgets are thin, and a model that runs on what is already collected can reach far more patients. Calibration assessment showed slopes ranging from 1.044 to 1.291 and intercepts from 1.030 to 1.222. These values indicate a tendency toward mild overdispersion, particularly in the validation set. While this suggests that the nomogram’s absolute probability estimates should be interpreted with caution and external recalibration is recommended, the model maintained strong rank-ordering ability (as reflected by the C-index and AUC) and demonstrated net benefit on decision curve analysis. The decision curves reinforced this picture: across thresholds up to 50%, the combined model net benefit stayed above the single-domain models, and all three models beat the treat-all and treat-none strategies.
Two external benchmarks help place our results in context. A multicenter study of 1026 HBV-related HCC patients without PVTT identified male sex, low albumin-to-alkaline-phosphatase ratio, elevated aspartate-to-platelet ratio, extrahepatic metastasis, and multiple tumors as risk factors for PVTT occurrence, and generated nomograms for PFS and OS with C-indexes around 0.72–0.80 [
48]. That study asked who will develop PVTT. Our study starts where that question ends, focusing on patients who already have the thrombus. In this population, our combined model reached C-indices of 0.843 (training) and 0.815 (validation), which are descriptively higher than their OS nomogram. Another study, also in HCC patients with PVTT, built a radiomics-based nomogram that included clinical and radiotherapy dosimetric variables for predicting OS after radiation, with a C-index of 0.73 [
49]. Our model also showed descriptively higher C-indices and did not require dosimetric data, so it can be applied before treatment selection rather than being tied to a specific modality. It should be noted, however, that these comparisons are descriptive only; formal statistical testing could not be performed because individual patient data from the external studies were not accessible. We acknowledge that the discriminative performance observed in our study—particularly the C-index of 0.843 in the training set and 0.815 in the validation set—may be partly optimistic. Several factors may contribute to this. First, the stringent inclusion criteria (pathological confirmation, exclusion of recent therapy, and complete follow-up) selected a relatively homogeneous patient subset, which can inflate performance relative to unselected real-world populations. Second, the validation set was derived from a random split of the same single-center cohort rather than an external dataset, which may underestimate performance degradation in independent settings. Third, the moderate sample size (
n = 94 for training) increases the risk of overfitting, despite the use of LASSO regularization and bootstrap optimism-correction. Finally, the multi-step modeling process—involving VIF screening, univariate selection, backward elimination, and LASSO—may have capitalized on chance associations specific to this cohort. To mitigate overfitting, we applied LASSO with 10-fold cross-validation and reported optimism-corrected C-index (0.830), which remains robust. Nevertheless, cautious interpretation is warranted, and external validation in larger, multicenter cohorts is essential before clinical implementation. Taken together, the two external comparisons suggest that blending routine laboratory values with CT texture features may sharpen survival prediction in this high-risk group, although head-to-head validation in a shared dataset is needed to confirm this observation.
Several additional limitations merit consideration. The study was retrospective and single-center, and the internal validation set was derived from the same cohort by random split, which cannot demonstrate external generalizability. The sample size was modest (n = 94 for training), and the model’s calibration showed a tendency toward overdispersion (slopes up to 1.291), suggesting that absolute survival probability estimates should be interpreted with caution and external recalibration is warranted. The positive intercepts (1.030–1.222) indicate that the model systematically underestimates survival probabilities, further supporting the need for caution when interpreting absolute survival estimates. Tumor thrombus extent was not graded by a formal system. Schoenfeld residual analysis detected time-varying effects for ALP and Rad-score. This is clinically plausible: ALP levels may fluctuate with biliary obstruction or bone involvement during disease progression, while Rad-score—derived from baseline CT—primarily reflects initial tumor heterogeneity and necrosis, which exert the strongest prognostic influence during the early post-treatment period. As time elapses, subsequent therapies and disease progression increasingly shape long-term survival. This pattern does not compromise the model’s intended use as a baseline risk stratification tool, as the strongest predictive signal occurs during the critical early treatment decision window. However, caution is warranted when extrapolating predictions to very long-term follow-up. Importantly, we did not systematically capture post-enrollment anticancer treatments or concomitant medications, which may have influenced long-term survival and metabolic markers. Despite these limitations, the model maintained strong discriminative performance and net benefit on decision curve analysis. Finally, multicenter prospective studies with external CT datasets are the logical next step to validate and refine these findings.