Sign in to use this feature.

Years

Between: -

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (537)

Search Parameters:
Journal = Cancers
Section = Cancer Informatics and Big Data

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 1971 KB  
Article
Hybrid Lexical–Semantic AI Architecture for Automated Cancer Registry Coding for the Vet-ICD-O-Canine-1 System from Free-Text Veterinary Pathology Reports
by Vitória Souza de Oliveira Nascimento, Marcello Vannucci Tedardi, Guilherme da Silva Rogério, Katia Cristina Pinello and Maria Lúcia Zaidan Dagli
Cancers 2026, 18(17), 2728; https://doi.org/10.3390/cancers18172728 (registering DOI) - 23 Aug 2026
Abstract
Background/Objectives: Free-text veterinary pathology diagnoses contain essential information for cancer registration but are difficult to convert into standardized ontology-based codes because of linguistic variability, contextual modifiers, and large ontology search spaces. This study evaluated a hybrid lexical–semantic architecture for the automated assignment of [...] Read more.
Background/Objectives: Free-text veterinary pathology diagnoses contain essential information for cancer registration but are difficult to convert into standardized ontology-based codes because of linguistic variability, contextual modifiers, and large ontology search spaces. This study evaluated a hybrid lexical–semantic architecture for the automated assignment of Vet-ICD-O-Canine-1 morphology codes. Methods: A retrospective single-registry benchmark included 211 diagnoses from the São Paulo Animal Cancer Registry. Of these, 190 contained sufficient morphological information for expert-reviewed reference coding, whereas 21 generic or insufficiently specified descriptions were retained as an exploratory challenge subset. Fuzzy lexical matching retrieved Top-10, Top-20, or Top-30 candidates from the complete 971-entry morphology ontology, followed by semantic selection using Claude Haiku 4.5 and structured JSON output. Performance and computational efficiency were compared to direct full-ontology inference. Results: Among the evaluated fuzzy metrics, token_set_ratio achieved the highest Top-30 reference-code retrieval rate of 89.5%. End-to-end exact-match agreement increased from 73.7% with Top-10 to 79.5% with Top-20 and 85.8% with Top-30 (95% CI, 80.1–90.0%). Top-30 generated non-null codes for 93.2% of the 190 evaluable diagnoses and achieved a conditional exact-match agreement of 92.1%. By contrast, the direct full-ontology baseline achieved 71.2% conditional exact-match agreement (42/59) among non-null predictions and 22.1% end-to-end exact-match agreement (42/190) when incorrect predictions, null outputs, and technical failures were considered non-concordant outcomes. Compared to direct full-ontology inference, Top-30 reduced input-token consumption by 92.4%, total token consumption by 92.2%, and inference cost by 91.3%, while avoiding the 118 API rate-limit failures observed with the direct baseline. Among the 21 insufficiently specified diagnoses, Top-30 returned null codes in 38.1% and non-null codes in 61.9%. Conclusions: Ontology-guided candidate reduction improved coding agreement, computational efficiency, and operational robustness within this retrospective single-registry benchmark. However, the reported performance estimates require confirmation in larger independent datasets, and an upstream data-sufficiency or abstention mechanism is needed before prospective operational deployment. Full article
Show Figures

Figure 1

50 pages, 31923 KB  
Article
A Fixed-Ratio Hybrid ARO-ALO Algorithm for Multi-Level Thresholding of Histopathological Colon Cancer Images
by Muhammed Faruk Şahin, Can Eyüpoğlu and Oktay Karakuş
Cancers 2026, 18(16), 2656; https://doi.org/10.3390/cancers18162656 - 17 Aug 2026
Viewed by 238
Abstract
Background/Objectives: Accurate segmentation of histopathological images while preserving cellular morphology in computer-aided diagnostic systems is critically important for the diagnosis and staging of colon cancer. However, conventional metaheuristic algorithms performing multi-level thresholding on such complex tissues often suffer from premature convergence by becoming [...] Read more.
Background/Objectives: Accurate segmentation of histopathological images while preserving cellular morphology in computer-aided diagnostic systems is critically important for the diagnosis and staging of colon cancer. However, conventional metaheuristic algorithms performing multi-level thresholding on such complex tissues often suffer from premature convergence by becoming trapped in local optima as the search space increases. To address this limitation, this study proposes a new label-independent hybrid optimization algorithm focused on colon adenocarcinoma segmentation. Methods: The proposed algorithm hybridizes the global exploration capability of the Artificial Rabbit Optimization (ARO) algorithm with the local exploitation ability of the Ant Lion Optimization (ALO) algorithm through an optimized fixed transition ratio, thereby enabling efficient localization of cellular density valleys. Results: The principal findings obtained from the LC25000 colon cancer dataset demonstrate that the ARO-ALO algorithm achieves stable performance with high SSIM (0.8043) and FSIM (0.8595) scores while preserving the histopathological hierarchy. Furthermore, the preservation of diagnostic morphology after segmentation is statistically validated by the high Pearson (0.9870) and Spearman (0.9948) correlation coefficients. In addition, supplementary generalization experiments are conducted on the Oral Squamous Cell Carcinoma (OSCC) and pulmonary circulation vessels datasets to verify the tissue-agnostic nature of the algorithm. Conclusions: Consequently, the ARO-ALO algorithm emerges as an efficient alternative for clinical decision support systems. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Figure 1

13 pages, 1143 KB  
Article
Using Decision Tree to Predict Cancer-Specific Mortality in Patients with Clear Cell Renal Cancer Treated with Nephrectomy
by Laura Martínez-Cayuelas, Pau Sarrio-Sanz, Jose-Vicente Segura-Heras, Milagros Muñoz-Montoya, Vicente-Francisco Gil-Guillen, Jesus Romero-Maroto and Luis Gomez-Perez
Cancers 2026, 18(16), 2644; https://doi.org/10.3390/cancers18162644 - 17 Aug 2026
Viewed by 205
Abstract
Background/Objectives: Accurate prognostic stratification after nephrectomy for clear cell renal carcinoma (ccRCC) remains challenging. Traditional models often lack the intuitive clinical application or the ability to handle non-linear interactions between variables. We aimed to develop and internally validate a decision tree-based model [...] Read more.
Background/Objectives: Accurate prognostic stratification after nephrectomy for clear cell renal carcinoma (ccRCC) remains challenging. Traditional models often lack the intuitive clinical application or the ability to handle non-linear interactions between variables. We aimed to develop and internally validate a decision tree-based model to predict cancer-specific survival in patients with ccRCC following nephrectomy. Methods: We analyzed 79,526 patients with ccRCC who underwent nephrectomy from the SEER database (2012–2018). Patients were randomized into development (2/3) and validation (1/3) cohorts. A conditional inference tree was constructed to predict cancer-specific survival. Multiple imputation by chained equations was used to handle missing data. Discriminatory ability was assessed using the C-index. Net clinical benefit was evaluated with decision curve analysis. The model was evaluated using CHARMS and PROBAST. Results: A decision tree with 15 risk groups is presented, further classified into high-, intermediate-, and low-risk categories according to observed median survival. The final predictors were tumor localization, tumor grade, TNM stage, age, and sarcomatoid differentiation. The model demonstrated excellent discriminatory performance, with a C-index of 0.846 (95% CI: 0.834–0.847). PROBAST assessment showed low risk of bias and low concern regarding applicability. Conclusions: The use of decision trees provides an interpretable alternative to conventional regression-based models. Three main risk categories and 15 subgroups are proposed based on tumor localization, tumor grade, TNM stage, age, and sarcomatoid differentiation. Our model demonstrates good applicability and a low risk of bias according to PROBAST guidelines; however, external validation in independent cohorts is required prior to clinical implementation. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Figure 1

26 pages, 11075 KB  
Article
Decitabine Reprograms Temozolomide-Resistant Glioblastoma Through Epigenetic Reactivation and Mesenchymal Attenuation: A Multi-Omics Study
by Itika Arora, Shamsa Hilal Saleh, Arshiya Akbar, Fareeha Arshad, Volodymyr Mavrych, Olena Bolgova, Faisal Abdulhameed Farrash, Ahmed Abu-Zaid, Andleeb Khan, Sheikh Muskan, Mohammed Imran Khan and Ahmed Yaqinuddin
Cancers 2026, 18(16), 2616; https://doi.org/10.3390/cancers18162616 - 14 Aug 2026
Viewed by 245
Abstract
Background/Objectives: Glioblastoma (GBM) is the most lethal primary brain malignancy in adults, with a median overall survival of approximately 15 months. Temozolomide (TMZ) resistance develops in virtually all patients, and no second-line regimen has improved outcomes over the past two decades. The [...] Read more.
Background/Objectives: Glioblastoma (GBM) is the most lethal primary brain malignancy in adults, with a median overall survival of approximately 15 months. Temozolomide (TMZ) resistance develops in virtually all patients, and no second-line regimen has improved outcomes over the past two decades. The DNA methyltransferase inhibitor decitabine (DAC) has attracted interest as a chemosensitizer, but whether it directly reverses the TMZ-resistance transcriptome or operates through distinct, complementary mechanisms has not been tested at multi-omics resolution. Methods: We performed an integrative six-layer multi-omics analysis across five public GEO datasets (bulk RNA-seq, EPIC 850K methylation, and 21,676 single cells) re-purposed from studies conducted for unrelated aims, formally tested DAC-mediated reversal of the TMZ-resistance transcriptome across 11,707 genes, mapped pharmacogenomic targets with DGIdb v5, and built an exploratory, hypothesis-generating 11-gene prognostic model internally validated in TCGA-GBM (n = 166) and externally tested in the independent CPTAC-GBM cohort (n = 96). Results: DAC reprogrammed transcription across 1114–1882 differentially expressed genes per cohort and reactivated 146 direct epigenetic targets, identifying INPP5D/SHIP1 as the top-ranked direct epigenetic-reactivation target. Genome-wide reversal analysis across 11,707 co-detected genes showed a negligible effect (Spearman ρ = 0.073), but single-cell analysis revealed significant per-cell attenuation of MES-like and stem-like programs (Δ = −0.071 and −0.135, respectively; both p < 0.001). The 11-gene risk model achieved a Harrell’s C-index of 0.706 (apparent); after correcting for the two-stage gene selection with a full-pipeline bootstrap, the optimism-corrected C-index was 0.63, and external validation in an independent cohort (CPTAC-GBM, n = 96) showed only near-chance discrimination (C-index 0.55), indicating that the signature does not generalize and is exploratory. Pharmacogenomic mapping yielded 734 unique therapeutic agents (230 FDA-approved) across 69 druggable targets after excluding AR. Most of these agents are not GBM-directed, so this catalog-level mapping is hypothesis-generating rather than a set of therapeutic recommendations. Conclusions: DAC does not broadly reverse the TMZ-resistant transcriptome but acts through three complementary mechanisms: epigenetic reactivation of INPP5D/SHIP1, cancer-testis-antigen and type I interferon induction, and per-cell attenuation of mesenchymal–stem-like transcriptional intensity, supporting hypotheses for rationally designed DAC-based combination therapy in TMZ-resistant GBM. Full article
Show Figures

Figure 1

19 pages, 2618 KB  
Article
Body Roundness Index and Incident Colorectal Cancer Risk: A Nationwide Population-Based Cohort Study
by Hyun Ho Kim, Changhyeok An, Kyu-na Lee, Kyung-do Han and Kang Woong Jun
Cancers 2026, 18(15), 2466; https://doi.org/10.3390/cancers18152466 - 31 Jul 2026
Viewed by 397
Abstract
Background/Objectives: Colorectal cancer (CRC) is among the most common cancers in Korea, and visceral adiposity is an established modifiable risk factor. The body roundness index (BRI), derived from waist circumference and height, estimates visceral adipose tissue more accurately than the body mass index [...] Read more.
Background/Objectives: Colorectal cancer (CRC) is among the most common cancers in Korea, and visceral adiposity is an established modifiable risk factor. The body roundness index (BRI), derived from waist circumference and height, estimates visceral adipose tissue more accurately than the body mass index (BMI) alone. We investigated the longitudinal association between the BRI and incident CRC in a large Korean national cohort. Methods: A nationwide cohort of 4,492,697 adults aged ≥20 years who participated in the 2012 Korean National Health Insurance Service health check-up was followed through December 2023 (mean 10.32 years). BRI values were categorized into quartiles (Q1–Q4). After applying a 1-year lag period, Cox proportional hazards regression estimated hazard ratios (HRs) and 95% confidence intervals (CIs), adjusting for age, sex, smoking, alcohol consumption, exercise, income, diabetes, hypertension, hypercholesterolemia, and chronic kidney disease. Subgroup and site-specific analyses were also performed. Results: During 46,382,686 person-years of follow-up, 57,644 incident CRCs were diagnosed. A higher BRI was independently and dose-dependently associated with CRC risk (Q4 vs. Q1: aHR 1.104, 95% CI 1.075–1.135; p for trend < 0.001). The association was significantly stronger in men (Q4 aHR 1.352, 95% CI 1.302–1.404) than in women (Q4 aHR 0.891, 95% CI 0.858–0.925; p for interaction < 0.001) and most pronounced in middle-aged adults (40–64 years). Site-specific analyses revealed a proximal-to-distal gradient: proximal colon (Q4 aHR 1.363), distal colon (Q4 aHR 1.244), and rectum (Q4 aHR 1.070). Conclusions: A higher BRI is independently and dose-dependently associated with increased CRC incidence in Korean adults, particularly in men and those with general obesity. The BRI may serve as a practical complement to the BMI for CRC risk stratification in populations with a high prevalence of metabolic obesity. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Figure 1

15 pages, 3738 KB  
Article
Burden, Trends, and Future Projections of Multiple Myeloma in China from 1990 to 2023
by Huiwen Liu, Jiazi Zhou, Hong Liu and Depei Wu
Cancers 2026, 18(15), 2449; https://doi.org/10.3390/cancers18152449 - 30 Jul 2026
Viewed by 377
Abstract
Background: Multiple myeloma (MM) is a malignant plasma cell disorder associated with substantial mortality and long-term disability. With rapid population aging in China, its disease burden may be increasing, yet comprehensive long-term assessments remain limited. Methods: Using data from the Global Burden of [...] Read more.
Background: Multiple myeloma (MM) is a malignant plasma cell disorder associated with substantial mortality and long-term disability. With rapid population aging in China, its disease burden may be increasing, yet comprehensive long-term assessments remain limited. Methods: Using data from the Global Burden of Disease Study 2023, we evaluated trends in prevalence, incidence, deaths, and disability-adjusted life years (DALYs) of MM in China from 1990 to 2023. Age-standardized rates were analyzed using Joinpoint regression to estimate temporal trends. Results: Between 1990 and 2023, prevalence, incidence, deaths, and DALYs of MM increased markedly in China, with consistent upward trends in age-standardized rates. The burden was concentrated among older adults and was higher in males than females. Decomposition analysis showed that epidemiological changes were the primary driver of increased DALYs and deaths. Forecasting analyses suggest that although age-standardized mortality and DALY rates may gradually decline, the absolute number of deaths is expected to continue rising. Conclusions: The burden of MM in China has increased substantially over the past three decades and is projected to remain high. Strengthening early detection, risk-oriented prevention, and long-term management strategies is essential to mitigate future disease burden. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Figure 1

16 pages, 16266 KB  
Article
Epidemiological Trends, Inter-Cancer Correlations, and Incidence Projections for 61 Cancer Types in Korea, 1999–2028: A Nationwide Population-Based Study
by Hyeran Jung and Minsun Jung
Cancers 2026, 18(14), 2341; https://doi.org/10.3390/cancers18142341 - 20 Jul 2026
Viewed by 436
Abstract
Background/Objectives: Korea has undergone rapid epidemiological transitions in cancer incidence over the past two decades. Using a 25-year nationwide dataset (1999–2023), we characterize long-term trends for 61 cancer types, examine inter-cancer correlations, and forecast incidence to 2028. Methods: Annual incidence counts, crude rates, [...] Read more.
Background/Objectives: Korea has undergone rapid epidemiological transitions in cancer incidence over the past two decades. Using a 25-year nationwide dataset (1999–2023), we characterize long-term trends for 61 cancer types, examine inter-cancer correlations, and forecast incidence to 2028. Methods: Annual incidence counts, crude rates, and age-standardized incidence rates (ASIRs) stratified by sex were obtained from the Korea Central Cancer Registry (KCCR) via the Korean Statistical Information Service (KOSIS). Annual percent change (APC) was estimated using log-linear regression. Pearson correlation coefficients were computed among cancer-specific ASIRs, with false-discovery-rate (FDR) correction for multiple comparisons. Multiple and hierarchical regression evaluated the statistical association of individual cancer types with the overall cancer rate, and variance inflation factors (VIFs) were used to quantify multicollinearity. Time series forecasting used damped Holt–Winters exponential smoothing; forecast accuracy was assessed with rolling-origin cross-validation (RMSE, MAE, MAPE) and benchmarked against ARIMA. A sensitivity analysis excluding the pandemic years (2020–2021) tested the robustness of trend estimates. Five-year prevalence data (2007–2023) were analyzed from the KCCR prevalence module. Results: Total cancer incidence increased from 101,854 in 1999 to 288,613 in 2023, a 183% increase. The overall ASIR rose from 402.7 to 522.9 per 100,000 (2020 standard population). The three fastest-growing cancers were thyroid (APC +7.56%, p < 0.001), prostate (+6.98%, p < 0.001), and breast (+5.03%, p < 0.001). Stomach (APC −2.20%) and liver (−2.90%) cancers showed significant declines. Hierarchical regression showed that adding thyroid, breast, and prostate to lung and stomach increased explained variance from R2 = 0.449 to 0.997; however, high VIF values (up to ~263) indicate substantial multicollinearity and compositional dependence, so these coefficients should not be read as independent causal contributions. Holt–Winters and ARIMA produced comparable accuracy (mean MAPE 4.6% vs. 4.7%). The five-year cancer prevalence pool reached 1,035,107 in 2023. Forecasting projects a total incidence of approximately 319,000 by 2028. Conclusions: Korean cancer epidemiology is undergoing a transition from infection-related cancers toward hormone-sensitive and screening-detectable malignancies. These findings support strategic resource allocation for high-growth cancers while maintaining vigilance over rising pancreatic and other emerging cancers. Full article
(This article belongs to the Special Issue Advances in Cancer Data and Statistics: 2nd Edition)
Show Figures

Figure 1

15 pages, 1827 KB  
Article
Initial Data Analysis for Cancer Registries: A Structured Framework and Demonstration Using Slovenian Cancer Registry Data
by Maja Jurtela, Lara Lusa, Tina Žagar, Nika Bric, Mojca Birk and Vesna Zadnik
Cancers 2026, 18(14), 2332; https://doi.org/10.3390/cancers18142332 - 20 Jul 2026
Viewed by 380
Abstract
Background/Objectives: Initial data analysis (IDA) is essential for valid and reproducible statistical analyses, but existing IDA frameworks were primarily developed for single-study datasets. Cancer registries (CRs) are extensive data systems characterized by continuous updates, repeated data extraction, multiple analytical uses, and evolving classification [...] Read more.
Background/Objectives: Initial data analysis (IDA) is essential for valid and reproducible statistical analyses, but existing IDA frameworks were primarily developed for single-study datasets. Cancer registries (CRs) are extensive data systems characterized by continuous updates, repeated data extraction, multiple analytical uses, and evolving classification systems, which create requirements not addressed by existing IDA frameworks. This study aims to develop a structured IDA framework adapted to CRs. Methods: We conceptualized IDA in CRs as a process spanning three data states: operational registry data, the extracted dataset and the analysis-ready dataset. The framework was developed by adapting existing IDA principles to the CR setting and organizing them into four stages: metadata, cleaning, screening, and reporting. The approach is based on predefined and versioned rule sets, structured recording of data processing, and metadata-based linkage between dataset definitions, data cleaning rules, screening outputs, the final report and dataset. A demonstrative survival dataset from the Slovenian Cancer Registry was used to illustrate implementation. Results: The framework is represented by the item set defining the activities and expected outputs of the structured IDA process in CRs. In the use case, metadata specified the dataset scope, intended use, variables, coding context, and applicable rules. Execution of the selected rules produced an analysis-ready dataset with traceable data cleaning steps, documented eligibility decisions, screening outputs describing the general and analysis-specific data properties, and the IDA report intended for external researchers to be delivered alongside the data. Conclusions: The proposed framework extends existing IDA approaches to meet the specific requirements of CRs. It supports consistent dataset preparation and with that improves transparent and reproducible data use. Full article
(This article belongs to the Special Issue Advances in Cancer Data and Statistics: 2nd Edition)
Show Figures

Figure 1

23 pages, 2922 KB  
Article
Explainable Machine Learning for Head and Neck Cancer Risk Stratification
by Amr Sayed Ghanem, Róbert Bata, Marianna Móré, Renáta Jávorné Erdei and Attila Csaba Nagy
Cancers 2026, 18(14), 2228; https://doi.org/10.3390/cancers18142228 - 11 Jul 2026
Viewed by 516
Abstract
Introduction: Head and neck cancers are frequently diagnosed at advanced stages, resulting in poor survival despite therapeutic advances, highlighting the need for earlier identification within routine clinical care. Routinely collected electronic health record data provide a potential source of longitudinal clinical signals, [...] Read more.
Introduction: Head and neck cancers are frequently diagnosed at advanced stages, resulting in poor survival despite therapeutic advances, highlighting the need for earlier identification within routine clinical care. Routinely collected electronic health record data provide a potential source of longitudinal clinical signals, although the clinical applicability of machine learning models remains limited by concerns regarding interpretability and real-world integration. Methods: This retrospective cohort study included 157,031 patients treated at the University of Debrecen Clinical Centre between 2007 and 2022, with head and neck cancer defined using ICD-10 codes C01 to C14. A total of 1397 variables, including demographic characteristics, ICD based comorbidities, and laboratory parameters, were reduced to 91 features using variance filtering and elastic net penalized Cox regression. Three survival modelling approaches were developed and compared, including CoxNet, Random Survival Forest, and XGBoost with a Cox objective. Results: The XGBoost model demonstrated the highest predictive performance with a mean concordance index of 0.916, followed by Random Survival Forest at 0.892 and CoxNet at 0.886, with acceptable calibration across models. Risk stratification showed clear separation between low, medium, and high-risk groups. Model interpretability using SHapley Additive exPlanations indicated that predictions were driven by a combination of demographic factors, laboratory markers, and clinically relevant diagnosis codes, reflecting both distal risk gradients and proximal clinical signals. Conclusions: These findings suggest that explainable machine learning applied to routine clinical data can support accurate and clinically interpretable risk stratification, with potential utility for opportunistic early identification of high-risk patients within existing healthcare pathways. Clinical Relevance: Explainable EHR-based survival models may support opportunistic identification of patients at increased head and neck cancer risk within routine clinical workflows, potentially improving triage and referral. Full article
(This article belongs to the Special Issue New Statistical and Machine Learning Methods for Cancer Research)
Show Figures

Figure 1

20 pages, 4457 KB  
Article
Early Classification of Bladder Cancer Using Spectrum-Aided Visual Enhancer (SAVE) and Deep Learning Models: A Non-Invasive Technology for Faster Detection
by Min-Hsin Yang, Yaswanth Nagisetti, Arvind Mukundan, Riya Karmakar, Chun-Feng Chang, Syna Syna, Ying-Jui Ni and Hsiang-Chen Wang
Cancers 2026, 18(13), 2147; https://doi.org/10.3390/cancers18132147 - 3 Jul 2026
Cited by 1 | Viewed by 507
Abstract
Background/Objectives: Bladder cancer (BC) is a significant global health issue, ranking as the ninth most prevalent cancer with a rising incidence. Conventional diagnostic methods, including cystoscopy and standard imaging techniques, possess limitations in identifying early cancer symptoms and accurately staging bladder cancer. Methods: [...] Read more.
Background/Objectives: Bladder cancer (BC) is a significant global health issue, ranking as the ninth most prevalent cancer with a rising incidence. Conventional diagnostic methods, including cystoscopy and standard imaging techniques, possess limitations in identifying early cancer symptoms and accurately staging bladder cancer. Methods: Consequently, this study developed a computer-aided diagnostic (CAD) system utilizing a novel, purely software-driven approach called Spectrum-Aided Vision Enhancer (SAVE) in conjunction with deep learning algorithms. Results: Our results demonstrate the profound potential of SAVE to expand diagnostic accessibility in medical imaging. While achieving an overall performance comparable to standard WLI (overall p = 0.41, indicating strong non-inferiority), SAVE demonstrated targeted absolute improvements in F1-scores for the most challenging early stage categories without requiring expensive optical equipment. For instance, in the ‘Above T1’ class, SAVE elevated the F1-score from 65% to 85% utilizing VGG16. Conclusions: These findings indicate that SAVE can provide reliable baseline detection while enhancing visual cues for specific complex lesions without requiring expensive optical equipment. For instance, in the ‘Above T1’ class, SAVE elevated the F1-score from 65% to 85% utilizing VGG16. These findings indicate that SAVE can provide resource-constrained clinical settings with advanced, high-contrast diagnostic capabilities, effectively decentralizing precision urological care. Full article
Show Figures

Figure 1

12 pages, 712 KB  
Article
Real-World Data of Comprehensive Genomic Profiles and Clinicopathological Characteristics of Duodenal Epithelial Neoplasms
by Marin Ishikawa, Hideyuki Hayashi, Kohei Nakamura, Ryutaro Kawano, Eriko Aimono and Hiroshi Nishihara
Cancers 2026, 18(13), 2097; https://doi.org/10.3390/cancers18132097 - 28 Jun 2026
Viewed by 335
Abstract
Background/Objectives: Duodenal epithelial neoplasms are rare; however, the widespread use of surveillance endoscopy and advances in endoscopic imaging technology have increased their incidental detection. Owing to their rarity, the clinicopathological characteristics and natural course of duodenal epithelial neoplasms have not been thoroughly [...] Read more.
Background/Objectives: Duodenal epithelial neoplasms are rare; however, the widespread use of surveillance endoscopy and advances in endoscopic imaging technology have increased their incidental detection. Owing to their rarity, the clinicopathological characteristics and natural course of duodenal epithelial neoplasms have not been thoroughly investigated. In this study, we aimed to clarify the genomic profile and clinicopathological characteristics of duodenal epithelial neoplasms. Methods: A total of 158 patients with duodenal epithelial neoplasms were enrolled. Comprehensive genomic profiling and immunohistochemical staining were performed. Immunophenotypes were categorized as gastric type (G-type), gastrointestinal type (GI-type), or intestinal type (I-type). The detection rate of potentially actionable genomic alterations and a high tumor mutational burden (TMB-H ≥ 10 Muts/Mb) were evaluated and compared across tumor types. Results: The median size of adenocarcinomas was larger than that of adenomas (p = 0.002). The age at diagnosis of G-type tumors was higher than that for the other two tumor types (p < 0.001). The median size of I-type tumors was smaller than that of the other two tumor types (p = 0.019). Compared with the other two types, G-type tumors were predominantly located in the superior region (p < 0.001), were macroscopic Type I (p = 0.002), and had significantly higher genomic alteration rates for KRAS (p < 0.001), GNAS (p < 0.001), CDKN2A (p = 0.004), and MDM2 (p < 0.001). Eighteen patients showed TMB-H. Conclusions: TMB-H was observed in >10% duodenal tumors. Additionally, the pathogenesis of G-type duodenal tumors differs from that of other immunophenotypic tumors. These findings could help in understanding the genomic profiles of duodenal tumors and in selecting treatment options. Full article
Show Figures

Graphical abstract

15 pages, 18043 KB  
Article
Breast Cancer Hormone Receptor Status Determination from H&E-Stained Biopsy Images Using Pixel-Level Classifiers
by Shuyang Wu, Ines P. Nearchou, Sandrine Prost, Jonathan A. Fallowfield, Hideki Ueno, Hitoshi Tsuda, Alastair Ironside, David J. Harrison and Timothy J. Kendall
Cancers 2026, 18(13), 2085; https://doi.org/10.3390/cancers18132085 - 27 Jun 2026
Viewed by 595
Abstract
Background: Analysis of digital images of histopathological sections is increasing due to widespread adoption of fully digitised workflows and the greater availability of whole-slide scanners. Currently, hormone receptor status in breast carcinoma is assessed by pathologists scoring separate immunohistochemically stained sections. Methods: In [...] Read more.
Background: Analysis of digital images of histopathological sections is increasing due to widespread adoption of fully digitised workflows and the greater availability of whole-slide scanners. Currently, hormone receptor status in breast carcinoma is assessed by pathologists scoring separate immunohistochemically stained sections. Methods: In this study, we employed pathologist-verified pixel-level annotations to train nested pixel classifiers capable of making case-level predictions of oestrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2) status directly from H&E-stained sections using biopsy cases alone. The model was evaluated on both an internal test set and an external international evaluation set from an institution in a different continent using different scanner hardware without the need for image normalisation. Results: In the internal test set, the models achieved AUCs of 0.8030, 0.7956 and 0.7488 for ER, PR, and HER2, respectively, with AUCs of 0.7008 and 0.7488 for ER and PR using an external cohort from an institution from which no cases were used for training. Conclusions: Our data highlight a potential strategy by which a pixel-based classifier, typically developed to quantify histological features within individual cases, could be used to make case/slide-level predictions but illustrate the challenges associated with this approach. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Figure 1

18 pages, 644 KB  
Article
Retrospective Cohort Study: Extracting Coexisting Background Breast-Lesion Features from Stage I–III Invasive Breast Cancer
by Ryan Jak Yang Lim, Phyu Nitar, Kah Weng Lau, Lester Chee Hao Leong, Veronique Kiak Mien Tan, Benita Kiat Tee Tan, Ern Yu Tan, Serene Si Ning Goh, Mikael Hartman, Fuh Yong Wong, Geok Hoon Lim, Jingmei Li and on behalf of the Joint Breast Cancer Registry
Cancers 2026, 18(12), 1965; https://doi.org/10.3390/cancers18121965 - 17 Jun 2026
Viewed by 477
Abstract
Background: Background breast features are frequently noted in pathology reports alongside invasive breast cancer but rarely factor into prognosis or treatment decisions. Their relationship to tumor characteristics and patient outcomes remains incompletely characterized. Methods: We conducted a retrospective cohort study of 7603 patients [...] Read more.
Background: Background breast features are frequently noted in pathology reports alongside invasive breast cancer but rarely factor into prognosis or treatment decisions. Their relationship to tumor characteristics and patient outcomes remains incompletely characterized. Methods: We conducted a retrospective cohort study of 7603 patients with Stage I–III invasive breast cancer (diagnosed 1991–2022, age < 80 years) from the Joint Breast Cancer Registry in Singapore. Natural language processing (NLP) was applied to 9754 free-text pathology reports to extract co-existing background breast features, with accuracy validated by dual-reviewer assessment of 200 reports. Because background features are most reliably assessed on excision specimens, the primary analytic cohort comprised 3988 patients with available excision pathology reports. Unsupervised hierarchical clustering grouped extracted features into three categories. Associations with tumor characteristics were assessed with multinomial logistic regression and ten-year overall survival by Cox proportional hazards models (median follow-up 9.6 years; 620 deaths). Results: Here, we show that NLP-based extraction of background breast features from routine pathology reports achieves an accuracy of over 90% across features. Lobular neoplasia and benign proliferative changes are associated with less aggressive tumor characteristics, whereas early neoplastic and papillary lesions are more prevalent in HER2-enriched and luminal B tumor subtypes. Benign proliferative changes are associated with better survival in age- and year-adjusted models (hazard ratio 0.91, 95% CI 0.86–0.97), but this association is attenuated after adjustment for stage and subtype. Conclusions: NLP-enabled extraction of background breast features from pathology text is feasible at scale. These features reflect tumor biology but do not independently add prognostic information beyond established clinical variables. Full article
(This article belongs to the Special Issue Advances in Cancer Data and Statistics: 2nd Edition)
Show Figures

Figure 1

14 pages, 1447 KB  
Article
Multi-Model Machine Learning for Survival Predictions for Castration-Resistant Prostate Cancer
by Tae Jin Kim, Jaeyun Jeong, Young Jin Ahn, Kwang Suk Lee, Jong Soo Lee, Seung Hwan Lee, Won Sik Ham, Byung Ha Chung, Jeong Hyun Lee and Kyo Chul Koo
Cancers 2026, 18(12), 1866; https://doi.org/10.3390/cancers18121866 - 7 Jun 2026
Viewed by 508
Abstract
Background: Accurate survival prediction is essential for optimizing treatment planning in patients with castration-resistant prostate cancer (CRPC). However, traditional statistical models often underperform because of limited variable inclusion and an inability to account for complex, multidimensional data interactions. Methods: We retrospectively collected 46 [...] Read more.
Background: Accurate survival prediction is essential for optimizing treatment planning in patients with castration-resistant prostate cancer (CRPC). However, traditional statistical models often underperform because of limited variable inclusion and an inability to account for complex, multidimensional data interactions. Methods: We retrospectively collected 46 clinical, laboratory, and pathological variables from 801 patients with CRPC, covering the disease course from initial diagnosis to CRPC progression. Multiple machine learning (ML) models, including random survival forests (RSF), XGBoost, LightGBM, and logistic regression, were developed to predict cancer-specific mortality (CSM), overall mortality (OM), and 2- and 3-year survival status. The dataset was divided into training and test cohorts (80:20), and 10-fold cross-validation was performed. Performance was assessed using the C-index for regression models and the area under the curve (AUC), accuracy, precision, recall, and F1-score for classification models. Model interpretability was evaluated using SHapley Additive exPlanations (SHAP). Results: Over a median follow-up of 24 months, 70.6% of patients experienced CSM. Although XGBoost with its own imputation method achieved the highest C-index in the validation set, RSF demonstrated more stable performance and achieved the highest C-index in the held-out test set for both CSM (0.772) and OM (0.771). For classification tasks, RSF demonstrated superior performance in predicting 2-year survival, whereas XGBoost achieved the highest F1-score for 3-year survival prediction. SHAP analysis identified time to first-line CRPC treatment, hemoglobin level, and alkaline phosphatase level as key predictors of survival outcomes. Conclusions: RSF demonstrated robust test-set performance for time-to-event prediction, whereas XGBoost showed complementary value for 3-year survival classification. These models provide accurate and interpretable prognostic tools that may support personalized treatment strategies. External validation and integration of emerging therapies are warranted to enhance broader clinical applicability. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Show Figures

Graphical abstract

10 pages, 224 KB  
Review
Current Clinical Utility of Gene Expression Panels in Primary Cutaneous Melanoma
by Taylor L. Garza, Jae Hwan Choi, Michelle McGee, Edmond Box, Timothy Nywening, Andreas Karachristos, Abraham Schwarzberg, Mayer Fishman and Richard Jacobson
Cancers 2026, 18(11), 1830; https://doi.org/10.3390/cancers18111830 - 3 Jun 2026
Viewed by 622
Abstract
Roughly 100,000 new cases of melanoma are diagnosed yearly in the United States, the majority of which are early-stage disease with excellent prognosis. However, a subset of patients harbors clinically occult aggressive biology that can go undetected on standard clinicopathologic analyses. Gene expression [...] Read more.
Roughly 100,000 new cases of melanoma are diagnosed yearly in the United States, the majority of which are early-stage disease with excellent prognosis. However, a subset of patients harbors clinically occult aggressive biology that can go undetected on standard clinicopathologic analyses. Gene expression profiling (GEP) assays have emerged as molecular adjuncts to risk stratification, aiming to improve sentinel lymph node biopsy (SLNB) decision-making and surveillance planning. Two commercial tests are widely available: the 31-gene expression profile (31-GEP, DecisionDx-Melanoma, Castle Biosciences) and the clinicopathologic gene expression profile (CP-GEP, Merlin, SkylineDx). This review summarizes the current evidence for each assay regarding their performance and utility, with a focus on potentially actionable use cases, limitations, and the practical context in which these tools are most valuable. Full article
(This article belongs to the Section Cancer Informatics and Big Data)
Back to TopTop