Refined Graph-Guided Fusion Network for Explainable Multimodal Lung Cancer Classification Using CT Imaging and Semantic Features
Abstract
1. Introduction
1.1. Motivation and Challenges
1.2. Novelty
- This work proposes a novel model, the Refined Graph-Guided Fusion Network (R-GGFN), to represent 3D CT images and doctors’ descriptions of pulmonary nodule structures and to fuse multiple images to classify pulmonary nodules.
- In multimodal feature fusion and gradient propagation, preserving discriminative radiologist-derived semantic information through a tabular skip connection is a key challenge, so we add a tabular skip-connection mechanism.
- We use an optimized training strategy that includes focal loss, weight decay, a reduced learning rate, and the ReduceLROnPlateau scheduler to ensure stable convergence, counteract class imbalance, and improve model generalization.
- The explainability and transparency of multimodal lung cancer classification are enhanced by designing a unified Explainable Artificial Intelligence (AI) model, which combines 3D Grad-CAM, SHAP, and graph attention visualization.
2. Related Work
2.1. Image-Based Deep Learning
2.2. Multimodal Information Fusion
2.3. Graph Neural Networks
2.4. Explainable Artificial Intelligence
2.5. Comparative Analysis of CNN, Graph Neural Network, Multimodal Fusion, and XAI
2.6. Feature Comparison of Existing Methods
2.7. Research Gap
3. Methodology
3.1. Experimental Setup
3.2. Dataset Description
3.3. Data Preprocessing
- Reconciled loading of volumetric CT and semantic annotation by the radiologist.
- Normalized the intensity distribution of images.
- Built a volume of images from a series of X-ray CT images.
- Multi-annotation as a single semantic feature vector for multiple training runs done by different radiologists.
- Eliminating data leakage by doing patient-level partitioning.
- Using only data augmentation for the training set.
- Weighted random sampling and focal loss to tackle class imbalance.
- Image Modality Preprocessing
3.3.1. Data Augmentation
3.3.2. Tabular Modality Preprocessing
3.3.3. Cross-Radiologist Annotation Aggregation
3.3.4. Patient-Level Dataset Partitioning
3.3.5. Class Imbalance Handling
4. Proposed Refined Graph-Guided Fusion Network (R-GGFN)
4.1. Overview of the Proposed Framework
- A 3D ResNet-18 backbone for extracting volumetric image features.
- A multilayer perceptron (MLP) for learning semantic representations from radiologist annotations.
- A Graph Attention Network (GAT) that models contextual relationships among pulmonary nodules.
- A refined classifier incorporating a tabular skip connection, enabling direct preservation of clinically meaningful semantic information throughout graph-based feature learning.
4.2. Image Feature Extraction
4.3. Semantic Feature Learning
4.4. Multimodal Feature Fusion
4.5. Graph-Guided Contextual Learning
4.6. Tabular Skip Connection
- Graph-enhanced contextual features;
- Original radiologist semantic information;
- Multimodal image representations.
4.7. Model Optimization
- Adam optimizer;
- Adaptive learning-rate scheduling (ReduceLROnPlateau);
- Weight decay regularization;
- Gradient clipping;
- Early convergence monitoring.
4.8. Explainable Artificial Intelligence Techniques
5. Experimental Results
5.1. Baseline Model Performance Result
5.2. Multimodal Fusion Results (Late Fusion and Gated Fusion)
5.2.1. Late Fusion
5.2.2. Gated Fusion
5.3. Naïve GGFN Training Results
5.4. R-GGFN Result
5.5. Comparative Performance: All Models
5.5.1. Statistical Significance of the Model Comparisons
5.5.2. Comparison with State-of-the-Art Methods
5.6. Results and Interpretation with Explainable AI
5.6.1. GradCAM-3D Image Branch Feature
5.6.2. SHAP: Tabular Branch Feature Importance
5.6.3. GAT Attention: Graph Branch Interpretability
6. Discussion
7. Conclusions
7.1. Limitations
7.2. Future Recommendation
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- World Health Organization. Lung Cancer (Fact Sheet). Available online: https://www.who.int/news-room/fact-sheets/detail/lung-cancer (accessed on 12 August 2026).
- Alsatari, E.S.; Smith, K.R.; Galappaththi, S.P.L.; Turbat-Herrera, E.A.; Dasgupta, S. The Current Roadmap of Lung Cancer Biology, Genomics and Racial Disparity. Int. J. Mol. Sci. 2025, 26, 3818. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gilad, S.; Lithwick-Yanai, G.; Barshack, I.; Benjamin, S.; Krivitsky, I.; Edmonston, T.B.; Bibbo, M.; Thurm, C.; Horowitz, L.; Huang, Y.; et al. Classification of the Four Main Types of Lung Cancer Using a MicroRNA-Based Diagnostic Assay. J. Mol. Diagn. 2012, 14, 510–517. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hussein, S.; Cao, K.; Song, Q.; Bagci, U. Risk Stratification of Lung Nodules Using 3D CNN-Based Multi-Task Learning. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Cham, Switzerland, 2017. [Google Scholar] [CrossRef] [Scilit]
- Armato, S.G., III; McLennan, G.; Bidaut, L.; McNitt-Gray, M.F.; Meyer, C.R.; Reeves, A.P.; Zhao, B.; Aberle, D.R.; Henschke, C.I.; Hoffman, E.A.; et al. The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A Completed Reference Database of Lung Nodules on CT Scans. Med. Phys. 2011, 38, 915–931. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
- Rahane, W.; Dalvi, H.; Magar, Y.; Kalane, A.; Jondhale, S. Lung Cancer Detection Using Image Processing and Machine Learning Healthcare. In 2018 International Conference on Current Trends Towards Converging Technology (ICCTCT); IEEE: New York, NY, USA, 2018; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
- Ioffe, S.; Szegedy, C. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In 32nd International Conference on Machine Learning (ICML); PMLR: Cambridge, MA, USA, 2015; pp. 448–456. [Google Scholar]
- Arevalo, J.; Solorio, T.; Montes-y-Gómez, M.; González, F.A. Gated Multimodal Units for Information Fusion. In Proceedings of the 5th International Conference on Learning Representations (ICLR) Workshop, Toulon, France, 24–26 April 2017. [Google Scholar]
- Rahman, M.; YongZhong, C.; Bin, L. Graph Attention Network-Based Multimodal Approach for Lung Diseases Classification. Sci. Rep. 2026, 16, 10914. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, W.; Zhou, M.; Yang, F.; Yang, C.; Tian, J. Multi-Scale Convolutional Neural Networks for Lung Nodule Classification. In International Conference on Information Processing in Medical Imaging (IPMI); Springer: Cham, Switzerland, 2015; pp. 588–599. [Google Scholar]
- Zhang, C.; Aamir, M.; Guan, Y.; Al-Razgan, M.; Awwad, E.M.; Ullah, R.; Bhatti, U.A.; Ghadi, Y.Y. Enhancing Lung Cancer Diagnosis with Data Fusion and Mobile Edge Computing Using DenseNet and CNN. J. Cloud Comput. 2024, 13, 91, Correction in J. Cloud Comput. 2024, 13, 111. https://doi.org/10.1186/s13677-024-00673-1. [Google Scholar] [CrossRef] [Scilit]
- Saihood, A.; Hasan, M.A.; Shnawa, S.M.; Fadhel, M.A.; Alzubaid, L.; Gupta, A.; Gu, Y. Multiside Graph Neural Network-Based Attention for Local Co-Occurrence Features Fusion in Lung Nodule Classification. Expert Syst. Appl. 2024, 252, 124149. [Google Scholar] [CrossRef] [Scilit]
- Sousa, J.V.; Matos, P.; Silva, F.; Freitas, P.; Oliveira, H.P.; Pereira, T. Single Modality vs. Multimodality: What Works Best for Lung Cancer Screening? Sensors 2023, 23, 5597. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oncu, E. Multimodal AI Framework for Lung Cancer Diagnosis: Integrating CNN and ANN Models for Imaging and Clinical Data Analysis. Preprint 2024. [Google Scholar] [CrossRef] [Scilit]
- Dubey, R.; Dhaka, A.; Nandal, A.; Sharma, A.K. An Attention-Guided Multimodal Deep Learning Framework by Integrating CT-PET Imaging and Clinical Data for Lung Cancer Detection. Sci. Rep. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sun, M.; Cui, C. Advanced AI-Driven Image Fusion Techniques in Lung Cancer Diagnostics: Systematic Review and Meta-Analysis for Precision Medicine. Robot. Intell. Autom. 2024, 44, 579–593. [Google Scholar] [CrossRef] [Scilit]
- Pushpa, M.; Gomathi, P.R.; Kota, P.S.; Kumar, D.A.; Pundir, S. Integrating Deep Learning and Graph Neural Networks for Multimodal Lung Tumor Analysis: A Novel Approach for Improved Classification and Predict. In International Conference on Smart Systems and Advanced Applications; IEEE: New York, NY, USA, 2023; pp. 346–352. [Google Scholar] [CrossRef] [Scilit]
- Waqas, A.; Tripathi, A.; Ramachandran, R.; Stewart, P.; Rasool, G. Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review. arXiv 2023, arXiv:2303.06471. [Google Scholar] [CrossRef] [Scilit]
- Das, S.P.; Mitra, S. Deep Ensembling with Multimodal Image Fusion for Efficient Classification of Lung Cancer. In International Conference on Computing, Communication and Networking Technologies (ICCCNT); IEEE: New York, NY, USA, 2024; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
- Fu, X.; Meng, X.; Zhou, J.; Ji, Y. High-Risk Factor Prediction in Lung Cancer Using Thin CT Scans: An Attention-Enhanced Graph Convolutional Network Approach. arXiv 2023, arXiv:2308.14000. [Google Scholar] [CrossRef] [Scilit]
- Priya, B.U.; Reddy, V.L. Hierlungxai: A Hierarchical and Explainable Deep Learning Framework for CT-Based Lung Cancer Classification. Asian Pac. J. Cancer Biol. 2026, 11, 649–667. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Chen, Y.; Wang, Y.; Ye, Y.; Sun, M.; Ren, H.; Cheng, W.; Zhang, H. Interpretable Pulmonary Disease Diagnosis with Graph Neural Network and Counterfactual Explanations. In IEEE International Conference on Systems, Man, and Cybernetics; IEEE: New York, NY, USA, 2023. [Google Scholar] [CrossRef] [Scilit]
- Sukumal, B.; Aueawatthanaphisut, A. Dual-Modal Lung Cancer AI: Interpretable Radiology and Microscopy with Clinical Risk Integration. arXiv 2026, arXiv:2604.16104. [Google Scholar] [CrossRef] [Scilit]
- Tong, G.; Xue, Z.; Dang, T.; Cao, T. Lung Nodule Malignancy Classification on 3D CT Images Using a Cosine Similarity-Enhanced Graph Attention Network. In Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 14–18 July 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Al-Shabi, M.; Lee, H.K.; Tan, M. Gated-Dilated Networks for Lung Nodule Classification in CT Scans. IEEE Access 2019, 7, 178827–178838. [Google Scholar] [CrossRef] [Scilit]
- Shivwanshi, R.R.; Nirala, N.S. A Hybrid AI Method for Lung Cancer Classification Using Explainable AI Techniques. Phys. Medica 2025, 134, 104985. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ying, Z.; Bourgeois, D.; You, J.; Zitnik, M.; Leskovec, J. GNNExplainer: Generating Explanations for Graph Neural Networks. In Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, USA, 8–14 December 2019. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019; Volume 32, pp. 8024–8035. [Google Scholar]
- Harris, C.R.; Millman, K.J.; van der Walt, S.J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N.J.; et al. Array Programming with NumPy. Nature 2020, 585, 357–362. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cardoso, M.J.; Li, W.; Brown, R.; Ma, N.; Kerfoot, E.; Wang, Y.; Murrey, B.; Myronenko, A.; Zhao, C.; Yang, D.; et al. MONAI: An Open-Source Framework for Deep Learning in Healthcare. arXiv 2022, arXiv:2211.02701. [Google Scholar]
- Gu, Y.; Chi, J.; Liu, J.; Yang, L.; Zhang, B.; Yu, D.; Zhao, Y.; Lu, X. A Survey of Computer-Aided Diagnosis of Lung Nodules from CT Scans Using Deep Learning. Comput. Biol. Med. 2021, 137, 104806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chang, C.; Zhao, Q.; Zhao, L.; Yang, X. Explainable AI for Lung Nodule Detection and Classification in CT Images. In SPIE Medical Imaging Conference; SPIE: Bellingham, WA, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31st Conference on Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph Attention Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Zhu, W.; Liu, C.; Fan, W.; Xie, X. DeepLung: 3D Deep Convolutional Nets for Automated Pulmonary Nodule Detection and Classification. In IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2018; pp. 673–681. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; Paluri, M. A Closer Look at Spatiotemporal Convolutions for Action Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2018. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 618–626. [Google Scholar]
- Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Rajkumar, K.V.; Kachapuram, B.R.; Sudhakar, G.; Bhukya, S.; Thota, P.; Bharat Siva Varma, P. Multi-Scale Attention-Driven Deep Learning Framework for Lung Cancer Classification from CT Images. Discov. Comput. 2026, 29, 371. [Google Scholar] [CrossRef] [Scilit]
- Lin, C.; Jiang, H.; Ma, S.; Tang, J.; Ning, Y.; Jin, L.; He, W.; Bai, J.; Xiong, Z.; Zhu, B.; et al. Multi-Center Validated Attention-BiFPN Deep Learning for CT-Based Lung Cancer Subtype Classification. Eur. J. Med. Res. 2026. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD); ACM: New York, NY, USA, 2016; pp. 785–794. [Google Scholar]
- Zhao, D.; Xi, J.; Guo, X.; Chai, J.; Xu, Z.; Li, L.; Xue, Y.; Sun, Q.; Zheng, Y.; Liu, S. Graphicalized Vision-Language Modeling for Comprehensive Lung Nodule Analysis and Risk Stratification. npj Digit. Med. 2026, 9, 442. [Google Scholar] [CrossRef] [Scilit] [PubMed]
















| Reference | Dataset | Sample Size | Reference Standard | Model Type | Validation Design | Explainability Method | Performance |
|---|---|---|---|---|---|---|---|
| Aamir et al. (2024) [12] | Public CT lung cancer dataset | 4500 CT images/cases | Lung cancer/tissue classification labels | DenseNet + CNN-based multimodal data fusion with Mobile Edge Computing | Reported in the original study | None | Accuracy 99.3%; Precision 99.3%; Recall 99.3%; F1-score 99.3% |
| Pushpa et al. (2023) [18] | Multimodal lung tumor data including CT images, clinical, and molecular information | Not reported | Not reported | CNN + Graph Neural Network for multimodal lung tumor analysis | Reported in the original study | None | Accuracy: 85% |
| Saihood et al. (2024) [13] | LIDC-IDRI; LUNGx for external evaluation | LIDC-IDRI: 1570 nodules (858 malignant, 712 benign); LUNGx: 73 nodules from 60 CT scans (37 benign, 36 malignant) | Radiologist-assigned LIDC-IDRI malignancy scores; score 3 treated as undetermined and excluded from the final binary set | Multi-side Graph Neural Network with attention-based local co-occurrence feature fusion | 10-fold cross-validation on LIDC-IDRI; external testing on unseen LUNGx | Graph/attention-based interpretability analysis | LIDC-IDRI: 87.17% Accuracy, 95.00% AUC; LUNGx: 69.86% Accuracy, 70.20% AUC |
| Tong et al. (2025) [25] | LIDC-IDRI | Not reported | LIDC-IDRI radiologist-assigned malignancy ratings | 3D CNN + cosine-similarity graph + Graph Attention Network (CSEGAT) | Patient-level classification; exact split not reported in accessible information | none | Accuracy: 90.85%. Sensitivity: 88.84%. Specificity: 90.65% |
| Shivwanshi & Nirala (2025) [27] | Public CT dataset | Not reported | Five-class lung-nodule malignancy classification | Radiomic features + InceptionNet + Vision Transformer + XGBoost feature selection + CatBoost | Not clearly reported in accessible information | SHAP | Accuracy: 96.74%; Precision: 93.68%; Recall: 96.74%; F1: 95.19%; AUC: 99.76% |
| Proposed R-GGFN (2026) | LIDC-IDRI | 867 patients; 2602 nodules | Consensus LIDC-IDRI radiologist-assigned malignancy rating; mean S ≥ 3 = malignant, S < 3 = benign | 3D ResNet-18 + MLP semantic encoder + gated multimodal fusion + GAT + tabular skip connection | Patient-level split: 70% training (606 patients, 1820 nodules), 15% validation (130 patients, 356 nodules), and 15% held-out test (131 patients, 426 nodules); test set accessed once after model selection | 3D Grad-CAM + SHAP + GAT attention visualization | Accuracy: 85.21%; AUROC: 0.9147; PR-AUC: 0.9213; F1: 0.8609 |
| Method | 3D CT | Clinical Features | Multimodal Fusion | Graph Learning | Skip Connection | Focal Loss | Multi-XAI | External Validation |
|---|---|---|---|---|---|---|---|---|
| Aamir et al. [12] | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Pushpa et al. [18] | ✓ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ |
| Saihood et al. [13] | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | Partial | ✗ |
| CSEGAT [25] | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | Partial | ✗ |
| Hybrid AI + SHAP [27] | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ |
| Proposed R-GGFN | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ |
| Method | Branch | Attachment Point | Question Answered |
|---|---|---|---|
| GradCAM-3D | Image | Layer 4 of R3D-18 forward + backward hook on the last convolutional block | Which 3D voxel regions most drove the predicted class? |
| SHAP Kernel-Explainer | Tabular | MLP backbone + tabular skip connection path, evaluated via a self-loop-only GATConv pass (no cross-nodule graph) | Which of the 8 radiologist annotation features pushed the output toward or away from malignancy? |
| Gat Attention | Graph | return_attention_weights = True, attention edge weights collected across the validation set | How much did each neighboring patient (co-patient) node contribute contextual information to the prediction? |
| Model | Accuracy | F1 | AUROC | PR-AUC | Precision | Recall | Specificity |
|---|---|---|---|---|---|---|---|
| 3D ResNet (best epoch 10) | 63.15% | 0.639 | 0.698 | 0.762 | 0.681 | 0.602 | 0.667 |
| Tabular MLP (best epoch 9) | 68.31% | 0.710 | 0.742 | 0.752 | 0.705 | 0.714 | 0.646 |
| XGBoost (best round 16) | 67.61% | 0.709 | 0.764 | 0.781 | 0.691 | 0.727 | 0.615 |
| Model | Best Epoch | Accuracy | F1 | AUROC | PR-AUC | Recall (Malignant) |
|---|---|---|---|---|---|---|
| Late Fusion | 9 | 67.61% | 0.723 | 0.743 | 0.769 | 0.779 |
| Gated Fusion | 8 | 67.84% | 0.714 | 0.746 | 0.772 | 0.740 |
| Epoch | Acc. | Precision | Recall | F1 | Sens. | Spec. | AUROC | PR-AUC |
|---|---|---|---|---|---|---|---|---|
| 1 | 0.4860 | 0.4756 | 1.0000 | 0.6447 | 1.0000 | 0.0368 | 0.6255 | 0.5791 |
| 2 | 0.4747 | 0.4703 | 1.0000 | 0.6397 | 1.0000 | 0.0158 | 0.6774 | 0.6630 |
| 3 | 0.6180 | 0.5688 | 0.7470 | 0.6458 | 0.7470 | 0.5053 | 0.6964 | 0.6719 |
| 4 | 0.6657 | 0.6374 | 0.6566 | 0.6469 | 0.6566 | 0.6737 | 0.7159 | 0.6846 |
| 5 | 0.6601 | 0.6098 | 0.7530 | 0.6739 | 0.7530 | 0.5789 | 0.7320 | 0.7038 |
| 6 | 0.6938 | 0.6541 | 0.7289 | 0.6895 | 0.7289 | 0.6632 | 0.7480 | 0.7040 |
| 7 | 0.6994 | 0.6612 | 0.7289 | 0.6934 | 0.7289 | 0.6737 | 0.7575 | 0.7215 |
| 8 | 0.6629 | 0.6027 | 0.8133 | 0.6923 | 0.8133 | 0.5316 | 0.7597 | 0.7265 |
| 9 | 0.6882 | 0.6410 | 0.7530 | 0.6925 | 0.7530 | 0.6316 | 0.7590 | 0.7148 |
| 10 | 0.6573 | 0.5909 | 0.8614 | 0.7010 | 0.8614 | 0.4789 | 0.7501 | 0.6987 |
| Epoch | Acc. | Precision | Recall | F1 | Sens. | Spec. | AUROC | PR-AUC |
|---|---|---|---|---|---|---|---|---|
| 1 | 0.5955 | 0.6833 | 0.2470 | 0.3628 | 0.2470 | 0.9000 | 0.7048 | 0.6921 |
| 2 | 0.6292 | 0.6250 | 0.5120 | 0.5629 | 0.5120 | 0.7316 | 0.7596 | 0.7483 |
| 3 | 0.6770 | 0.6759 | 0.5904 | 0.6302 | 0.5904 | 0.7526 | 0.8041 | 0.7925 |
| 4 | 0.7472 | 0.7289 | 0.7289 | 0.7289 | 0.7289 | 0.7632 | 0.8413 | 0.8338 |
| 5 | 0.7725 | 0.7485 | 0.7711 | 0.7596 | 0.7711 | 0.7737 | 0.8627 | 0.8564 |
| 6 | 0.7837 | 0.7987 | 0.7169 | 0.7556 | 0.7169 | 0.8421 | 0.8789 | 0.8721 |
| 7 | 0.8090 | 0.8182 | 0.7590 | 0.7875 | 0.7590 | 0.8526 | 0.8904 | 0.8856 |
| 8 | 0.8146 | 0.8247 | 0.7651 | 0.7938 | 0.7651 | 0.8579 | 0.8982 | 0.8946 |
| 9 | 0.8174 | 0.8435 | 0.7470 | 0.7923 | 0.7470 | 0.8789 | 0.9056 | 0.8915 |
| 10 | 0.8230 | 0.8456 | 0.7590 | 0.8000 | 0.7590 | 0.8789 | 0.9101 | 0.8884 |
| Model | Type | Acc. | AUROC | PR-AUC | F1 |
|---|---|---|---|---|---|
| 3D ResNet | Image-only | 63.15% | 0.698 | 0.762 | 0.639 |
| Late Fusion | Multimodal | 67.61% | 0.743 | 0.769 | 0.723 |
| Gated Fusion | Multimodal | 67.84% | 0.746 | 0.772 | 0.714 |
| Naïve GGFN | Multimodal + Graph | 61.97% | 0.693 | 0.701 | 0.695 |
| Tabular MLP | Tabular-only | 68.31% | 0.742 | 0.752 | 0.710 |
| XGBoost | Tabular-only | 67.61% | 0.764 | 0.781 | 0.709 |
| R-GGFN | Multimodal + Graph | 85.21% | 0.9147 | 0.9213 | 0.861 |
| Reference | Dataset | Image Modality | Multimodal Fusion | Graph Learning | Explainable AI | Key Contribution | Reported Performance | Limitation |
|---|---|---|---|---|---|---|---|---|
| Aamir et al. (2024) [12] | Public CT Dataset | CT | ✓ | ✗ | ✗ | CNN-based multimodal feature fusion with Mobile Edge Computing | Reported in original paper | No graph reasoning or interpretable predictions |
| Pushpa et al. (2023) [18] | Public Lung Cancer Dataset | CT | ✓ | (GNN) | ✗ | Early multimodal GNN framework | Reported in original paper | Limited graph modeling and no explainability |
| Saihood et al. (2024) [13] | LIDC-IDRI | CT | ✗ | (Attention GNN) | Partial (Attention) | Attention-based graph learning for lung nodules | Reported in original paper | Image-only learning without radiologist annotations |
| CSEGAT (2025) [25] | LIDC-IDRI | 3D CT | ✗ | (Cosine Similarity GAT) | Partial (Attention) | Graph Attention with cosine similarity graph construction | Reported in original paper | No multimodal fusion or comprehensive XAI |
| Hybrid AI + SHAP (2025) [27] | Public CT Dataset | CT + Clinical | ✓ | ✗ | (SHAP) | Explainable multimodal CNN framework | Reported in original paper | No graph-based contextual learning |
| My Proposed R-GGFN | LIDC-IDRI | 3D CT + Radiologist Semantic Features | ✓ | (Graph Attention Network) | (Grad-CAM, SHAP, GAT Attention) | Graph-guided multimodal fusion with tabular skip connection, optimization strategy, and multi-level explainability | Accuracy = 85.21%, AUROC = 0.9147, PR-AUC = 0.9213, F1 = 0.8609 | Requires external multi-center validation |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jafar, A.; Asif, R.; Jameel, S.M. Refined Graph-Guided Fusion Network for Explainable Multimodal Lung Cancer Classification Using CT Imaging and Semantic Features. Information 2026, 17, 839. https://doi.org/10.3390/info17090839
Jafar A, Asif R, Jameel SM. Refined Graph-Guided Fusion Network for Explainable Multimodal Lung Cancer Classification Using CT Imaging and Semantic Features. Information. 2026; 17(9):839. https://doi.org/10.3390/info17090839
Chicago/Turabian StyleJafar, Adiba, Raheela Asif, and Syed Muslim Jameel. 2026. "Refined Graph-Guided Fusion Network for Explainable Multimodal Lung Cancer Classification Using CT Imaging and Semantic Features" Information 17, no. 9: 839. https://doi.org/10.3390/info17090839
APA StyleJafar, A., Asif, R., & Jameel, S. M. (2026). Refined Graph-Guided Fusion Network for Explainable Multimodal Lung Cancer Classification Using CT Imaging and Semantic Features. Information, 17(9), 839. https://doi.org/10.3390/info17090839

