Open AccessArticle
Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI
by
Wenhao Guo
Wenhao Guo
and
Golrokh Mirzaei
Golrokh Mirzaei *
Department of Computer Science and Engineering, The Ohio State University, Columbus, OH 43210, USA
*
Author to whom correspondence should be addressed.
Submission received: 13 November 2025
/
Revised: 2 December 2025
/
Accepted: 3 December 2025
/
Published: 22 December 2025
Simple Summary
Glioblastoma is an aggressive brain cancer, and follow-up MRI scans are used to determine whether changes after treatment represent real tumor growth or temporary treatment effects. This decision is difficult, especially at the first follow-up. We analyzed 180 patients and compared eleven deep learning models across two follow-up timepoints. Overall accuracy was similar at both timepoints, ranging from about 70% to 74%. However, the second follow-up provided clearer separation between the three clinical outcomes, with the best model improving its F1 score from 0.44 at the first follow-up timepoint to 0.53 at the second follow-up timepoint. A model that combines convolutional features with a state-space sequence method consistently gave the best balance of accuracy and efficiency, while some transformer models reached higher AUC values but required much more computation. These findings offer a practical benchmark to guide future research and clinical tool development.
Abstract
Background: Differentiating true tumor progression (TP) from treatment-related pseudoprogression (PsP) in glioblastoma remains challenging, especially at early follow-up. Methods: We present the first timepoint-specific, cross-sectional benchmarking of deep learning models for follow-up MRI using the Burdenko GBM Progression cohort (n = 180). We analyze different post-RT scans independently to test whether architecture performance depends on timepoint. Eleven representative DL families (CNNs, LSTMs, hybrids, transformers, and selective state-space models) were trained under a unified, QC-driven pipeline with patient-level cross-validation. Across both timepoints, accuracies were comparable (~0.70–0.74), but discrimination improved at the second follow-up, with F1 and AUC increasing for several models, indicating richer separability later in the care pathway. Results: A Mamba+CNN hybrid consistently offered the best accuracy–efficiency trade-off, while transformer variants delivered competitive AUCs at substantially higher computational cost, and lightweight CNNs were efficient but less reliable. Performance also showed sensitivity to batch size, underscoring the need for standardized training protocols. Notably, absolute discrimination remained modest overall, reflecting the intrinsic difficulty of TP vs. PsP and the dataset’s size and imbalance. Conclusions: These results establish a timepoint-aware benchmark and motivate future work incorporating longitudinal modeling, multi-sequence MRI, and larger multi-center cohorts.
Share and Cite
MDPI and ACS Style
Guo, W.; Mirzaei, G.
Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI. Cancers 2026, 18, 36.
https://doi.org/10.3390/cancers18010036
AMA Style
Guo W, Mirzaei G.
Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI. Cancers. 2026; 18(1):36.
https://doi.org/10.3390/cancers18010036
Chicago/Turabian Style
Guo, Wenhao, and Golrokh Mirzaei.
2026. "Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI" Cancers 18, no. 1: 36.
https://doi.org/10.3390/cancers18010036
APA Style
Guo, W., & Mirzaei, G.
(2026). Timepoint-Specific Benchmarking of Deep Learning Models for Glioblastoma Follow-Up MRI. Cancers, 18(1), 36.
https://doi.org/10.3390/cancers18010036
Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details
here.
Article Metrics
Article Access Statistics
For more information on the journal statistics, click
here.
Multiple requests from the same IP address are counted as one view.