Background/Objectives: Tetralogy of Fallot (TOF) is the most common cyanotic congenital heart defect, and cardiac computed tomography (CT) is increasingly central to its anatomical and pre-procedural assessment. Artificial-intelligence research in TOF is dominated by MRI; deep learning on cardiac CT in congenital heart disease exists but addresses multi-class diagnosis and segmentation, and the one binary TOF-versus-control CT study used slice-level validation without confounder control, and, to our knowledge, no CT study reports controlling the confounding intrinsic to a TOF-versus-control comparison. This confounding is structural: TOF is imaged predominantly in infancy, so a naive classifier can learn age, body size, and acquisition protocol rather than pathology. We develop and internally evaluate a confounder-matched, anatomy-guided deep-learning pipeline for TOF on cardiac CT.
Methods: Contrast-enhanced cardiac CT from a single scanner was de-identified and restricted to one reconstruction (FC15 kernel, 0.5 mm), then matched 1:1 on age and sex, yielding 42 TOF and 42 controls (n = 84); controls were children imaged for suspected but excluded cardiovascular disease, so scan indication, unlike age and sex, was not matched. Standardized volumes were decomposed into four fixed sub-volumes positioned to approximate the components of the diagnostic tetrad: malalignment ventricular septal defect (VSD), overriding aorta, right-ventricular outflow tract (RVOT), and right-ventricular hypertrophy (RVH). Whether each sub-volume contains its named target was audited against independent physician region-of-interest annotations. Per region, a 2.5D transfer-learning classifier (ImageNet ResNet18) and a 3D CNN (DenseNet121) were trained with leak-free patient-level five-fold cross-validation and the branches fused. Optimism was assessed by repeated cross-validation and, for model selection, by nested cross-validation with the component subset and operating point chosen inside an inner loop. Discrimination was reported with bootstrap 95% confidence intervals (CIs); AUROCs were compared by DeLong test, with Benjamini–Hochberg correction applied to a seven-member family (the four within-component comparisons, two hybrid-versus-VSD contrasts, and hybrid versus whole-heart) and other comparisons reported uncorrected.
Results: Matching removed the age difference (median 0.33 years, IQR 0.17–0.92 vs. 0.33, IQR 0.27–0.73;
p = 0.86) with balanced sex (
p = 1.00). The pre-specified four-component hybrid reached AUROC 0.829 (95% CI 0.74–0.91); the VSD region alone reached 0.828 (0.74–0.91), so the tetrad decomposition did not improve accuracy, and the containment audit shows it does not deliver the intended anatomical interpretability either. The 2.5D model exceeded the 3D CNN for every component (0.769–0.828 vs. 0.573–0.656; raw DeLong
p = 0.007–0.037, Benjamini–Hochberg q up to 0.065 under a seven-member family, the weakest comparison (RVH) not surviving correction). Repeated cross-validation gave 0.811 ± 0.026 and nested cross-validation 0.787 ± 0.029; a stronger backbone with multi-phase data, handcrafted radiomics, and a large CT foundation model did not significantly improve on the matched pipeline. Grad-CAM maps were sensitive to both model weights and labels and superior to a centred-blob null in all eight comparisons and significantly so in seven, but not consistently superior to a resolution-matched random attribution, so no localization claim is made. Calibration was imperfect (slope 0.67) and recalibration gave no net gain; at an in-sample Youden threshold sensitivity was 0.93 and specificity 0.64. Occlusion sensitivity on the whole-heart baseline model showed it relies on the physician-marked septal, aortic and right-ventricular sites 1.8–4.1 times more than distance-matched surrounding tissue, while gross morphometry alone reached 0.651–0.663. Two of the four sub-volumes did not contain their target: the RVOT prior contained the physician annotation in 48.1% of cases and, because the model samples only the central band, excluded it in 99.4%; the RVH box was offset toward the midline, containing the marked target in 43.4% of annotations. Repositioning the priors, leak-free and derived from controls only, did not change discrimination (all
p≥ 0.10), and boxes placed at random positions inside the standardized heart reached 0.765 on average against 0.796 for the published priors, a difference this cohort cannot resolve.
Conclusions: As a proof of concept, confounder-matched deep learning can recognize TOF on cardiac CT. Increasing model capacity did not significantly improve on the matched pipeline; the separate contribution of matching itself was not isolated against an unmatched comparator. Given the small, single-centre sample and the absence of external validation, these findings are hypothesis-generating and require external, multi-centre confirmation before any clinical use.
Full article