3.1. Benchmark Evaluation
Table 2,
Table 3 and
Table 4 present the quantitative performance comparison between the proposed
framework and several state-of-the-art LDCT denoising methods [
24,
38,
39] on the abdomen, head, and chest datasets, respectively. To provide a more rigorous assessment of model generalization, all experiments were conducted using patient-level data partitioning, ensuring that all CT slices from a given patient were assigned exclusively to either the training, validation, or testing subset. The reported values represent the mean performance with the corresponding 95% confidence intervals computed over the independent test images. Across all anatomical regions, the proposed
consistently achieved the highest SSIM and PSNR values while producing the lowest RMSE among the evaluated methods.
To further assess the robustness of the observed improvements and account for intra-patient correlation across slices, statistical significance testing was performed using a two-sided paired Wilcoxon signed-rank test on patient-level aggregated measurements (
independent patient cases for each anatomical dataset). For each patient volume, quantitative metrics were first averaged across constituent slices, and paired tests were conducted across independent subjects. The resulting
p-values are reported in
Table 5,
Table 6 and
Table 7. The proposed
demonstrated statistically significant improvements over all comparison methods across the evaluated anatomical regions (
), indicating that the observed performance gains were consistent across independent patient samples.
For the abdomen dataset, achieved the highest reconstruction quality with an SSIM of 0.9020, a PSNR of 42.40 dB, and the lowest RMSE of 0.0076. Compared with SADiff and GEDFormer, the proposed framework improved structural similarity while simultaneously reducing reconstruction error, demonstrating more accurate attenuation recovery under the patient-level evaluation protocol.
A similar trend was observed for the head dataset, where achieved the highest SSIM (0.7180), PSNR (42.25 dB), and the lowest RMSE (0.0078). Despite the increased anatomical complexity and the presence of fine, high-frequency cranial structures, the proposed framework preserved structural information more effectively than the competing approaches.
The chest dataset remained the most challenging evaluation scenario because of the large attenuation differences between lung parenchyma, vasculature, and surrounding soft tissues. Nevertheless, achieved the best overall performance, with an SSIM of 0.7980, a PSNR of 36.10 dB, and an RMSE of 0.0152. These results demonstrate that the proposed framework effectively suppresses projection-domain noise while maintaining structural consistency following image reconstruction.
Figure 4 presents qualitative comparisons of the benchmark denoising methods for representative abdomen, head, and chest CT slices. Although all methods reduced the noise present in the LDCT images, noticeable differences were observed in structural preservation and attenuation fidelity. DRL exhibited residual noise and localized texture distortions, while SADiff produced smoother reconstructions at the expense of subtle anatomical details. GEDFormer preserved global structures more effectively but introduced slight over-smoothing around tissue boundaries. In contrast, the proposed
framework generated images that most closely resembled the NDCT reference while preserving both low-contrast soft-tissue structures and fine anatomical details. These qualitative observations are consistent with the quantitative results reported in
Table 5,
Table 6 and
Table 7.
To further evaluate attenuation preservation, intensity profiles were extracted along the reference lines shown in the CT slices.
Figure 5 illustrates the corresponding intensity distributions for the abdomen, head, and chest datasets. The profile lines were selected to traverse representative anatomical structures and tissue boundaries within each CT slice. Because the anatomical characteristics of the three datasets differ substantially, separate profile locations were selected to capture clinically relevant structural transitions while maintaining the same profile location across all denoising methods for each corresponding image. The LDCT profiles exhibit substantial fluctuations caused by quantum noise, resulting in noticeable deviations from the NDCT reference. Although DRL, SADiff, and GEDFormer reduce these fluctuations, discrepancies remain around high-gradient anatomical transitions. In comparison, the proposed
framework produces intensity profiles that most closely follow the NDCT reference throughout the sampling path, indicating improved attenuation recovery and preservation of tissue boundaries. In the abdomen dataset,
Figure 5a,b, the proposed framework accurately reproduces attenuation transitions at liver–soft tissue interfaces and vascular structures. For the head dataset,
Figure 5e,f,
better preserves the rapid intensity transitions associated with cortical bone and intracranial anatomy. Similarly, for the chest dataset,
Figure 5c,d, the proposed framework maintains profile consistency across the strong attenuation transitions between lung parenchyma, vessels, and surrounding soft tissues while effectively suppressing noise-induced oscillations.
The improved performance can be attributed to three complementary design components. First, the projection-domain sequential state-space modeling captures long-range dependencies between neighboring projections more effectively than conventional convolutional processing. Second, the photon-aware transformer adaptively emphasizes reliable projection measurements according to the estimated photon statistics, enabling more robust feature extraction under low-dose acquisition conditions. Finally, the reconstruction-consistency optimization jointly constrains the projection and image domains through differentiable filtered backprojection, improving attenuation fidelity and preserving anatomical structures in the reconstructed CT images.
3.2. Ablation Analysis
To evaluate the contribution of each component within the proposed framework, an ablation study was conducted using the patient-level data partitioning protocol. As summarized in
Table 8, the proposed architecture was constructed incrementally, beginning with a baseline Vision Transformer (ViT) and progressively incorporating the projection-domain sequential state-space module, the photon-aware transformer, and the reconstruction consistency loss.
The baseline ViT achieved SSIM values of 0.8580, 0.6420, and 0.7210 for the abdomen, head, and chest datasets, respectively. Introducing the projection-domain sequential state-space module (P, without the photon-aware transformer) consistently improved all quantitative metrics across the three anatomical regions. For example, the abdomen SSIM increased from 0.8580 to 0.8760, while the corresponding RMSE decreased from 0.0112 to 0.0098. Similar improvements were observed for the head and chest datasets, indicating that sequential state-space modeling effectively captures long-range dependencies and angular continuity within the projection domain.
Replacing the state-space module with the photon-aware transformer (PT, without ) produced additional performance improvements over the baseline network. Compared with the P configuration, the photon-aware transformer achieved higher SSIM and PSNR values across all datasets, demonstrating that adaptively weighting projection features according to estimated photon statistics improves robustness under low-dose acquisition conditions.
When both components were combined within the proposed PS3T framework and optimized using only the conventional mean squared error objective, further improvements were observed across all evaluation metrics. These results suggest that the sequential state-space module and photon-aware transformer provide complementary benefits by jointly modeling projection-domain dependencies while enhancing feature representation according to measurement reliability.
The final configuration incorporated the proposed reconstruction consistency loss, , resulting in the highest quantitative performance across all anatomical regions. Compared with the MSE-only configuration, the proposed optimization improved SSIM from 0.8980 to 0.9020, PSNR from 41.80 dB to 42.40 dB, and reduced RMSE from 0.0080 to 0.0076 for the abdomen dataset. Similar improvements were consistently observed for the head and chest datasets. These findings indicate that explicitly enforcing consistency between the denoised projection data and the reconstructed CT images improves attenuation fidelity and anatomical preservation beyond optimization performed solely in the projection domain.
The contribution of each architectural component is also evident in the qualitative comparisons shown in
Figure 6. The baseline ViT exhibits noticeable texture smoothing and loss of fine structural detail. Incorporating the sequential state-space module improves edge continuity and structural consistency, while the photon-aware transformer further suppresses residual noise by adaptively emphasizing reliable projection measurements. The complete
PS3T framework, together with the proposed reconstruction consistency loss, produces images with the sharpest anatomical boundaries and tissue contrast that most closely resemble the NDCT reference.
These observations are further supported by the corresponding attenuation profiles shown in
Figure 7. The baseline ViT exhibits larger deviations from the NDCT reference, particularly around regions containing rapid attenuation transitions. The addition of either the sequential state-space module or the photon-aware transformer improves profile alignment by reducing noise-induced fluctuations while preserving structural transitions. The complete
PS3T framework demonstrates the closest agreement with the NDCT reference across all anatomical regions, illustrating the complementary contributions of projection-domain sequential state-space modeling, photon-aware attention, and reconstruction-consistency optimization.
3.3. Reconstruction Consistency
The reconstruction consistency trends reported in
Table 9 and
Table 10 are consistent with the intensity profile analyses presented in
Figure 5 and
Figure 7. Models achieving lower reconstruction consistency losses produce attenuation profiles that more closely match the NDCT reference, indicating improved preservation of tissue intensity distributions after reconstruction. In particular, the proposed
framework achieves the lowest
values throughout the entire training process, demonstrating superior reconstruction fidelity compared with both benchmark methods and ablated configurations.
For the benchmark comparison in
Table 9, all methods exhibit a reduction in reconstruction consistency loss as training progresses, indicating improved alignment between the reconstructed images and the reference NDCT images. However,
demonstrates the fastest convergence, as shown in
Figure 8a, and maintains the lowest loss throughout training. At the beginning of training,
achieves an initial
value of 0.1750, already lower than DRL (0.2180), SADiff (0.1920), and GEDFormer (0.1870). This performance gap becomes increasingly evident as optimization progresses. By epoch 50,
reaches a final reconstruction consistency loss of 0.0128, outperforming GEDFormer (0.0190), SADiff (0.0225), and DRL (0.0460). These results demonstrate that the proposed framework learns reconstruction-consistent representations more effectively than existing approaches.
The improved convergence behavior of can be attributed to its physics-guided projection-domain learning strategy. Unlike conventional image-domain denoising approaches that primarily optimize image similarity, incorporates projection-domain representations with photon-aware attention and sequential state space evolution, enabling the network to preserve acquisition-related characteristics while reducing noise. As a result, the learned features remain more consistent with the underlying CT reconstruction process, leading to lower reconstruction discrepancies throughout training.
The ablation study in
Table 10, illustrated in
Figure 8b, further highlights the contribution of each proposed component. The baseline ViT exhibits the highest reconstruction consistency loss throughout training, reaching 0.0710 at epoch 50. Introducing projection-domain processing through PT (no
) substantially reduces the reconstruction error to 0.0320, demonstrating the benefit of incorporating projection information. Similarly, the
(no T) configuration achieves a lower final loss of 0.0275, indicating that sequential state space modeling improves reconstruction consistency. However, neither component alone achieves the performance of the complete
framework.
The full model achieves the lowest loss at every evaluated epoch, decreasing from 0.2110 at epoch 1 to 0.0128 at epoch 50. Compared with the baseline ViT, this represents an approximately 82% reduction in reconstruction consistency error, demonstrating the complementary effect of combining projection-domain learning, sequential state space evolution, and transformer-based refinement. The progressively increasing separation between and its ablated variants during later training stages suggests that the proposed components provide a stronger reconstruction-aware learning signal and improve optimization stability.
Hence, the reconstruction consistency analysis confirms that provides superior convergence behavior and reconstruction fidelity compared with existing benchmark methods and individual component variations. The consistently lower values indicate that the proposed physics-guided framework effectively constrains the denoising process toward physically plausible solutions, resulting in improved attenuation preservation and anatomical consistency after reconstruction.
3.4. Physical and Structural Advantages
The improved performance of can be attributed to its architectural design, which integrates physics-guided constraints within a sequential state-space transformer framework. Unlike conventional deep learning approaches that rely primarily on data-driven feature extraction, the proposed framework incorporates knowledge of the CT acquisition process into the denoising pipeline. This enables the network to exploit both the statistical characteristics of the measured projections and the underlying physics governing X-ray attenuation.
Operating directly in the projection domain (sinogram space), models sequential relationships between neighboring projection views using a state-space formulation. This facilitates the learning of long-range angular dependencies that are inherently present during CT acquisition. Furthermore, the photon-aware transformer adaptively emphasizes projection measurements according to their estimated photon statistics, allowing the network to place greater emphasis on more reliable measurements while reducing the influence of noise-contaminated projections. By performing denoising before image reconstruction, the proposed framework suppresses projection-domain noise prior to filtered backprojection, thereby reducing the propagation of noise-related artifacts into the reconstructed CT images.
The transformer component further enhances feature representation by jointly modeling local contextual information and long-range dependencies within the projection data. Compared with conventional self-attention mechanisms, the sequential state-space formulation provides a computationally efficient alternative for modeling long projection sequences while avoiding the quadratic complexity associated with standard transformer architectures. In combination with the proposed reconstruction consistency optimization, these components encourage agreement between the denoised projection data and the reconstructed image domain, resulting in improved attenuation preservation, structural fidelity, and quantitative reconstruction performance across all evaluated anatomical regions.
3.5. Clinical Implications and Future Work
The improvements demonstrated by indicate its potential for enhancing low-dose CT image quality. The observed gains in SSIM and PSNR, particularly in anatomically challenging regions such as the Head and Chest where bone structures and respiratory motion may complicate image reconstruction, suggest that the proposed framework could contribute to future dose reduction strategies while maintaining image quality. Furthermore, the patient-level evaluation performed in this study provides a more rigorous assessment of model generalization by ensuring complete separation of patient data between the training, validation, and testing subsets. Nevertheless, additional clinical validation remains necessary before establishing diagnostic equivalence.
While the proposed framework demonstrates consistent improvements in quantitative image quality metrics, visual assessments, and attenuation profile preservation, these evaluations do not directly measure diagnostic performance. Recent recommendations from the medical imaging and medical physics communities have emphasized that conventional metrics such as SSIM, PSNR, and RMSE may not fully characterize the clinical utility of deep learning-based CT reconstruction and denoising methods, particularly with respect to low-contrast lesion detectability and the preservation of clinically relevant anatomical structures [
40]. Although confidence intervals and paired statistical significance testing were incorporated to assess the robustness of the reported quantitative improvements, these analyses remain limited to image-quality metrics and do not directly evaluate diagnostic effectiveness. Future studies will therefore incorporate task-based image quality assessment using observer-performance metrics, such as the detectability index (
), together with blinded reader studies involving radiologists to further evaluate the diagnostic reliability and clinical applicability of the proposed framework.
Although the proposed framework demonstrated consistent performance across the evaluated anatomical regions under a patient-level evaluation protocol, the study was conducted using a single publicly available benchmark dataset. Future work will therefore investigate the generalizability of using additional multi-institutional and multi-vendor clinical datasets acquired under diverse imaging protocols and dose levels.
Overall, demonstrates promising potential for projection-domain LDCT denoising by integrating acquisition physics with sequential state-space modeling to improve image quality while preserving anatomical fidelity. Further validation through larger multicenter datasets, task-based image quality assessment, and prospective reader studies will be important to establish its clinical utility and facilitate translation into routine clinical practice.