4.4. Comparisons with the State-of-the-Art Methods
We conducted a comprehensive and holistic comparison analysis between our FDPR-DBNet and a range of current state-of-the-art methods on the DRIVE, CHASE_DB1, and STARE datasets. These methods mainly involved several representative general medical image segmentation models (including U-Net [
8], CE-Net [
15], LadderNet [
14], U-Net3+ [
16], TransUNet [
21], VM-UNet [
31], and VM-UNet-v2 [
32]), and models specifically designed for retinal vessel segmentation (containing SA-UNet [
17], MSTP-Net [
19], PA-Net [
24], Serp-Mamba [
33], and CRMA-UNet [
34]). We also embraced comparisons with two advanced frequency-domain-based segmentation methods, comprising FreqUNet [
45] and FSE-Mamba [
44]. Except for data for PA-Net, FreqUNet, MSCR-Net, SDHM
2T, and CRMA-UNet, which were cited directly from the original papers, the results for the remaining methods were obtained experimentally using the released source code and the default parameter configurations provided in their literature. The comparison results are summarized in
Table 4,
Table 5 and
Table 6, where “–” denotes unavailable data, while boldface and underlined text indicate optimal and second-ranked results, respectively. It is evident from the data in
Table 4,
Table 5 and
Table 6 that the proposed FDPR-DBNet demonstrated superior segmentation performance across most evaluation metrics. To support the quantitative data, we also provide an intuitive visual comparison of segmentation results between our model and other cutting-edge methods in
Figure 5,
Figure 6 and
Figure 7 on these datasets.
Quantitative analysis: Initially, in terms of the sensitivity (SE) indicator, Our FDPR-DBNet established new state-of-the-art performance on the DRIVE dataset, reaching a top SE score of 87.21%, an improvement of 2.59% over the second-best-performing method, CRMA-UNet. On the CHASE_DB1 and STARE datasets, it achieved sub-optimal performance, with competitive scores of 86.35% and 88.12%, respectively, nearly approaching the scores of the top-performing SDHM
2T and PA-Net. The SE metric reflected its capability in segmenting vessel target regions, where a higher SE score directly correlated with a reduction in the incidence of missing vessel pixels. This could be further confirmed by qualitative segmentation comparison results. As shown in
Figure 5 and
Figure 6, our model identified more vascular pixels when compared with other approaches. It should be emphasized that the SE indicator did not account for the misclassification of non-vessel pixels as vessels, and other indicators were required for further performance assessment.
Concerning the specificity (SP) metric, our model achieved scores of 98.80%, 98.63%, and 99.02% on the DRIVE, CHASE_DB1, and STARE datasets, respectively. Its SP scores on the DRIVE and CHASE_DB1 datasets were 0.05% and 0.29% lower than those of CRMA-UNet, respectively. Nevertheless, our model outperformed CRMA-UNet on the other reported metrics for these datasets and across all reported metrics on the STARE dataset. Given that the single SP metric may ignore the omission of vascular pixels, it was necessary to combine it with the SE indicator for a more comprehensive insight into the model’s performance. As illustrated in
Figure 5,
Figure 6 and
Figure 7, it was again confirmed that missegmentations occurred infrequently in our FDPR-DBNet, underscoring its effectiveness in discriminating target vascular pixels from non-vessel pixels.
As a threshold-independent metric, the Area Under the Curve (AUC) provided a holistic evaluation of the model’s discriminative ability. On the DRIVE dataset, our model achieved an AUC score of 98.43%, outperforming other state-of-the-art methods like PA-Net (98.33%) and Serp-Mamba (98.32%). In contrast, it was narrowly surpassed only by PA-Net (98.66% vs. 98.75%, and 98.70% vs. 99.08%) on the CHASE_DB1 and STARE datasets, and maintained a clear advantage by 0.31% and 0.29% over Serp-Mamba. In conclusion, the AUC range of 98.43% to 98.70% achieved by our model across diverse datasets demonstrated its exceptional robustness in distinguishing vascular pixels, underscoring its state-of-the-art performance in current retinal vessel segmentation tasks.
In the case of Accuracy (ACC), our model performed best on the CHASE_DB1 and STARE benchmarks, leading the second-best-performing CRMA-UNet by approximately 0.21% and 0.18%, respectively. On the DRIVE dataset, our model scored 97.75% and obtained a competitive second-place ranking, which only lagged behind the top-performing FreqUNet by about 0.56%. However, in other indicators, FreqUNet was inferior to it by a large margin. Similarly, regarding the precision (PR) indicator, our model consistently maintained a substantial lead on the DRIVE and CHASE_DB1 datasets, with—at least—additional gains of 0.70% and 1.65%. Even on the STARE benchmark, our model still secured the second position, only slightly falling short of Serp-Mamba by 0.65%. These outcomes highlighted that our model demonstrated a superior capability to accurately distinguish vascular regions. Given that vascular pixels occupied only a small fraction of the image, the resulting class imbalance rendered ACC and PR insufficient metrics for evaluating overall performance. Consequently, the Dice coefficient was necessary to further assess the model’s segmentation performance.
As for Dice, a harmonic mean of precision and sensitivity, our model demonstrated exceptional results at 86.47% and 86.33% on the DRIVE and STARE datasets, respectively. Specifically, it significantly outperformed the competing CRMA-UNet and PA-Net models by 2.54% on the DRIVE dataset, which demonstrated its superior capability in segmenting blood vessels both accurately and completely. At the same time, it also showed a margin gain of 0.24% over the second-best Serp-Mamba on the STARE benchmark, while building a solid advantage of 0.72% over PA-Net. Conversely, on the CHASE_DB1 dataset, it ranked second with a Dice score of 85.55% and delivered a narrow drop by only 0.25% when compared to CRMA-UNet, whereas it kept an obvious advantage over other recent models like Serp-Mamba (83.73%) and PA-Net (83.08%). Compared to MSCR-Net and SDHM
2T, our model consistently maintained a highly competitive advantage, outperforming these very recent methods in Dice by 3.26% and 3.57% on the DRIVE dataset, 3.91% and 4.50% on the CHASE_DB1 dataset, as well as 3.23% and 3.06% on the STARE dataset, respectively. Of note, CRMA-UNet did not exhibit a performance superiority over our model across other assessment indicators. The consistently high Dice (ranging from 85.55% to 86.47%) substantiated that our FDPR-DBNet accurately recognized the foreground (vessels) while minimizing false positives. Our model achieved MCC scores 8.02%, 2.15%, and 2.61% higher than those of SDHM
2T on the DRIVE, CHASE_DB1 and STARE datasets, respectively. While our model’s MCC score on the CHASE_DB1 dataset was marginally lower than that of Serp-Mamba by 0.28%, our model established a clear dominance in clDice, outperforming it by approximately 1.85%. These results validated our model’s superior ability to preserve vascular connections and prevent capillary fragmentation. Beyond that, we could observe from
Figure 5,
Figure 6 and
Figure 7 that our model effectively recognized refined retinal blood vessels, further confirming our model’s benefits in precisely segmenting vascular pixels and attenuating irrelevant noise.
To further analyze performance trends across various metrics among different methods, we carried out visualizations on comparative data using line plots (See
Figure 8,
Figure 9,
Figure 10,
Figure 11,
Figure 12 and
Figure 13). In these plots, each data point corresponded to a specific method, with yellow markers highlighting the proposed FDPR-DBNet, which achieved top-tier scores with respect to the corresponding indicator. While the general trends indicated that numerical variances among most state-of-the-art methods were relatively narrow, our FDPR-DBNet exhibited a pronounced performance edge in ACC, SE, and AUC across most datasets. Furthermore, the model maintained a highly competitive standing in Dice, collectively sustaining the robustness and superiority of the proposed FDPR-DBNet.
By scrutinizing the experimental results for the three datasets, we found that, compared with CNN-based segmentation methods, models combining Transformers with CNNs (such as PA-Net and TransUNet) generally achieved better segmentation performance, which may be due to their ability to effectively model long-range global and local dependencies among vessel features. Moreover, Mamba-based methods like Serp-Mamba generally performed better than CNN-based and Transformer-based methods. Concretely, Mamba-based models showed consistent improvements over TransUNet (78.41%, 79.55% and 76.70%) in terms of Dice, with Serp-Mamba reaching up to 83.83%, 83.73%, and 86.09%, respectively. Other key indices also showed a similar trend on these three datasets, demonstrating a strong capacity to accurately identify complex vessel structures. These advantages of Mamba-based methods could be attributed to more effective long-distance modeling power and an efficient selective state-space layer. An interesting observation was that frequency-domain-based methods possessed a certain degree of merit over spatial domain-based ones. For example, a frequency-domain method, FSE-Mamba, offered highly competitive AUCs of 98.31%, 98.41%, and 98.55% across all three datasets, rising by around 0.30%, 0.23%, and 0.38% over the spatial-domain-based model MSTP-Net, respectively. This affirmed the benefits of integrating frequency-domain modeling to preserve and process intricate vascular structural details. Ultimately, our FDPR-DBNet, a novel frequency-domain-based method, delivered near-optimal performance gains across all three datasets and reemphasized the efficacy of integrating frequency-domain information to capture fine-grained vessel details. Compared with other recently reported frequency-domain methods like FreqUNet and FSE-Mamba, the Dice score improved by 10.65%, 8.32%, and 19.12% over FreqUNet across the three benchmarks, while increasing from 81.18% to 86.47%, 81.07% to 85.55%, and 82.61% to 86.33% compared with FSE-Mamba, respectively. Moreover, it also maintained a highly stable AUC across all datasets, ranging from 98.43% to 98.70%, whereas FreqUNet showed more volatility, dropping from a competitive 98.00% on the CHASE_DB1 dataset down to 95.00% on the STARE dataset. Meanwhile, FSE-Mamba also displayed a similar fluctuating trend. The observed performance gains originated from several synergistic modules. Theoretically, our model bypassed this via DWT, isolating fine boundary details through PACA in the high-frequency branch and macro-structures via SFCA in the low-frequency branch. Furthermore, the SARF and CFF modules enabled adaptive fusion and mutual guidance of frequency-aware region and boundary features at different scales by dual-branch interaction based on a residual self-attention mechanism, which suppressed irrelevant noisy information, learned refined boundary features, and optimized region representations. Through the CSE block, our model established an explicit mathematical trade-off between localized pixel clarity (via LDE) and long-range structural continuity (via GSE), which prevented vessel fragmentation and elevated topological continuity. The MPCR module leveraged high-order prototype alignment for dynamic feature calibration, which enabled the network to distinctively suppress non-vascular confounding artifacts and preserve structural integrity in feature representations. Beyond that, a tailored connectivity loss function that penalized fragmentation and disconnection phenomena within the predicted vascular network further guided the model to better preserve the underlying topological structure of blood vessels.
Qualitative analysis: As can be seen from the visual comparison results between our model and other advanced methods in
Figure 5,
Figure 6 and
Figure 7, our model attained superior segmentation performance in some challenging scenarios. Traditional convolutional methods such as CE-Net and SA-UNet exhibited omissions of details at the distal branches of retinal blood vessels and suffered from significant structural disconnections or broken vessel segments in low-contrast regions of the DRIVE dataset. Even UNet3+ and TransUNet occasionally introduced noise or lost fine details. Although Mamba-based methods like VM-UNet and Serp-Mamba performed well, they still showed slight over-segmentation in dense regions. On the CHASE_DB1 dataset featuring higher resolution images, while models such as MSTP-Net and FSE-Mamba exhibited strong discriminative capacity to identify blood vessels by focusing on multi-scale features, false positives in areas with low contrast still occurred. The STARE dataset often contained pathological characteristics (like exudates), which complicated the task of vessel segmentation. These advanced models like TransUNet and Serp-Mamba tended to lose structural integrity in highly curved vessel paths and failed to maintain the topological continuity of fine-grained capillaries. Even in the presence of a frequency-guided attention mechanism, FSE-Mamba exhibited a tendency to prioritize dominant global structures, but occasionally missed the most minute terminal capillaries or showed over-smoothing at the junctions where multiple vessels crossed. It still struggled to retrieve delicate high-frequency signals, which inevitably led to the undesirable omission of fine terminal vessels. In contrast, our model predicted a relatively intact and continuous blood vessel network that aligned more closely with the ground truth, especially in the peripheral regions where vessel signals are weakest. In a nutshell, the proposed FDPR-DBNet stood out as a high-performing model that bridged the gap between global structural awareness and local fine details relevant to retinal blood vessels through the mutual coordination of SFCA, PACA, SARF, CFF, LDE, GSE, and MPCR modules. Its innovative architecture and hierarchical feature extraction strategy enabled the acquisition of more nuanced and comprehensive characteristic representations to successfully overcome the discontinuity issue that appeared in Transformer and Mamba models, effectively preserving vascular integrity and continuity.
To further investigate various network performance metrics, we additionally employed the ROC and precision–recall (P-R) curves for the DRIVE, CHASE_DB1, and STARE datasets, as displayed in
Figure 14 and
Figure 15. The ROC curves provided a comprehensive assessment of discrimination by considering both positive and negative samples, whereas the precision–recall curves emphasized positive samples and were therefore well suited to imbalanced data. Upon observation, it was apparent that the ROC curves for our model across all three datasets consistently occupied the upper-left corner and enclosed the curves of other models, indicating that the model effectively distinguished vessel from non-vessel pixels while minimizing false alarms. Similarly, the P-R curves were nearly the outermost across these datasets, indicating a favorable trade-off between precision and recall. In summary, the proposed FDPR-DBNet consistently outperformed alternative methods in terms of both ROC and P-R curves. This dual-curve dominance suggested that our model was not only statistically robust but also practically effective for the high-precision requirements of retinal vessel segmentation.
4.7. Cross-Dataset Generalization Evaluations
To assess the generalizability of FDPR-DBNet, we conducted cross-dataset vessel segmentation experiments on the DRIVE, CHASE_DB1, and STARE datasets and compared our model with current state-of-the-art methods. We trained each model on one dataset and tested it on the other two, alternating the training dataset.
Table 8,
Table 9 and
Table 10 list the generalization comparison results of different methods across these three datasets. Herein, we define experimental scheme 1 as training on the STARE dataset and testing on the DRIVE and CHASE_DB1 datasets. Experimental scheme 2 is defined as training on the CHASE_DB1 dataset and testing on the DRIVE and STARE datasets. The remaining configuration, which involves training on the DRIVE dataset and testing on the STARE and CHASE_DB1 datasets, is termed experimental scheme 3. It was obvious from these tables that the proposed model consistently delivered overall optimal performance across most evaluation metrics in all experimental scenarios, validating its powerful generalizability and robustness. Among them, Serp-Mamba integrated a serpentine interwoven adaptive scan mechanism into Mamba networks to capture long-range dependency correlations in vessels. Nevertheless, it showed weaker generalization capabilities compared with our designed FDPR-DBNet. In Experimental Scheme 1, the ACC of Serp-Mamba exhibited a negligible improvement (0.03%) compared to ours on the DRIVE dataset, while its SP value rose by roughly 1.24% on the CHASE_DB1 dataset. However, our model achieved substantial improvements of 5.15%, 4.17%, 3.21%, and 0.11% in SE, Dice, PR, and AUC, respectively, on the DRIVE dataset, as well as increases of 1.54%, 2.15%, 2.72%, and 0.25% on the CHASE_DB1 benchmark. In Experimental Scheme 2, a similar trend could also be observed between our model and Serp-Mamba. Recently, FSE-Mamba introduced multi-scale axial attention and frequency-domain guided attention into the Mamba network to strengthen the global perception of vascular features and model inter-frequency relationships. However, it still showed limited cross-dataset generalization capability relative to our model.
For example, in experimental scheme 3, although FSE-Mamba obtained a higher overall ACC (97.93% vs. 96.32% for the STARE dataset, and 97.81% vs. 96.48% for the CHASE_DB1 dataset), our model maintained a notable gain in other critical segmentation indicators. The SE, Dice, and AUC of our model outweighed those of the baselines by 3.97%, 2.06%, and 1.21%, as well as 1.66%, 2.32%, and 0.54%, on these two datasets, respectively.These results showed that our model still retained high performance even under domain shifts, successfully suppressing noise while maintaining fine vessel skeleton structures on different external datasets. Upon further analysis of these tables, it could be seen that when testing on the STARE dataset after training on the DRIVE dataset, nearly all models brought a performance boost in terms of the SE indicator compared to their reverse scenario. This phenomenon could primarily be attributed to two factors: on the one hand, there existed inherent visual differences between these two datasets. The images from the DRIVE dataset were characterized by darker tones and lower contrast, whereas those from the STARE dataset were noticeably brighter with higher contrast. This clear visual property of fundus images from the STARE dataset contributed to vessel recognition. On the other hand, relatively few thin vessels were annotated in the STARE benchmark, and when the models trained on this dataset were transferred to the DRIVE dataset for testing, they could not accurately segment thin vessels.