Abstract
Accurate segmentation of ischemic stroke lesions from multimodal magnetic resonance imaging (MRI) is fundamental for quantitative assessment, treatment planning, and outcome prediction; yet, it remains challenging due to highly heterogeneous lesion morphology, low lesion–background contrast, and substantial variability across scanners and protocols. This work introduces Tri-UNetX-2D, a large-kernel and scale-aware 2D convolutional network with explicit boundary refinement for automated ischemic stroke lesion segmentation from DWI, ADC, and FLAIR MRI. The architecture is built on a compact U-shaped encoder–decoder backbone and integrates three key components: first, a Large-Kernel Inception (LKI) module that employs factorized depthwise separable convolutions and dilation to emulate very large receptive fields, enabling efficient long-range context modeling; second, a Scale-Aware Fusion (SAF) unit that learns adaptive weights to fuse encoder and decoder features, dynamically balancing coarse semantic context and fine structural detail; and third, a Boundary Refinement Head (BRH) that provides explicit contour supervision to sharpen lesion borders and reduce boundary error. Squeeze-and-Excitation (SE) attention is embedded within LKI and decoder stages to recalibrate channel responses and emphasize modality-relevant cues, such as DWI-dominant acute core and FLAIR-dominant subacute changes. On the ISLES 2022 multi-center benchmark, Tri-UNetX-2D improves Dice Similarity Coefficient from 0.78 to 0.86, reduces the 95th-percentile Hausdorff distance from 12.4 mm to 8.3 mm, and increases the lesion-wise F1-score from 0.71 to 0.81 compared with a plain 2D U-Net trained under identical conditions. These results demonstrate that the proposed framework achieves competitive performance with substantially lower complexity than typical 3D or ensemble-based models, highlighting its potential for scalable, clinically deployable stroke lesion segmentation.
1. Introduction
Ischemic stroke is one of the leading causes of mortality and long-term disability worldwide, accounting for approximately 85% of all stroke cases [1,2]. Accurate delineation of ischemic lesions from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, and prognostic assessment. In clinical practice, multimodal MRI—including diffusion-weighted imaging (DWI), apparent diffusion coefficient (ADC), and fluid-attenuated inversion recovery (FLAIR)—plays a crucial role in characterizing the spatial and temporal evolution of infarction [3,4,5]. DWI and ADC are sensitive to cytotoxic edema and the acute ischemic core, while FLAIR highlights subacute vasogenic edema, together providing complementary information for lesion staging. Precise segmentation of these lesions is therefore indispensable for quantitative volumetry, perfusion-diffusion mismatch assessment, and automated decision-support systems in stroke care.
Despite substantial progress, ischemic lesion segmentation remains challenging due to several factors: high heterogeneity in lesion size and morphology, inter-patient anatomical variability, low tissue contrast at lesion borders, and intensity inconsistencies across scanners and imaging protocols [6,7]. Lesions may appear scattered, fragmented, or diffuse, especially in chronic or multi-territorial infarcts, making automated delineation difficult even for expert readers. Consequently, deep learning models must learn to integrate local texture cues with large-scale contextual information while remaining computationally efficient and generalizable.
The Ischemic Stroke Lesion Segmentation (ISLES 2022) dataset [8] provides an ideal benchmark for this task, comprising 400 multimodal MRI subjects (DWI, ADC, FLAIR) from multiple clinical centers with standardized voxel alignment and expert-annotated lesion masks. Due to its diversity in scanner types, protocols, and stroke subtypes, ISLES 2022 is widely adopted to evaluate the generalization ability of segmentation models. Recent state-of-the-art methods such as nnU-Net [9], SegResNet [10], St-RegSeg [11], TransUNet [12], UNet++ [13], Attention U-Net [14], and DeepMedic [15] have achieved competitive results on this benchmark using 3D convolutional networks. While 3D architectures capture volumetric continuity effectively, they incur substantial computational cost and often require extensive memory and ensemble averaging to achieve robust results [16,17,18]. Conversely, lightweight 2D U-Net-based models [19,20,21] are easier to train and deploy but suffer from limited receptive fields, restricting their ability to capture global spatial dependencies critical for large-territory infarcts [22].
Traditional U-Net decoders concatenate encoder features uniformly across scales, implicitly assuming equal contribution from coarse and fine representations. However, this assumption neglects the semantic gap between low-level structural features and high-level contextual maps, resulting in over-smoothed predictions and discontinuities in multifocal lesions [23,24]. Moreover, most segmentation losses such as Dice or Binary Cross-Entropy prioritize region overlap over contour accuracy, leading to blurred lesion borders and suboptimal Hausdorff scores [25,26,27]. These limitations highlight the need for an architecture that balances contextual awareness, multi-scale adaptivity, and boundary precision [27,28].
To address these challenges, this paper proposes Tri-UNetX-2D, a novel multimodal framework for ischemic stroke lesion segmentation that unifies wide-context modeling, adaptive scale fusion, and boundary-aware refinement within an efficient 2D U-shaped architecture. The name reflects its three core innovations: first, Large-Kernel Inception (LKI) that uses factorized depthwise separable convolutions to emulate very large receptive fields (k = 7 to 13), capturing long-range dependencies without increasing computational cost. Second, Scale-Aware Fusion (SAF), which is a decoder mechanism that learns adaptive scale weights between encoder and decoder features, enabling dynamic fusion across coarse and fine spatial representations. Third, the Boundary Refinement Head (BRH), which is an auxiliary output branch that explicitly supervises contour prediction to sharpen lesion boundaries and reduce Hausdorff distance. Additionally, Squeeze-and-Excitation (SE) attention modules are integrated throughout the network to adaptively recalibrate feature channels, emphasizing modality-specific cues (e.g., DWI for acute infarction and FLAIR for subacute edema). This combination allows the model to jointly capture long-range context, scale adaptivity, and anatomical boundary precision while maintaining a low memory footprint and fast inference.
- The main contributions of this work can be summarized as follows:
- A novel Large-Kernel Inception (LKI) module that enables efficient long-range context extraction within 2D CNNs.
- An adaptive Scale-Aware Fusion (SAF) mechanism that dynamically balances encoder and decoder features for multi-scale lesion reconstruction.
- A Boundary Refinement Head (BRH) that introduces explicit contour supervision to enhance boundary precision.
- Comprehensive benchmarking and ablation on ISLES 2022, showing consistent improvement over standard U-Net and recent transformer-based or 3D methods.
2. Materials and Methods
Accurate segmentation of ischemic stroke lesions from multimodal magnetic resonance imaging (MRI) remains challenging due to anatomical variability, heterogeneous lesion morphology, and the complementary yet inconsistent contrast across modalities. Traditional 2D convolutional neural networks (CNNs), such as U-Net, provide efficient pixel-wise localization but are constrained by limited receptive fields, hindering their ability to capture global contextual cues. Meanwhile, 3D CNNs better exploit volumetric structure but require significantly greater computational resources and are more prone to overfitting when annotated data are limited.
2.1. Clinical Datasets Description
The proposed Tri-UNetX-2D framework was developed and evaluated on the Ischemic Stroke Lesion Segmentation (ISLES 2022) dataset. This benchmark clinical dataset facilitates automated segmentation of ischemic stroke lesions from multimodal MRI, providing diffusion-weighted imaging (DWI), apparent diffusion coefficient (ADC), and fluid-attenuated inversion recovery (FLAIR) sequences for each subject, alongside expert-annotated ground-truth lesion masks.
The dataset comprises 400 subjects in total (250 for training, 150 for testing), collected from multiple clinical centers. This multi-institutional diversity ensures broad coverage of acquisition parameters, stroke subtypes (lacunar, cortical, and large-territory infarcts), and imaging contrasts, thereby promoting model robustness and generalization. All MRI volumes are co-registered and aligned to a common neuroanatomical space. The provided ground-truth labels represent binary infarct-core masks, delineating regions of irreversible ischemic injury determined by neuroradiological consensus.
ISLES 2022 is particularly relevant for studying acute-to-subacute ischemia. The DWI sequence is the most sensitive marker of acute ischemia, the ADC map helps discriminate true infarction from T2-shine-through artifacts, and the FLAIR sequence accentuates vasogenic edema and subacute parenchymal changes. By integrating these modalities, the dataset provides a comprehensive and physiologically meaningful representation of stroke progression. Figure 1 illustrates a representative example.
Figure 1.
An example of the ISLES 2022 dataset shows the DWI, FLAIR, ADC, and the mask.
The ISLES 2022 dataset provides an official predefined split consisting of 250 training cases and 150 testing cases, which was adopted in this study without modification to ensure direct comparability with prior work. The testing set corresponds to the official ISLES 2022 test partition and was not used during training, model selection, or ablation experiments. All robustness analyses and architectural ablations were conducted exclusively on the training set with strict patient-level separation.
2.2. Data Pre-Processing and Input Representation
A standardized preprocessing pipeline was applied to mitigate inter-scanner variability, unify spatial geometry, and enhance contrast while preserving informative tissue characteristics. First, all modalities were rigidly co-registered to the DWI reference space to ensure voxel-wise anatomical correspondence across channels. The aligned 3D volumes were resampled to 1.0 mm3 isotropic resolution using linear interpolation for image data and nearest-neighbor interpolation for binary labels. Each volume was subsequently zero-padded or cropped to a fixed in-plane resolution of 320 × 320, maintaining consistent spatial dimension across cases.
To standardize intensity profiles, z-score normalization was applied within brain masks generated via Otsu thresholding on DWI, followed by morphological refinement. Mild histogram clipping further attenuated extreme outlier intensities (e.g., motion or reconstruction artifacts) while retaining clinically meaningful signal dynamics. For modality-coupled interpretation, DWI and ADC were normalized jointly to preserve their biophysical relationship, whereas FLAIR intensities were normalized independently due to differing enhancement dynamics in subacute tissue.
Following normalization, the three modalities were stacked channel-wise to form a 3-channel 2D slice representation, enabling multimodal fusion within a 2D computational framework that avoids the substantial overhead of full 3D architectures. To enhance local tissue contrast and increase lesion visibility, especially in ADC and FLAIR, contrast-limited adaptive histogram equalization (CLAHE) was applied independently to each modality.
Data augmentation was performed online during training to improve robustness and reduce overfitting. Augmentations included random flips, ±10° in-plane rotations, elastic deformations, intensity jittering, and modality dropout with a probability of 0.1, which simulates scenarios where one modality is degraded or unavailable. To address class imbalance, slices containing infarcted tissue were oversampled. The resulting dataset consists of standardized, contrast-enhanced multimodal 2D inputs suitable for effective representation learning.
Axial slice decompositions were used to generate 2D input tensors of dimension 3 × 320 × 320. Although 3D processing could exploit through-plane context, the proposed 2D formulation remains advantageous in several respects. First, it reduces memory footprint and computational overhead, enabling larger batch sizes. Second, it mitigates overfitting in limited-data settings. Finally, it allows contextual expansion via LKI modules, which emulate global receptive field growth despite 2D processing. By presenting DWI, ADC, and FLAIR jointly, the model learns relationships between acute diffusion-restricted lesions and delayed FLAIR hyperintensity, thereby improving ischemic core delineation.
Although the proposed framework operates on axial 2D slices, evaluation was performed at the 3D volume level. Slice-wise predictions were reconstructed by stacking axial slices in their original spatial order. Dice Similarity Coefficient (DSC) and 95th-percentile Hausdorff distance (HD95) were computed using the original voxel spacing provided by the ISLES 2022 dataset.
Lesion-wise F1-score was calculated by identifying three-dimensional connected components using 26-connectivity, where a predicted lesion was considered correctly detected if it overlapped with at least one voxel of a ground-truth lesion. This evaluation protocol ensures consistent and reproducible assessment across cases while preserving volumetric lesion integrity.
2.3. Network Architecture Overview
Tri-UNetX-2D follows a U-shaped encoder–decoder design composed of three encoder stages, a bottleneck, and three symmetric decoder stages, as shown in Figure 2. Skip connections propagate fine-resolution encoder features to the decoder to support spatial recovery. Each encoder stage integrates an LKI module to extract wide contextual information, an SE unit for channel recalibration, and a downsampling operation.
Figure 2.
Architecture of the proposed Tri-UNetX-2D model.
Each decoder stage receives two inputs: the upsampled decoder map and the corresponding encoder skip feature map. These are combined using SAF to learn adaptive Scale-Aware Fusion, followed by SE attention and a refinement LKI module. The final decoder output feeds two parallel output heads: a main segmentation head producing lesion probability maps, and an auxiliary Boundary Refinement Head (BRH) for contour supervision during training.
While the proposed Large-Kernel Inception (LKI), Scale-Aware Fusion (SAF), and Boundary Refinement Head (BRH) are inspired by common design principles in medical image segmentation, their formulations differ from prior approaches in several key aspects. LKI employs factorized depthwise large kernels with progressive kernel expansion to efficiently approximate global receptive fields within a 2D framework, rather than relying on fixed large convolutions or attention mechanisms. SAF introduces input-adaptive weighting between encoder and decoder features instead of static skip concatenation, and BRH integrates explicit contour supervision directly into the decoder to guide boundary-aware feature learning during training, rather than post-processing refinement. Together, these components are jointly optimized within a lightweight multimodal 2D architecture specifically designed for heterogeneous stroke lesions.
2.3.1. Large-Kernel Inception (LKI) Module
A principal limitation of conventional CNNs is their reliance on small (3 × 3) kernels, which confine the network’s receptive field and restrict the ability to capture global context. This limitation is particularly detrimental for ischemic lesions, which may extend across large cortical or subcortical regions. The Large-Kernel Inception (LKI) module alleviates this issue by using depthwise separable and factorized convolutions that emulate very large kernels without increasing computational load. Each LKI module, shown in Figure 3, comprises two parallel branches.
Figure 3.
Large-Kernel Inception (LKI) module structure.
The first branch performs sequential depthwise convolutions of sizes (k × 1) followed by (1 × k), where (k) increases with network depth (7, 9, 11, 13), effectively modeling long horizontal and vertical dependencies. The sequential depthwise convolutions (k × 1 followed by 1 × k) emulate a full k × k receptive field while maintaining spatial resolution through same-padding. This factorized design captures long-range vertical and horizontal dependencies efficiently, avoiding the computational burden of explicit large kernels while preserving contextual integrity.
The second branch applies a 3 × 3 dilated depthwise convolution with dilation = 2, enabling the extraction of mid-range texture and local structural details. The two branches are concatenated and fused by a 1 × 1 pointwise convolution to integrate contextual and local cues. This is followed by an SE block that recalibrates the importance of feature channels and a residual shortcut to preserve low-level details. The SE inside the LKI block adaptively reweights channel responses after feature fusion, before residual addition.
By combining large-scale context with local textural fidelity, the LKI block allows the encoder to interpret global brain regions (e.g., vascular territories and inter-hemispheric asymmetry) while still detecting fine, bulbous lesion structures. This configuration thus balances context comprehension and spatial precision in a computationally efficient manner.
2.3.2. Scale-Aware Fusion (SAF) in the Decoder
As the encoder extracts progressively abstract and coarse features, the decoder must reintegrate these with finer-resolution information to enable precise spatial localization. In standard U-Net designs, skip connections transfer encoder features to the decoder via direct concatenation, implicitly assuming that encoder and decoder feature maps contribute equally. However, this uniform treatment overlooks differences in semantic resolution, where the encoder features capture localized fine-scale structure, whereas decoder features represent broader contextual understanding.
This imbalance may limit performance, especially when ischemic lesions vary greatly in size, shape, and anatomical location. To overcome this limitation, we introduce a Scale-Aware Fusion (SAF) unit at each decoder stage to adaptively regulate the contribution of encoder and decoder feature streams. The SAF unit takes the following inputs: first, the skip-connected encoder feature map , and the upsampled decoder feature map , both with dimensions ().
As illustrated in Figure 4, both inputs are first projected into a shared feature space using 1 × 1 convolutions, ensuring channel alignment. Each projected feature map is then passed through global average pooling (GAP) to produce compact global descriptors and , capturing the global contextual significance of each scale.
Figure 4.
Schematic of the Scale-Aware Fusion (SAF) unit.
These descriptors are concatenated and fed into a lightweight fully connected network (FCN) with softmax activation that computes two normalized weights satisfying that their sum is one . These weights determine how strongly each feature stream contributes to final fusion. The fused feature representation is then computed as a weighted sum, as shown in Equation (1).
Because the fusion weights are input-adaptive, the SAF block dynamically amplifies fine spatial cues for small lesions (larger ) or emphasizes broader, contextual information for large infarcts (larger ). This allows the decoder to better handle lesion heterogeneity across subjects. The fused map () is then processed by an SE block and an LKI module for feature recalibration and large-kernel contextual refinement before proceeding deeper into the decoder.
2.3.3. Squeeze-and-Excitation (SE) Attention
In multimodal MRI, the diagnostic relevance of each modality changes depending on lesion phase and location. For this aim, the network employs the Squeeze-and-Excitation (SE) mechanism that allows the network to automatically learn which feature channels are most informative. The SE block is integrated within each LKI module, immediately after the multi-branch fusion, to enable adaptive recalibration of channel responses. This placement allows the model to selectively emphasize the most informative contextual features while maintaining local–global balance. By embedding SE internally, each LKI becomes a self-contained attention-aware context extraction unit, ensuring that hierarchical encoder representations remain anatomically meaningful and modality-adaptive.
The SE unit is shown in Figure 5; it adopts the channel-wise Squeeze-and-Excitation formulation, performing global spatial pooling followed by channel reweighting to emphasize modality-specific features. It performs global average pooling (GAP) to produce channel descriptors, with GAP across H × W producing one scalar per channel. Then, this is followed by two fully connected layers that model channel interdependencies and generate activation weights through a sigmoid function. These weights are applied to rescale the feature channels adaptively.
Figure 5.
Channel-wise Squeeze-and-Excitation (SE) block.
Integrating SE within both the encoder and the decoder allows modality-specific attention, emphasizing ADC-derived features in the acute diffusion-restricted core and FLAIR-derived features in the subacute tissue. This dynamic channel reweighting improves generalization across heterogeneous patient populations and enhances interpretability by aligning the network’s focus with clinically relevant contrasts.
In the decoder, SE attention is applied both before and within the refinement LKI module, serving distinct but complementary roles. The pre-LKI SE, positioned immediately after Scale-Aware Fusion, performs global channel rebalancing across fused multi-scale features from encoder skip links and upsampled maps. The internal SE within LKI operates at a finer granularity, recalibrating the channel responses of the spatially transformed features. This hierarchical attention strategy ensures both inter-scale and intra-feature adaptivity, enhancing reconstruction fidelity and boundary precision in the final segmentation output.
2.3.4. Boundary Refinement Head
Accurate delineation of ischemic stroke lesion boundaries remains one of the most challenging aspects in automated segmentation. Although the primary encoder–decoder pathway in Tri-UNetX-2D is trained using region-based objectives (e.g., Dice, BCE), such objectives emphasize volumetric overlap rather than contour accuracy. As a result, predictions often exhibit smoothed or imprecise borders, especially along thin or irregular lesion margins. These inaccuracies undermine clinical utility, as subtle boundary variations can influence lesion volumetry, ischemic core estimation, and downstream prognostic analysis.
To mitigate this limitation, Tri-UNetX-2D integrates a dedicated Boundary Refinement Head (BRH), designed to explicitly enhance spatial precision by forcing the network to model lesion contours. The BRH is attached at the end of the decoder and operates in parallel with the main segmentation output. Specifically, the decoder feature map is passed through a lightweight refinement branch composed of a 3 × 3 convolution, which captures local edge context, followed by a 1 × 1 convolution that compresses multi-channel information into a single boundary likelihood representation. A sigmoid activation then converts this representation into a probability map that highlights candidate contour pixels.
During training, this boundary prediction branch is supervised using a contour-specific loss function. In our implementation, contours are derived by applying morphological edge detection to ground-truth masks, and the BRH output is compared against these contours using a combination of Binary Cross-Entropy and L1 distance losses. The final training objective is a weighted sum of the main segmentation loss ( and the boundary loss (, controlled by a scalar weight ( = 0.2), as explained in the equation below:
The auxiliary loss encourages the decoder to preserve sharp transitions at lesion interfaces, strengthening the model’s sensitivity to narrow, fragmented stroke regions that are often difficult to capture by region-based losses alone. Consequently, the BRH contributes to reductions in boundary-error metrics such as Hausdorff distance (HD95) and enhances local geometric fidelity without compromising global volume overlap.
Although the boundary head is only used during training, its influence persists at inference time. By guiding intermediate representations toward edge-aware feature encoding, it produces segmentation maps that exhibit improved structural consistency and sharper delineation of stroke tissue, particularly along irregular cortical or subcortical boundaries. Ultimately, this design improves both visual quality and quantitative reproducibility, supporting more reliable lesion characterization in downstream clinical or research pipelines.
2.4. Encoder–Decoder Processing Pipeline
To clarify how the encoder and decoder operate within the Tri-UNetX-2D framework shown in Figure 1, this section details the flow of information through each stage, describing how local and global features are progressively extracted, compressed, and reconstructed.
The encoder consists of three downsampling stages and one bottleneck layer. Each stage applies an LKI block followed by SE attention and a downsampling operation. Kernel sizes expand with depth (k = 7, 9, 11, 13) to capture increasing contextual range. Stage 1 extracts low-level features such as intensity gradients and local tissue contrast that preserve the detailed textures from DWI, ADC, and FLAIR. Second, stage 2 captures mid-level patterns representing tissue organization and hemispheric symmetry, while SE attention emphasizes modality-specific cues. Stage 3 encodes higher-level semantics, such as vascular-territory structure and infarct distribution. Finally, the bottleneck integrates full-slice contextual understanding via a large kernel (k = 13), capturing global infarct topology and inter-hemispheric relationships.
The decoder path mirrors the encoder, performing three upsampling stages. Each stage receives two inputs—the upsampled decoder feature and the corresponding encoder skip feature, which are adaptively fused via Scale-Aware Fusion (SAF). The fused map is refined using SE attention (for modality weighting) and LKI modules (for spatial context restoration).
Decoder stage 1 (SAF-1, k = 11) merges bottleneck output with encoder stage 3 features to recover large-scale infarct structure. After that, decoder stage 2 (SAF-2, k = 9) integrates intermediate features for cortical and subcortical refinement. Then, decoder stage 3 (SAF-3, k = 7) fuses fine-resolution details from encoder stage 1, restoring sharp lesion borders. Two output heads finalize the reconstruction: a main segmentation head generating the lesion probability map and a Boundary Refinement Head producing contour maps to enforce crisp anatomical delineation.
2.5. Implementation and Training Details
Tri-UNetX-2D follows a three-stage encoder–decoder design. The encoder channel widths are {32, 64, 128}, with a bottleneck width of 256 channels, and the decoder mirrors this configuration symmetrically. In the Scale-Aware Fusion (SAF) unit, encoder and decoder features are first projected using 1 × 1 convolutions, followed by global average pooling and a lightweight fully connected network with a hidden dimension of 32 and softmax activation to generate adaptive fusion weights.
Squeeze-and-Excitation (SE) blocks employ a channel reduction ratio of 16. The segmentation loss consists of a weighted combination of Dice loss and Binary Cross-Entropy loss. The boundary loss weight was set to based on empirical sensitivity analysis, which provided the best trade-off between Dice overlap and boundary accuracy (HD95).
All models were trained using the Adam optimizer with an initial learning rate of , which was reduced by a factor of 0.5 upon validation performance plateau. Training was performed for 100 epochs with a batch size of 8. Experiments were conducted on a standard workstation laptop equipped with an Intel Core i7 CPU and 24 GB of RAM, without GPU acceleration. Under this configuration, the average training time per model was approximately 15 h.
The proposed Tri-UNetX-2D contains approximately 4.9 million trainable parameters. On the same CPU-only setup, the average inference time was approximately 35 ms per axial slice, corresponding to almost 4 s per 3D volume, depending on the number of slices.
2.6. Evaluation Metrics
To evaluate the segmentation performance of the proposed Tri-UNetX-2D framework on the ISLES 2022 dataset, three widely accepted quantitative metrics were employed, consistent with previous ischemic stroke segmentation studies. These metrics jointly assess spatial overlap, lesion detectability, and boundary precision between predicted and reference lesion masks.
The Dice Similarity Coefficient (DSC) was adopted as the primary measure of volumetric agreement, quantifying the spatial overlap between the predicted segmentation and the manually annotated ground truth:
where (A) and (B) denote the predicted and reference masks, respectively. A Dice value of 1 indicates perfect overlap, while 0 represents complete mismatch. High DSC values reflect accurate delineation of both the lesion’s extent and shape.
The Hausdorff distance at the 95th percentile (HD95) evaluates the boundary accuracy between the predicted and true lesion surfaces. It measures the 95th-percentile bidirectional distance between boundary voxels, providing robustness against outliers or small spurious predictions. Lower HD95 values indicate more precise and anatomically faithful boundaries—an essential property in clinical follow-up and volumetric quantification.
Finally, the lesion-wise F1-Score assesses detection sensitivity at the component level, reflecting the model’s ability to correctly identify each ischemic focus:
where (TP), (FP), and (FN) are the numbers of true-positive, false-positive, and false-negative lesions, respectively. A lesion is considered correctly detected if any voxel of the predicted component overlaps with the ground-truth lesion.
Together, these three metrics provide complementary perspectives on model performance: DSC captures spatial agreement, HD95 quantifies boundary precision, and lesion-wise F1 evaluates detection reliability across small and large infarcts. This combination ensures a balanced and clinically meaningful evaluation of segmentation accuracy for ischemic stroke lesions.
3. Results and Discussion
This section presents both quantitative and qualitative evaluations of the proposed Tri-UNetX-2D framework for ischemic stroke lesion segmentation. First, we benchmark Tri-UNetX-2D against a baseline U-Net to assess performance gains provided by the architectural components. Then, we conduct an ablation study to analyze the effect of LKI, SAF, SE, and boundary refinement on segmentation accuracy. Finally, we compare Tri-UNetX-2D with current state-of-the-art models on ISLES 2022 before discussing the clinical implications of the findings.
3.1. Baseline Comparison: U-Net vs. Tri-UNetX-2D
To establish reference performance, a standard 2D U-Net was trained and evaluated on the ISLES 2022 dataset using identical preprocessing and optimization settings. Its performance was compared with the proposed Tri-UNetX-2D, which incorporates enlarged receptive-field modeling (LKI), adaptive multi-scale fusion (SAF), and channel-wise attention (SE), alongside a boundary refinement head.
As summarized in Table 1, Tri-UNetX-2D outperformed the baseline U-Net across all evaluation metrics. The Dice score was consistently higher, demonstrating improved volumetric overlap with the ground truth. Likewise, HD95 was substantially reduced, indicating sharper and more anatomically accurate lesion contours. The lesion-wise F1 demonstrated improved detection of multifocal lesions—particularly small or low-contrast lesions that baseline U-Net frequently missed.
Table 1.
Quantitative comparison of U-Net vs. Tri-UNetX-2D.
Figure 6 shows qualitative examples from representative ISLES subjects, comparing the ground truth segmentation (a column) with plain U-Net (b column) and the proposed Tri-UNetX-2D (c column). In many cases, U-Net underestimated irregular and fragmented lesions, especially within cortical territories exhibiting subtle diffusion abnormalities. Tri-UNetX-2D produced more complete delineations and fewer false-negative islands, particularly in deep white matter and periventricular regions. These results confirm that combining large-kernel context modeling with adaptive multi-scale fusion enhances the network’s ability to distinguish infarcted tissue from normal parenchyma across diverse anatomical and contrast conditions.
Figure 6.
Qualitative comparison of stroke segmentation (red areas). Columns: (a) ground-truth masks, (b) baseline U-Net, and (c) Tri-UNetX-2D.
3.2. Ablation and Robustness Analysis
To quantify the contribution of each architectural component within Tri-UNetX-2D, we performed a stepwise ablation beginning from a baseline U-Net and progressively adding the Large-Kernel Inception (LKI), Scale-Aware Fusion (SAF), and Boundary Refinement Head (BRH). The quantitative results are summarized in Table 2.
Table 2.
Ablation analysis of Tri-UNetX-2D design components.
Introducing LKI atop the U-Net backbone yielded the largest single-component improvement in Dice (from 0.78 to 0.82), confirming that expanding the receptive field enables more accurate representation of spatially extensive ischemic territories and reduces fragmentation in the predicted masks. Adding SAF produced further gains in DSC and lesion-wise F1, particularly benefiting cases with multifocal or heterogeneous lesions. By dynamically modulating contributions from encoder and decoder features, SAF enhanced topological continuity and improved recovery of small satellite infarcts. While the improvement in overlap metrics was modest, BRH notably enhanced contour accuracy, further reducing HD95 from 8.7 mm to 8.3 mm, indicating sharper boundary localization and reduced surface deviation from the reference contours. Although the DSC and lesion-F increased modestly (+0.01 each), the primary value of BRH lies in enforcing contour precision—critical for downstream volumetric quantification and follow-up analysis.
Figure 7 qualitatively illustrates segmentation evolution across ablation configurations. The addition of LKI reduces omission errors in large cortical infarcts; inclusion of SAF strengthens structural coherence and improves delineation of smaller infarct foci; and BRH sharpens lesion boundaries, mitigating over-smoothing present in earlier variants. The complete Tri-UNetX-2D configuration yields the most anatomically faithful lesion shapes with superior spatial consistency.
Figure 7.
Qualitative ablation results of the stroke segmentation (red). From left to right: (a) ground-truth infarct mask, (b) baseline U-Net, (c) U-Net + LKI, (d) U-Net + LKI + SAF, and (e) full Tri-UNetX-2D.
In addition to architectural ablations, we evaluated the robustness of Tri-UNetX-2D to data partitioning. Five independent random 80/20 splits were generated from the ISLES 2022 training set with strict patient-level separation to prevent data leakage. Across all splits, Tri-UNetX-2D consistently outperformed the baseline 2D U-Net with low inter-split variance in the Dice Similarity Coefficient (DSC), as summarized in Table 3. Paired Wilcoxon signed-rank tests on per-case DSC values confirmed that the improvements were statistically significant (p < 0.01), demonstrating that the reported gains are not dependent on a specific data partition.
Table 3.
Robustness analysis across five random train–validation splits on the ISLES 2022 training set (mean ± std).
To assess the sensitivity of the Boundary Refinement Head (BRH), we evaluated different boundary loss weights λb ∈ {0.1, 0.2, 0.3} as shown in Table 4. The setting λb = 0.2 yielded the best balance between the Dice Similarity Coefficient and boundary accuracy (HD95). We additionally compared contour generation using morphological edge detection and Sobel-based operators and observed comparable segmentation performance, with morphological contours providing slightly more stable boundary refinement. These results indicate that the BRH is robust to reasonable variations in contour extraction and loss weighting.
Table 4.
Sensitivity analysis of the boundary loss weight λb.
3.3. Discussion
Recent studies in ischemic stroke lesion segmentation have increasingly adopted complex 3D convolutional networks and hybrid transformer–CNN architectures, motivated by their ability to model volumetric continuity and long-range dependencies. As shown in Table 5, widely used models such as SegResNet, St-RegSeg, and TransUNet achieve Dice scores in the range of 0.787–0.801, as reported in their respective ISLES 2022 publications, and were not reproduced or retrained in this study. These methods are often accompanied by strong boundary precision, reflected by HD95 values of approximately 3–4 mm. However, such architectures typically require patch-based inference, high GPU memory, and multi-stage training schemes.
Table 5.
Comparison of the proposed Tri-UNetX-2D with representative state-of-the-art methods on multimodal ischemic stroke segmentation.
In contrast, advanced 2D U-Net variants—including UNet++, Attention U-Net, and nnU-Net 2D—offer improved computational efficiency but generally underperform on heterogeneous ischemic lesions, reporting Dice scores between 0.75 and 0.78 and higher HD95 values (10–12 mm). This behavior reflects their limited receptive fields and insufficient contextual modeling.
Tri-UNetX-2D addresses these limitations by integrating large-kernel contextual modeling, adaptive multi-scale fusion, and explicit boundary supervision within a lightweight 2D framework. With a Dice score of 0.86, lesion-wise F1 of 0.81, and HD95 of 8.3 mm, the proposed method outperforms all listed 2D baselines and demonstrates performance comparable to recently reported 3D models, while requiring only a fraction of their computational resources. This outcome confirms that careful architectural design can compensate for the lack of volumetric dimensionality inherent in slice-wise 2D processing. Results reported in Table 5 for 3D and transformer-based methods are quoted from their original ISLES 2022 publications and are intended strictly for contextual comparison rather than protocol-matched evaluation, while all baseline and ablation experiments in this work were conducted under identical preprocessing, data splits, and training settings.
The experimental findings consistently demonstrate that the proposed architecture offers substantial improvements over a standard 2D U-Net and over progressively simplified variants of Tri-UNetX-2D. These results underscore the importance of combining wide-context modeling, multi-scale adaptivity, and boundary supervision when addressing the heterogeneous, dispersed, and morphologically complex nature of ischemic lesions.
First, the introduction of the Large-Kernel Inception (LKI) module contributes the largest single improvement in segmentation accuracy. Stroke lesions often span large cortical or subcortical territories and exhibit irregular or fragmented shapes that cannot be adequately captured by conventional 3 × 3 convolutions. By factorizing very large kernels (k = 7–13) into efficient depthwise separable operations, LKI significantly expands the receptive field without increasing model complexity. The improvement in Dice (0.78 → 0.82) and reduction in HD95 (12.4 → 10.6 mm) confirm that LKI enables better integration of long-range context, reduces missed peripheral regions, and enhances reconstruction of global infarct topology.
Second, the Scale-Aware Fusion (SAF) unit provides notable gains in both structural consistency and lesion detectability, particularly for scattered or small lesions. In contrast to the uniform skip concatenation of traditional U-Nets, SAF learns adaptive weights that modulate the relative contribution of encoder and decoder features. This adaptivity improves lesion continuity and enhances sensitivity to subtle ischemic changes, as reflected by the increase in lesion-wise F1 (0.75 → 0.80) and visually more coherent segmentations in periventricular and deep white matter regions.
Third, the Boundary Refinement Head (BRH) primarily enhances contour accuracy rather than volumetric overlap. While the Dice gains are modest, the reduction in HD95 from 8.7 mm to 8.3 mm highlights improved boundary sharpness and reduced deviation from expert annotations. Because small boundary errors can lead to clinically meaningful discrepancies in volumetric assessment—especially for small cortical infarcts—this improvement is particularly relevant. BRH’s edge-sensitive supervision complements region-based losses and yields a persistent beneficial effect on final predictions, even though BRH is used only during training.
An additional strength of Tri-UNetX-2D is its computational efficiency. Unlike 3D convolutional networks, which commonly require patch-based extraction, multi-GPU setups, and substantial memory overhead, Tri-UNetX-2D operates entirely in 2D with rapid inference and low resource consumption. Despite this simplicity, it surpasses the performance of many 2D CNNs and achieves accuracy competitive with recently reported transformer-based and 3D architectures. This efficiency makes Tri-UNetX-2D suitable for time-sensitive clinical workflows and deployment on resource-limited systems.
The architecture also aligns with emerging trends in medical image segmentation. Contemporary work increasingly emphasizes the importance of large receptive fields, adaptive multi-scale integration, and precise boundary refinement. Tri-UNetX-2D incorporates all three principles within a unified framework. In addition, the embedded SE modules introduce modality-adaptive channel recalibration, enabling the network to adjust its reliance on DWI, ADC, and FLAIR depending on lesion stage and scanner characteristics—an important advantage in heterogeneous clinical environments.
Despite these strengths, certain limitations remain. The slice-wise 2D formulation does not explicitly capture through-plane continuity, potentially limiting performance for lesions with complex 3D morphology. Furthermore, although preprocessing reduces scanner variability, domain shifts across unseen institutions may still impact generalization. Future work may explore lightweight 2.5D enhancements, cross-slice attention mechanisms, or domain adaptation strategies to further improve robustness in multi-center deployment.
Overall, Tri-UNetX-2D effectively balances accuracy, efficiency, and robustness, demonstrating performance competitive with complex 3D or transformer-based models while retaining a compact computational footprint. While no protocol-matched efficiency comparison with 3D models is performed, the reported parameter count and CPU inference time provide quantitative evidence supporting the lightweight and resource-efficient nature of Tri-UNetX-2D.
4. Conclusions
This work introduced Tri-UNetX-2D, a large-kernel and scale-aware convolutional framework with explicit boundary refinement for multimodal ischemic stroke lesion segmentation. By unifying three complementary innovations; Large-Kernel Inception (LKI) for expanded contextual representation, Scale-Aware Fusion (SAF) for adaptive multiscale feature integration, and a Boundary Refinement Head (BRH) for contour-focused supervision. The proposed architecture successfully addresses the challenges posed by heterogeneous lesion morphology, low tissue contrast, and multi-center variability commonly encountered in stroke MRI. Evaluation on the ISLES 2022 benchmark demonstrated substantial improvements over a standard 2D U-Net, increasing the Dice score from 0.78 to 0.86, reducing HD95 from 12.4 mm to 8.3 mm, and boosting lesion-wise F1 from 0.71 to 0.81. Ablation results confirm the distinct contributions of each component, with LKI delivering the largest gains in volumetric overlap, SAF improving sensitivity to multifocal lesions, and BRH markedly enhancing boundary accuracy. Compared with recent state-of-the-art 2D, 3D, and transformer-based approaches, Tri-UNetX-2D achieves competitive or superior performance while requiring far lower computational resources, making it suitable for real-time or resource-limited clinical settings. Although the slice-wise formulation does not explicitly encode through-plane continuity, future extensions may incorporate 2.5D reasoning, lightweight volumetric modules, cross-slice attention, and domain adaptation techniques to improve robustness across diverse scanners and institutions. Overall, these findings demonstrate that integrating wide-context convolution, adaptive multiscale fusion, and boundary-aware supervision within a streamlined 2D architecture enables accurate, efficient, and clinically meaningful ischemic stroke lesion segmentation, establishing Tri-UNetX-2D as a strong foundation for scalable deployment in real-world neuroimaging pipelines and intelligent stroke-assessment systems.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable. This study uses a publicly available dataset with fully de-identified imaging data; therefore, ethical review and approval were not required.
Informed Consent Statement
Not applicable. The study does not involve human participants or identifiable patient data.
Data Availability Statement
The data used in this study are publicly available from the ISLES 2022—Ischemic Stroke Lesion Segmentation Challenge Dataset, accessible at: https://www.isles-challenge.org/ (accessed on 15 March 2025).
Conflicts of Interest
The author declares no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| CNN | Convolutional Neural Network |
| BR | Boundary Refinement |
| CT | Computed Tomography |
| MRI | Magnetic Resonance Imaging |
| DWI | Diffusion-Weighted Imaging |
| FLAIR | Fluid-Attenuated Inversion Recovery |
| ROI | Region of Interest |
| DSC | Dice Similarity Coefficient |
| HD95 | 95th Percentile Hausdorff Distance |
| ReLU | Rectified Linear Unit |
References
- Feigin, V.L.; Stark, B.A.; Johnson, C.O.; Roth, G.A.; Bisignano, C.; Abady, G.G.; Abbasifard, M.; Abbasi-Kangevari, M.; Abd-Allah, F.; Abedi, V.; et al. Global, Regional, and National Burden of Stroke and Its Risk Factors, 1990–2019: A Systematic Analysis for the Global Burden of Disease Study 2019. Lancet Neurol. 2021, 20, 795–820. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Donnan, G.A.; Fisher, M.; Macleod, M.; Davis, S.M. Stroke. Lancet 2008, 371, 1612–1623. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chalela, J.A.; Kidwell, C.S.; Nentwich, L.M.; Luby, M.; Butman, J.A.; Demchuk, A.M.; Hill, M.D.; Patronas, N.; Latour, L.; Warach, S. Magnetic Resonance Imaging and Computed Tomography in Emergency Assessment of Patients with Suspected Acute Stroke: A Prospective Comparison. Lancet 2007, 369, 293. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Warach, S. Use of Diffusion and Perfusion Magnetic Resonance Imaging as a Tool in Acute Stroke Clinical Trials. Curr. Control. Trials Cardiovasc. Med. 2001, 2, 38–44. [Google Scholar] [CrossRef] [Scilit]
- Heit, J.J.; Zaharchuk, G.; Wintermark, M. Advanced Neuroimaging of Acute Ischemic Stroke: Penumbra and Collateral Assessment. Neuroimaging Clin. N. Am. 2018, 28, 585–597. [Google Scholar] [CrossRef] [Scilit]
- Maier, O.; Menze, B.H.; von der Gablentz, J.; Häni, L.; Heinrich, M.P.; Liebrand, M.; Winzeck, S.; Basit, A.; Bentley, P.; Chen, L.; et al. ISLES 2015—A Public Evaluation Benchmark for Ischemic Stroke Lesion Segmentation from Multispectral MRI. Med. Image Anal. 2017, 35, 250–269. [Google Scholar] [CrossRef] [Scilit]
- Alirr, O.I. Ischemic Stroke Lesion Core Segmentation from CT Perfusion Scans Using Attention ResUnet Deep Learning. J. Imaging Inform. Med. 2025, 38, 3507–3516. [Google Scholar] [CrossRef] [Scilit]
- ISLES: Ischemic Stroke Lesion Segmentation Challenge. 2022. Available online: https://isles-challenge.org/ (accessed on 17 November 2025).
- Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. NnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [Scilit]
- Mahfuzur, M.; Siddiquee, R.; Yang, D.; He, Y.; Xu, D.; Myronenko, A. Automated Ischemic Stroke Lesion Segmentation from 3D MRI. arXiv 2022, arXiv:2209.09546. [Google Scholar] [CrossRef] [Scilit]
- Gui, C.; An, X.; Li, T.; Liu, S.; Ming, D. St-RegSeg: An Unsupervised Registration-Based Framework for Multimodal Magnetic Resonance Imaging Stroke Lesion Segmentation. Quant. Imaging Med. Surg. 2024, 14, 9459–9476. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Siddiquee, M.M.R.; Tajbakhsh, N.; Liang, J. UNet++: Redesigning Skip Connections to Exploit Multiscale Features in Image Segmentation. IEEE Trans. Med. Imaging 2020, 39, 1856–1867. [Google Scholar] [CrossRef] [Scilit]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention U-Net: Learning Where to Look for the Pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar] [CrossRef] [Scilit]
- Kamnitsas, K.; Ledig, C.; Newcombe, V.F.J.; Simpson, J.P.; Kane, A.D.; Menon, D.K.; Rueckert, D.; Glocker, B. Efficient Multi-Scale 3D CNN with Fully Connected CRF for Accurate Brain Lesion Segmentation. Med. Image Anal. 2017, 36, 61–78. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Myronenko, A. 3D MRI Brain Tumor Segmentation Using Autoencoder Regularization. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); 11384 LNCS; Springer: Cham, Switzerland, 2019; pp. 311–320. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); Springer: Cham, Switzerland, 2015; Volume 9351, pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
- Alirr, O.; Alshatti, R.; Altemeemi, S.; Alsaad, S.; Alshatti, A. Automatic Brain Tumor Segmentation from MRI Scans Using U-Net Deep Learning. In Proceedings of the 2023 5th International Conference on Bio-Engineering for Smart Technologies (BioSMART 2023), Paris, France, 7–9 June 2023. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Chen, S.; Ji, C.; Fan, J.; Li, Y. Boundary-Aware Context Neural Network for Medical Image Segmentation. Med. Image Anal. 2022, 78, 102395. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alirr, O.I. Coronary Artery Segmentation in CTA Images: Evaluating Automated Segmentation of Coronary Arteries Using U-Net Variants and Vesselness Enhancement. Eur. J. Pure Appl. Math. 2025, 18, 6300. [Google Scholar] [CrossRef] [Scilit]
- Alirr, O.; Khalifa, T. Anatomically Guided Cascaded U-Net Ensemble for Coronary Artery Calcification Segmentation in Cardiac CT. Bioengineering 2025, 12, 1243. [Google Scholar] [CrossRef] [Scilit]
- Alirr, O.I.; Al-Absi, H.R.H.; Ashtaiwi, A.; Khalifa, T. Efficient Extraction of Coronary Artery Vessels from Computed Tomography Angiography Images Using ResUnet and Vesselness. Bioengineering 2024, 11, 759. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Wu, Y.; Wu, J.; Zhang, X.; Wang, D.; Zhu, S.; Chen, Y.; Wu, Y.; Wu, J.; Zhang, X.; et al. Redefining Contextual and Boundary Synergy: A Boundary-Guided Fusion Network for Medical Image Segmentation. Electronics 2024, 13, 4986. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Nan, Y.; Del Ser, J.; Yang, G. Large-Kernel Attention for 3D Medical Image Segmentation. Cognit. Comput. 2022, 16, 2063–2077. [Google Scholar] [CrossRef] [Scilit]
- Zhao, L.; Ma, J.; Shao, Y.; Jia, C.; Zhao, J.; Yuan, H. MM-UNet: A Multimodality Brain Tumor Segmentation Network in MRI Images. Front. Oncol. 2022, 12, 950706. [Google Scholar] [CrossRef] [Scilit]
- Garcia-Salgado, B.P.; Almaraz-Damian, J.A.; Cervantes-Chavarria, O.; Ponomaryov, V.; Reyes-Reyes, R.; Cruz-Ramos, C.; Sadovnychiy, S. Enhanced Ischemic Stroke Lesion Segmentation in MRI Using Attention U-Net with Generalized Dice Focal Loss. Appl. Sci. 2024, 14, 8183. [Google Scholar] [CrossRef] [Scilit]
- Al Irr, O.I.; Abd Rahni, A.A. Automatic Volumetric Localization of the Liver in Abdominal CT Scans Using Low Level Processing and Shape Priors. In Proceedings of the 2015 IEEE International Conference on Signal and Image Processing Applications (ICSIPA), Kuala Lumpur, Malaysia, 19–21 October 2015. [Google Scholar]
- Alirr, O.I. Dual Attention U-Net for Liver Tumor Segmentation in CT Images. Int. J. Comput. Commun. Control 2024, 19, 6226. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.








