Next Article in Journal
Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research
Previous Article in Journal
BERT-Based Models for Normalization of Adverse Drug Event Expressions in Social Media to Standard Medical Terminology for Drug Safety Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings

by
Emmanuel Chibuikem Nnadozie
1,2,*,
Susana Merino-Caviedes
1,
Daniel A. de Luis-Román
3,4,5,
Marcos Martín-Fernández
1,5 and
Carlos Alberola-López
1,5
1
Image Processing Laboratory, School of Telecommunications Engineering, University of Valladolid, 47011 Valladolid, Spain
2
Mechatronics Research Group, University of Nigeria, Nsukka 410001, Nigeria
3
Endocrinology and Nutrition Department, Clinical University Hospital of Valladolid, 47003 Valladolid, Spain
4
Endocrinology and Nutrition Research Centre, School of Medicine, University of Valladolid, 47005 Valladolid, Spain
5
Valladolid Health Research Institute (IBioVALL), C. Rondilla, Sta. Teresa, 47010 Valladolid, Spain
*
Author to whom correspondence should be addressed.
Big Data Cogn. Comput. 2026, 10(5), 142; https://doi.org/10.3390/bdcc10050142
Submission received: 2 February 2026 / Revised: 13 April 2026 / Accepted: 17 April 2026 / Published: 2 May 2026

Abstract

High-performance medical segmentation models are often benchmarked on high-end GPUs. Such benchmarks do not provide useful performance insights for point-of-care low-end devices. This work, firstly, posits that to achieve improved clinical adoption of AI-powered segmentation models, especially in reduced manpower settings like rural hospitals, we need benchmarks that provide actionable insights on the degree to which high-performance models address five deployment constraints viz: resource-effectiveness for low-end computing devices, clinically acceptable accuracy, clinically compatible execution times, localization of user data, and user-based finetuning. In this work, five state-of-the-art foundation segmentation models and one target-specific model were systematically evaluated on three multi-organ medical datasets. Furthermore, the best-ranking foundation model and target-specific model were benchmarked on three low-end devices. Our findings show that lightweight foundation models provided the best performance trade-off and are easily user-fine-tuned on custom datasets. Target-specific models provide high accuracy out-of-the-box, but may require significant optimisation to deliver comparably fast execution times and user-based finetuning on low-end devices. The methods and results from this research provide actionable insights on high-performance medical segmentation models for low-end computing devices, as a necessary step towards improved adoption in resource-limited clinical settings.

1. Introduction

Medical image segmentation, the leading application of Artificial Intelligence (AI) in healthcare [1], enables the effective diagnosis, treatment, and monitoring of disease progression [2]. Ab initio, the gold standard medical segmentation approaches involved manual feature extraction based on traditional image processing techniques [3]. Traditional approaches, though offering a level of interpretability and controllability [4], are very intensive, time-consuming, and require high expertise. Furthermore, the techniques are not suitable for the analysis of large-scale datasets.
Deep learning (DL)-based medical segmentation models have become the go-to models because of their ability to learn discrete image features and produce accurate segmentation results over a range of medical segmentation tasks covering specific anatomical structures and pathological regions [5].
Powerful DL segmentation models have been proposed for medical image segmentation. Until recently, the performance of the UNet with a U-shaped architecture and skip connections surpassed most medical segmentation models [6]. Improvements have been made to the original UNet, including extending the network to 3D applications [7]; introducing attention modules to focus on target structures [8,9]; incorporating inception modules [10], residual networks [11,12], dense blocks [13,14], using them in an ensemble network [15], and so forth. UNet and its variants, based on convolutional neural networks (CNN), are poor at global feature extraction [16]. Later models like TransUNet [17] and Swin-Unet [18] incorporate transformer networks into UNet-based architectures to learn long-range dependencies in medical images, leading to better local and global feature extraction.
Historically, medical image segmentation models have been designed for specific data or targets [19]. Models based on CNN [20], Fully Convolutional Networks (FCN) [21], U-Net [6], and V-Net [22], etc. have been trained to segment specific anatomical structures, including kidneys [23], tumours [23,24,25,26], stroke lesions [27,28], and skin lesions [29]. A recent UNet-based segmentation model is nnU-Net [30]. It is a self-configuring model built specifically for the biomedical domain. The preprocessing, network architecture, training, and post-processing configurations are automatically selected to match the current data set. The authors report that nnU-Net, without manual interventions, outperforms most state-of-the-art (SOTA) models on 23 publicly available biomedical datasets [30]. Task-specific models hardly generalize well to new datasets or tasks without extensive and time-consuming retraining on new data. There is a demand for universal segmentation models: the so-called foundation models, which are capable of few-shot learning (requiring minimal retraining on very few samples), or zero-shot learning (requiring no retraining). The Segment Anything Model (SAM) [31] and its derivatives, such as MedSAM [19], MobileSAM [32], etc., can generalize to new data with little or no fine-tuning. Foundation models have also been applied for medical image segmentation. Mattjie et al. [33], Mazurowski et al. [34], Huang et al. [35] evaluated the robustness of SAM for medical image segmentations in zero-shot mode, where SAM performed better when provided with bounding box prompts. Chen et al. [36] applied SAM for segmentation of ultrasound (US) images from cardiac tissue, thyroid nodules, and the fetal head. The results were good for US images with clear tissue structures but poor for images with shadow artifacts and ambiguous boundaries. Zhang et al. [37] introduced SAM-path, where a pathology encoder eliminated the need for manual prompting. This approach achieved a 5.12% increase in the Dice score [38]. Other applications of foundation models in medical image segmentation can be found in recent surveys [39,40].
Ongoing research attempts to leverage AI systems to improve the capabilities of medical imaging software [41,42], thus creating advanced functionalities beyond traditional medical image processing. Incorporation of DL-based segmentation into medical imaging software enables automatic location of regions of interest, thus minimizing the physician’s task. With the interactive tools provided in these software packages, the physician can interact with the generated segmentations until satisfactory results are achieved. This implies that the more accurate the automatic segmentation, the fewer corrections the physician has to make.
Notable computer platforms that incorporate DL-based image segmentation include Imfusion Suite [43], ITK-SNAP [44], 3D-Slicer [45], and the web-based NORA [46], among others. These tools support features like automatic and interactive segmentations using machine learning. Some of their drawbacks include the need to upload user datasets to online servers, raising possible privacy issues. Also, the incorporated DL segmentation models are limited to a few organs/structures. In addition, we observe no functionality for user-based fine-tuning of the AI segmentation models on the user’s data for custom applications. We envisage that a functionality for user-based model fine-tuning will overcome the situation where the provided segmentation model is limited to a specific organ or image modality.
Whereas deep learning-based medical segmentation models have demonstrated significant potential to improve diagnostic accuracy, and enhance operational efficiency, as well as support data-driven decision-making [47], this progress has largely been achieved under laboratory conditions or using computational environments characterized by high-performance hardware and reliable network infrastructure [48]. This assumption of a high-end computational infrastructure constitutes a critical limitation to the wider deployment and adoption of clinical deep learning solutions. Current state-of-the-art DL models are often characterized by high memory demands and computational complexity, which collectively constitute a bottleneck in model deployment on low-cost, energy-efficient devices [49,50]. Thus, it becomes infeasible to deploy many DL solutions in real-world contexts where resource constraints are rather intrinsic. Zhou et al. [51] noted that the transformer-based image encoder of the high-performance SAM segmentation model was too computationally intensive to run on low-end devices. They highlighted that mitigating this deployment constraint required distilling the SAM model into a computationally light network. Furthermore, Zhou et al. [52] reported that not only the large SAM2 segmentation model encoders but also its novel memory attention blocks posed latency bottlenecks for deployment on low-cost devices. More specifically for medical segmentation, Li et al. [53] reported the computational expensiveness of vision transformer-based models, which would include the likes of SAM and ViT-powered UNet-based segmentation models, thus limiting their deployment on resource-limited devices. Iqbal et al. [54] pointed out that conventional UNet-based architectures fail to meet the speed and efficiency requirements of real-time clinical applications. Furthermore, issues like hardware mismatch, little or no support for high parallelism, absence of tensor cores, and the need for model compression/optimisation imply that model benchmarks on high-end GPUs are not a reliable predictor of model performance on low-end devices [55,56].
While many SOTA medical segmentation models would benefit real use-case deployment, we establish that it is crucial to provide model benchmarks that are a true representation of the model performance on the target end devices. For instance, in resource-constrained clinical settings such as rural and medium-sized hospitals, care homes, and ambulances, where a doctor with a small computer may want to semiautomatically segment structures in medical scans/images, AI-assisted image segmentation tools will be useful if they demonstrate good segmentation accuracy, are light enough to run on low-end devices, and are fast enough to execute within clinically compatible times. Furthermore, many of the existing AI-powered segmentation platforms require user data to be uploaded to external servers, creating privacy concerns. Therefore, we conclude that to improve the adoption of AI-powered medical segmentation in real clinical scenarios, there is a need to develop model benchmarks on low-cost devices that provide insights on how the models comply with the following restraints:
1.
Software and models should be sufficiently lightweight to run on resource-limited devices.
2.
Accuracy must be acceptable for clinical use.
3.
Performance must be within clinically compatible times.
4.
Models should be easy for physicians to extend to other segmentation tasks.
5.
Privacy of medical data should be maintained.
This paper aims to provide systematic model benchmarks that provide insights into how state-of-the-art medical segmentation models comply with the above-mentioned constraints for deployment in resource-limited clinical settings. In particular, we provide insightful outcomes of our experiments on the suitability or otherwise of foundation segmentation models in contrast to task-specific models for deployment on computer platforms for clinical use. In addition, we also make available on our GitHub page the code for easy and quick user-based finetuning of the compliant segmentation model to extend the segmentation capabilities on new data.
The rest of the paper is structured as follows. Section 2 outlines the materials and methods of the research, including the description and justification of the data and the selected segmentation models. In Section 3, we present the results of our experiments. Section 4 discusses the results of the experiments in light of the objectives of the investigations. The work is concluded in Section 5 while highlighting areas for further work.

2. Materials and Methods

2.1. Dataset Description and Justification

For this research, multiple datasets were selected to encompass different imaging modalities and target organs. The first two datasets are publicly available magnetic resonance imaging (MRI) acquisitions of the kidneys [57] and left ventricle myocardium (LVM) [58]. The third is a custom dataset consisting of ultrasound (US) images of the Rectus Femoris (RF) muscle. Refer to the study by [59] for more information on the study that generated the dataset. We also considered structural variation of the organs; thus, the kidney and RF muscle vary from the LVM task in the sense that the LVM has a cavity in its structure.

2.1.1. Kidney Dataset

This publicly available dataset [57] contains T2-weighted abdominal magnetic resonance images, with accompanying manually defined binary masks of the kidneys. The dataset contains 60 volumes, with each volume having between 9 and 13 slices. Healthy control subjects comprise half of the scans, while chronic kidney disease patients comprise the remaining half. To increase the potential for more accurate segmentations, the MRI sequence was designed in such a way as to optimize the contrast between the kidneys and the tissues around them. The ground truth segmentation masks were generated by three specialists trained in kidney segmentation with an average of two years of experience.
Total Kidney Volume (TKV) is routinely used as a parameter for a variety of renal pathologies [60]. For Autosomal Dominant Polycystic Kidney Disease (ADPKD), progression can be monitored by recording TKV: an increase in the TKV parameter is associated with a decrease in renal function in this case. Otherwise, in chronic kidney disease (CKD), a decrease in TKV is associated with a decrease in renal functionality. Thus, renal segmentation is a very important step in many medical procedures related to the kidneys.

2.1.2. Left Ventricle Myocardium Dataset

These data were sourced from the MICCAI 2019 left ventricular full quantification challenge (https://lvquan19.github.io/, accessed on 2 December 2024). It consists of 2D short-axis cine magnetic resonance images of 56 subjects. For each subject, 20 frames were provided for the whole cardiac cycle. The epicardial and endocardial borders were manually contoured and double-checked by two experienced cardiac radiologists [58].
Contours in these types of images are customarily used in clinical practice to assess global functional parameters of the heart, such as the ejection fraction. Magnetic resonance imaging is considered the gold standard that other methods are compared with [61].

2.1.3. Rectus Femoris Muscle Dataset

This data set consists of 269 2D ultrasound images of the Rectus Femoris (RF) muscle in adult patients. The borders of the cross-sectional areas of the belly of the muscle were delineated in the images with dashed lines. This study was approved by the Clinical Trials Committee of our area (code PI233341). All patients signed a written informed consent document before their inclusion in the study, and all images were anonymized for use.
In Disease Related Malnutrition (DRM), the evaluation of the effects of treatment is usually based on changes in muscle function and composition [62,63]. Muscle loss is routinely assessed by morphofunctional measurements. Muscular ultrasound allows us to measure several parameters such as the area and thickness of the RF muscle. In addition, image echogenicity defines thresholds for separating different components within a given tissue: muscle, fat, collagen, connective tissue, and fibrosis. In both cases, contouring of the muscle is necessary.

2.2. Selected Segmentation Models

Keeping in mind the already mentioned restrictions for clinical deployment, foundation segmentation models were considered, since these models, having already learned the general features of images, can be easily and quickly fine-tuned on custom data. Moreover, their zero-shot performance is competitive with many task-specific models. The three versions of the SAM [31]: vit-h, vit-l, and vit-b, as well as MedSAM [19] and MobileSAM [32], were selected. We also evaluated the performance of the foundation models alongside a high-performance target-specific medical segmentation model, the so-called nnU-Net [30]. Whereas some of the selected models are capable of 3D segmentation, the presented experiments in this work were 2D-based. This is also clinically useful given that radiologists often analyse medical images in 2D. These models are described in detail in Appendix A. Ablation studies are critical to understanding how different model components affect model performance [64]. For the interested reader, though beyond the scope of our work, ablation studies for these models exist that evaluate network architecture components, including architecture adaptation and adaptation modules [19,65], prompting strategies [19,66], finetuning strategies [19], feature fusion mechanisms [67], and distillation approaches [32].

2.3. Experiment Environment

For training, the NVIDIA RTX A5000 (24GB) GPU(NVIDIA, Santa Clara, CA, USA) was used. To evaluate the models for user-enabled fine-tuning on low-end computing devices, we configured the best-performing foundation model as well as the trained nnU-Net model on a server CPU (128 CPUs, AMD EPYC 7513 32-Core Processor, AMD, Santa Clara, CA, USA), a desktop CPU (8 core Intel i7 4th Gen, Intel, Santa Clara, CA, USA), and a laptop CPU (quadcore Intel i7 10th Gen, Intel, Santa Clara, CA, USA). The Pytorch framework and other dependencies were set up to fine-tune/train and run the models.

2.4. Segmentation Models Training and Fine-Tuning

2.4.1. Data Preprocessing

This work utilised some preprocessing steps to improve the robustness of the models. All dataset images were rotated at random angles between 0 and 360 degrees to introduce diversification of the training data as well as improve the robustness of the models to object orientations. Intensity scaling was also applied to the images to encourage faster convergence. For the kidney and RFM dataset, the intensity was rescaled to the 2nd and 98th percentiles. For the LV dataset, intensity rescaling was between the 2nd and 95th percentiles. All images were normalised to intensity values within 0–255. The left ventricle and kidney datasets were split into 60:20:20 for training, validation, and testing, respectively. The rectus femorare 2 set, which had a fewer number of images, was split into 80:10:10 for training, validation, and testing, respectively. The above splits were used for all models. The dataset splits are summarised in Table 1.

2.4.2. Training and Validation

According to Gu et al. [68], the SAM encoder could be too parameter-rich for finetuning on small medical image segmentation datasets, thus making the finetuning process difficult. Moreover, Asokan et al. [69], while working to adapt SAM to medical image segmentation, reported that finetuning the encoder on a small medical image segmentation dataset tends to distort the inherent capabilities of the underlying foundation model; hence, their advocation for the retention of the original SAM encoder during finetuning. These conclusions agree with our attempts to finetune the SAM decoder on our comparatively small datasets, which, in our case, failed to converge. Furthermore, the encoder acts as a general-purpose feature extractor, sufficiently learning representations that can be transferred across segmentation tasks, whereas the mask decoder is responsible for task-specific mask prediction. Thus, we focused on fine-tuning the mask decoder of the foundation models. Moreover, the results from fine-tuning just the mask decoder were motivating enough to avoid the time and resource cost of fine-tuning the entire model pipeline, keeping in mind the objective of model deployment on low-end devices.
Figure 1 shows a schema of the mask decoder finetuning pipeline. Our mask decoder finetuning approach entailed freezing the layers of the image and prompt encoders and updating only the weights of the mask decoder. In the forward pass, image embeddings and prompt embeddings are generated. The image encoder generates image embeddings from the input images, while the prompt embeddings are generated by the prompt encoder. The prompt encoder can take different prompt formats, which include points, boxes, and/or masks. In our case, we use box prompts, which we calculated from the ground truth annotations. To simulate user inaccuracies, perturbations were added to the bounding boxes by jittering the boxes in all directions using random integer values ranging from 0 to 20 pixels. The image and prompt embeddings were then fed as inputs to the mask decoder. The segmentation loss is computed on the basis of hybrid dice and cross-entropy loss functions. Thereafter, only the mask decoder weights are updated using Adam optimisation. We implemented early stopping to avoid overfitting.

2.4.3. Model Testing

Model testing involved running inference on the test set and determining accuracy and inference time. The accuracy metrics consisted of the Dice Similarity Coefficient (DSC), 95 percentile Hausdorff distance (HD95), and Hu moments. While DSC and HD are well-known metrics for calculating the similarity between predicted and ground-truth segmentations [70], the less used Hu Moments consist of seven numbers derived from central moments that are invariant to translation, scale, rotation, and reflection [71]. They are useful for shape-matching in image segmentation. Anghelache Nastase et al. [72] incorporated Hu’s moments for classification of malignant and benign breast tumours because of its ability to discriminate shape differences between predicted and ground truth masks. Moreover, Damian et al. [73] reported Hu’s moments as powerful shape descriptors for classifying malignant melanoma and benign lesions. For the HD measure, we opted for HD95 because of its robustness to outliers and noisy boundary predictions as noted by Kim et al. [74] and Wu et al. [75]. This robustness of HD95 to outliers and less sensitivity to noise is the motivation for its use in the popular MICCAI challenges, such as the MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024 [76] and the MICCAI Brain and Tumour Segmentation (BraTS) Challenge [77]. For HD and Hu Moments, lower values mean better accuracy. Conversely, greater values indicate better accuracy for DSC. Library functions from the MedSAM [19] were used to compute the dice coefficients. For the HD95, functions from the Skimage python library were used. And, for the Hu moments, functions from the OpenCV python library were used. The mean inference time per image was also calculated. The testing involved two stages. First, inference was performed only for the foundation models on the server GPU. Based on the three parameters–model size, accuracy, and inference speed–the foundation model with the best trade-off was selected. The second stage involved inferences for the nnU-Net model and comparing the results with the best-performing foundation model. Furthermore, the inference times for the two models were evaluated, as indicated above, on various low-end devices, including the server CPU, a desktop CPU, and a laptop CPU for the two representative datasets.

3. Results

3.1. Training and Inference on GPU

The model names have been given aliases to minimize cluttering on the plots. The aliases are presented in Table 2, which also includes the sizes of the models. Figure 2 shows boxplots of the DSC, HD, and Hu moments for the fine-tuned and original foundation models on the 3 datasets. Inference times for the foundation models on all three datasets are captured in Figure 3.
A radar chart, as shown in Figure 4, shows the trade-off of the foundation models in terms of model accuracy, inference time, and model size. The plot for each model forms a triangle in the radar chart, each vertex corresponding to the value for the model on each of the three axes. The values are normalized to the range [0, 1] by dividing by the maximum among all models in each coordinate. On the three axes, lower values mean better performance. Thus, the closer the triangle is to the centre, the better the trade-off.
Boxplots for the DSC, HD, and Hu moments for the best foundation model (finetuned MobileSAM) and nnU-Net, using LVM and RFM datasets, are shown in Figure 5a,b. The model sizes for the finetuned MobileSAM and nnU-Net are also shown in Table 2.

3.2. Training and Inference on CPU

Having considered the finetuned MobileSAM as the foundation model with the best trade-off between size, accuracy, and speed, evaluation of the training and inference times on the non-GPU devices was carried out for the finetuned MobileSAM and nnU-Net. The finetuned MobileSAM was successfully trained on all CPU devices. However, nnU-Net could only be trained on the server CPU. The desktop training terminal kept crashing, whereas on the laptop, the training was aborted after about 3 h of not progressing beyond the first epoch. Training times per epoch are shown in Table 3. Furthermore, we present in Table 3 also the floating point operations per second (FLOPS) and the number of parameters of the network. The inference times on all devices for the finetuned MobileSAM and nnU-Net are shown in Figure 5c,d for the corresponding datasets. To visually assess the performance of the model, colour-coded masks of segmentation selected samples from the 75th, 50th, and 25th percentiles of HD shown in Figure 2 were plotted for each of the datasets for the finetuned MobileSAM in Figure 6.

4. Discussion

The results of this work shall be discussed in light of the research objectives, especially as regards the above-mentioned constraints that should bound segmentation models for deployment on low-cost devices. We first discuss the accuracy, inference speed, and model size of the foundation models, after which we compare the selected foundation model with the nnU-Net.

4.1. On Foundation Models

4.1.1. Accuracy

Figure 2 shows that the fine-tuned models consistently outperform the models without fine-tuning for the three datasets. This performance is expected, since the fine-tuned models have learned the features specific to the target datasets. All models perform better on the kidney and RFM datasets than on the LVM dataset. This could be attributed to the structure of the target organ/object of interest. Unlike the other datasets, the LVM has a cavity in its structure, which might have presented an additional complexity for the models. The larger fine-tuned foundation models, including f_h and f_l, show no significantly better performance over the smaller models like f_Mb despite having more learnable parameters and deeper architectures. This might be attributed to the possibility that the current segmentation tasks were not overly complex for any of the models. We expect that as the segmentation tasks become more complicated, such as multiple organ/structure segmentation having a wide range of shapes and sizes, the superiority of the larger models might be manifested in an improved accuracy. Given the numerical accuracy figures (see Figure 2) in addition to visual inspection of the segmentation masks (see Figure 6), we consider the accuracy of the fine-tuned models to be acceptable in a clinical environment.

4.1.2. Inference Speed and Model Size

Figure 3 presents the inference speed of the foundation models on the server GPU. Among the three datasets, the models show a consistent speed trend relative to each other. The inference time for the MobileSAM models is as little as 0.07 s per image, which is over 31 times faster than for SAM vit-h models. The direct influence of model size on inference time is evident in the similarity between how sizes and speed vary relative to each other (last column of Table 2). The inference time increases with size, with f_Mb 63 times smaller than f_h, at just 39 MB.

4.1.3. Accuracy-Size-Speed Trade-Off

All fine-tuned models show good accuracy, with no model performing extremely better than others. Thus, we look for model size and speed to determine the model with the best trade-off. The inference time of the fine-tuned MobileSAM is significantly better than other models. The same holds for model size. The tradeoff can be visualized on the radar chart in Figure 4. Therefore, the fine-tuned MobileSAM was chosen as the model with the best accuracy-size-speed trade-off, thus becoming the selected model for performance comparison with the nnU-Net model on low-cost devices.

4.2. MobileSAM vs. nnU-Net on Low-Cost Devices

4.2.1. Accuracy

According to Figure 5c,d, inasmuch as the accuracies of both models are good, nnU-Net shows better accuracy. This underscores an advantage of target/domain-specific models, as the architecture of nnU-Net is designed specifically for the medical imaging domain and thus will tend to outperform on medical imaging datasets. However, training time, size, and inference time for both models must also be taken into consideration to determine which model offers the best solution for low-cost devices.

4.2.2. Memory and Computational Intensity

The size of the model is a critical consideration for low-cost devices. The less memory a model occupies, the better. Table 2 shows that the finetuned MobileSAM at 39 MB is 42 times smaller than nnU-Net. At this size, the finetuned MobileSAM can be deployed to almost any low-end device without the need to apply model quantization. The FLOPS count and number of network parameters, as presented in Table 3, highlight the higher computational intensity of nnUNet over the finetuned MobileSAM, thus favouring the deployment of the lighter network.

4.2.3. Model Tasks Expansion

To understand the ease with which the models can be extended to cover more segmentation tasks, it is necessary to assess how quickly the models can be retrained or fine-tuned on new datasets, not just on the server GPU, but, very importantly, on a user’s low-cost device. The training times per epoch for the RF muscle dataset on the four machines of interest are shown in Table 3. nnU-Net was successfully trained on the server GPU and server CPU but failed on the desktop and laptop CPUs. On the desktop CPU, the training terminal was unable to handle the training and shut down. On the laptop, the training froze in the first epoch without any output results. The MobileSAM was successfully fine-tuned on all devices and at much lower training times. For instance, on the server CPU, the training time per epoch for MobileSAM is close to 1000 times less than that of nnU-Net.

4.3. Device Execution Times

In real clinical settings, it is desirable that segmentations are done on the fly—the faster the segmentation model, the better. As observed in Figure 5c,d, MobileSAM outclasses nnU-Net in model inference time. On the laptop, the finetuned MobileSAM segments an image in as little as 1.5 s, about 12 times and 75 times faster than nnU-Net on the LVM and RFM datasets, respectively. We present in Table 4 a ranking of the two models with respect to the parameters under consideration, with some remarks.

4.4. Case Visualisation

In Figure 6, we present sample segmentations for the finetuned MobileSAM—the model with the best tradeoff for deployment on resource-constrained devices. Sample masks are selected from the 75th, 50th, and 25th percentiles of the HD for all three datasets. Thus, we are able to visualize the best case and worst-case segmentations of the finetuned MobileSAM. At this point we wish to highlight that the MobileSAM model has an inbuilt feature to automatically deal with ambiguity, wherein the network generates three segmentation masks. Thereafter, the network automatically outputs the mask with the best intersection over union value. This feature, combined with adequately annotated data, minimises cases of total failure of the model. We also note that from Figure 6, the finetuned MobileSAM tends to exhibit uncertainties around the object boundaries. This is where the prompt feature of MobileSAM becomes more useful. The user can add more prompts to get better segmentation results at the boundaries. For an in-depth study on uncertainty estimation for SAM-based models, including the MobileSAM, see Deng et al. [78] and Zhang et al. [79].

5. Conclusions and Future Work

The adoption of efficient and lightweight AI-assisted medical imaging is important in a clinical setup such as small- to medium-size hospitals, rural hospitals, and on-board equipment in ambulances, where radiologists and high-end computing resources are limited in supply. However, the challenge of adopting AI-assisted radiological solutions can be traced to the constraints established in this paper. In this work, several state-of-the-art foundation segmentation models were benchmarked on various medical image datasets to evaluate their candidacy to deploy on low-cost devices in a real clinical scenario. The MobileSAM model provides the best trade-off in terms of model accuracy, model size, and inference speed. Subsequently, MobileSAM was evaluated alongside the task-specific nnU-Net segmentation model to determine how much they overcome the set requirements. We conclude that in a situation where a clinician with a laptop wants to semiautomatically segment structures in medical scans/images, MobileSAM offers the best solution with good accuracy, minimal memory footprint, real-time inference speed, and the ability to be quickly fine-tuned on additional user data on low-end computing devices. We will make the finetuning scripts freely available on our Github page.
For the future, we foresee possibilities for extending this work in several directions. First, given the existence of numerous types of low-end devices, it is necessary to evaluate the benchmarked models on at least 10 representative resource-constrained devices, while including energy per inference values in the reported performance metrics. The second focus would be to investigate model optimisation strategies like pruning, quantisation, and distillation to reduce the inference latency by up to 50% while maintaining model accuracy within 2% of the baseline score. The third direction would focus on extending the benchmark studies to cover at least 5 additional segmentation tasks and 3 imaging modalities, including CT, MRI, and ultrasound. A final future direction would be to validate model usability in real clinical environments by conducting user studies with up to 10 clinicians. Reported performance indices would include annotation time reduction using AI-assisted segmentation as well as clinically acceptable thresholds.

Author Contributions

Conceptualisation: E.C.N., S.M.-C., M.M.-F. and C.A.-L.; Data curation: E.C.N., S.M.-C., D.A.d.L.-R., M.M.-F. and C.A.-L.; Methodology: E.C.N., S.M.-C., M.M.-F. and C.A.-L.; Writing original draft: E.C.N., M.M.-F. and C.A.-L.; Writing review and editing: E.C.N., S.M.-C., D.A.d.L.-R., M.M.-F. and C.A.-L.; Funding acquisition: E.C.N., M.M.-F. and C.A.-L. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the Spanish Agencia Estatal de Investigación, under Grants PID2020-115339RB-I00 and TED2021-130090B-I00 and CPP2021-008880. E. C. Nnadozie has been funded by the call for UVa 2023 predoctoral contracts co-funded by Banco Santander.

Institutional Review Board Statement

The study that produced the RF dataset was conducted in accordance with the Declaration of Helsinki of 1975, revised in 2013. The study protocol received approval from the Ethics Committee for Clinical Research of the Health Council of the Clinical University Hospital of Valladolid (HCUVA) (protocol code PI233341, approval date 9 November 2023), as well as from the individual Institutional Review Boards of the participating hospitals.

Informed Consent Statement

This study was approved by the Clinical Trials Committee of our area (code PI233341). All patients signed a written informed consent document before their inclusion in the study, and all images were anonymized for use.

Data Availability Statement

The Rectus Femoris Muscle dataset will be made available upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Description of Selected Foundation Segmentation Models

Appendix A.1. Segment Anything Model

The creation of SAM [31] was inspired by the success recorded with Large Language Models (LLM) for zero-shot and few-shot tasks. The SAM model is a foundation model for segmentation tasks that is able to return a valid segmentation mask given any segmentation prompt. A segmentation prompt, in this case, specifies what to segment in the image and could be a bounding box, one or more points, or text. In the absence of prompts, SAM will attempt to segment all objects in the scene. The SAM model consists of an image encoder for computing image embeddings, a prompt encoder for embedding prompts, and a mask decoder, where the information from the image and prompt encoders is combined to predict the output masks. The separation of the mask decoder from the image encoder ensures that the mask encoder can be prompted multiple times to generate masks for different objects in the same image without the need to recompute the image embeddings. This approach provides real-time usability of the model. SAM was trained on the SA-1B dataset, containing more than 11 million images and up to 1 billion masks [31].

Appendix A.2. MedSAM

SAM was trained mostly on natural images and exhibited reduced performance on medical images, especially for images with weak boundaries or low contrast [34]. MedSAM was proposed as a foundation model for medical images, achieved by fine-tuning SAM for more than 1.5 million medical image-mask pairs. This dataset, which covers a wide spectrum of anatomical structures and lesions across different imaging modalities and protocols, enables the MedSAM model to learn a rich representation of medical images. The authors reported a consistent outperformance of MedSAM over other state-of-the-art foundation segmentation models [19].

Appendix A.3. MobileSAM

The pre-trained weights of the SAM models are large, with the largest version, vit-h, in excess of 2 GB, and the smallest, vit-b, being about 300 MB. This could pose a problem in low-cost devices with limited memory and computing capacity. MobileSAM [32] addressed this challenge by distilling the knowledge from the SAM image encoder to a smaller one, while keeping the prompt encoder and mask decoder from the original SAM model. The authors reported competitive segmentation accuracy with respect to the original SAM and more than a 30-times increase in inference speed at much lower computational resources.

Appendix B. Data Preprocessing and Model Training

Several pre-processing steps were carried out on the datasets to match the input requirements of each model, as well as to boost model training performance. The SAM model takes a 2D image as input from the original 3D magnetic resonance data. Other preprocessing steps included rotations and intensity rescaling.
Data sets were split into training, validation, and test (inference) sets. Furthermore, we performed a 10-fold cross-validation strategy. The SAM, MedSAM, and MobileSAM had similar fine-tuning steps. The image embeddings were computed, and bounding boxes were generated from the ground truth masks, which were used as input prompts during the training stage. The image encoder was frozen while the mask decoder was fine-tuned. Validation was performed simultaneously with training on each epoch. Early stopping was implemented to prevent model over-fitting to current data.
Fine-tuning nnU-Net involved a complex process of having to match the data fingerprints for the pre-trained model and the new data. This required training first on the original data before fine-tuning on the new data. Thus, given the long training time for nnU-Net, we considered it more useful to train nnU-Net from scratch using our datasets since the model was capable of generating good performance on small datasets as well. Due to the long training time for nnU-Net, we also switched to 5-fold rather than 10-fold cross-validation. We trained nnU-Net on two of our datasets, representing two imaging modalities.

References

  1. Bajwa, J.; Munir, U.; Nori, A.; Williams, B. Artificial intelligence in healthcare: Transforming the practice of medicine. Future Healthc. J. 2021, 8, e188–e194. [Google Scholar] [CrossRef] [Scilit]
  2. De Fauw, J.; Ledsam, J.R.; Romera-Paredes, B.; Nikolov, S.; Tomasev, N.; Blackwell, S.; Askham, H.; Glorot, X.; O’Donoghue, B.; Visentin, D.; et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nat. Med. 2018, 24, 1342–1350. [Google Scholar] [CrossRef] [Scilit]
  3. Jain, A.K. Fundamentals of Digital Image Processing; Prentice-Hall, Inc.: Englewood Cliffs, NJ, USA, 1989; Available online: https://dl.acm.org/doi/book/10.5555/59921 (accessed on 19 March 2024).
  4. Yao, W.; Bai, J.; Liao, W.; Chen, Y.; Liu, M.; Xie, Y. From CNN to Transformer: A Review of Medical Image Segmentation Models. J. Imaging Inform. Med. 2024, 37, 1529–1547. [Google Scholar] [CrossRef] [Scilit]
  5. Antonelli, M.; Reinke, A.; Bakas, S.; Farahani, K.; Kopp-Schneider, A.; Landman, B.A.; Litjens, G.; Menze, B.; Ronneberger, O.; Summers, R.M.; et al. The Medical Segmentation Decathlon. Nat. Commun. 2022, 13, 4128. [Google Scholar] [CrossRef] [Scilit]
  6. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F., Eds.; Springer International Publishing: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
  7. Çiçek, Ö.; Abdulkadir, A.; Lienkamp, S.S.; Brox, T.; Ronneberger, O. 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention—MICCAI 2016; Ourselin, S., Joskowicz, L., Sabuncu, M.R., Unal, G., Wells, W., Eds.; Springer International Publishing: Cham, Switzerland, 2016; pp. 424–432. [Google Scholar]
  8. Zhang, Z.; Fu, H.; Dai, H.; Shen, J.; Pang, Y.; Shao, L. ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2019; Shen, D., Liu, T., Peters, T.M., Staib, L.H., Essert, C., Zhou, S., Yap, P.T., Khan, A., Eds.; Springer International Publishing: Cham, Switzerland, 2019; pp. 442–450. [Google Scholar]
  9. Schlemper, J.; Oktay, O.; Schaap, M.; Heinrich, M.; Kainz, B.; Glocker, B.; Rueckert, D. Attention gated networks: Learning to leverage salient regions in medical images. Med. Image Anal. 2019, 53, 197–207. [Google Scholar] [CrossRef] [Scilit]
  10. Chen, L.; Bentley, P.; Mori, K.; Misawa, K.; Fujiwara, M.; Rueckert, D. DRINet for Medical Image Segmentation. IEEE Trans. Med. Imaging 2018, 37, 2453–2462. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Alom, M.Z.; Yakopcic, C.; Taha, T.M.; Asari, V.K. Nuclei Segmentation with Recurrent Residual Convolutional Neural Networks based U-Net (R2U-Net). In Proceedings of the NAECON 2018—IEEE National Aerospace and Electronics Conference, Dayton, OH, USA, 23–26 July 2018; IEEE: New York, NY, USA, 2018; pp. 228–233. [Google Scholar] [CrossRef] [Scilit]
  12. Ibtehaz, N.; Rahman, M.S. MultiResUNet: Rethinking the U-Net architecture for multimodal biomedical image segmentation. Neural Netw. 2020, 121, 74–87. [Google Scholar] [CrossRef] [Scilit]
  13. Zhang, Z.; Wu, C.; Coleman, S.; Kerr, D. DENSE-INception U-Net for medical image segmentation. Comput. Methods Programs Biomed. 2020, 192, 105395. [Google Scholar] [CrossRef] [Scilit]
  14. Wang, Z.H.; Liu, Z.; Song, Y.Q.; Zhu, Y. Densely connected deep U-Net for abdominal multi-organ segmentation. In Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan, 22–25 September 2019; IEEE: New York, NY, USA, 2019; pp. 1415–1419. [Google Scholar] [CrossRef] [Scilit]
  15. Rahman, H.; Ben Aoun, N.; Bukht, T.F.N.; Ahmad, S.; Tadeusiewicz, R.; Pławiak, P.; Hammad, M. Automatic liver tumor segmentation of CT and MRI volumes using ensemble ResUNet-InceptionV4 model. Inf. Sci. 2025, 704, 121966. [Google Scholar] [CrossRef] [Scilit]
  16. Xiao, B.; Xu, B.; Bi, X.; Li, W. Global-Feature Encoding U-Net (GEU-Net) for Multi-Focus Image Fusion. IEEE Trans. Image Process. 2021, 30, 163–175. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Chen, J.; Mei, J.; Li, X.; Lu, Y.; Yu, Q.; Wei, Q.; Luo, X.; Xie, Y.; Adeli, E.; Wang, Y.; et al. TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers. Med. Image Anal. 2024, 97, 103280. [Google Scholar] [CrossRef] [Scilit]
  18. Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation. In Proceedings of the Computer Vision—ECCV 2022 Workshops: Tel Aviv, Israel, 23–27 October 2022; Proceedings, Part III; Springer: Berlin/Heidelberg, Germany, 2022; pp. 205–218. [Google Scholar] [CrossRef] [Scilit]
  19. Ma, J.; He, Y.; Li, F.; Han, L.; You, C.; Wang, B. Segment anything in medical images. Nat. Commun. 2024, 15, 654. [Google Scholar] [CrossRef] [Scilit]
  20. Gu, J.; Wang, Z.; Kuen, J.; Ma, L.; Shahroudy, A.; Shuai, B.; Liu, T.; Wang, X.; Wang, G.; Cai, J.; et al. Recent advances in convolutional neural networks. Pattern Recognit. 2018, 77, 354–377. [Google Scholar] [CrossRef] [Scilit]
  21. Zhou, X.; Takayama, R.; Wang, S.; Hara, T.; Fujita, H. Deep learning of the sectional appearances of 3D CT images for anatomical structure segmentation based on an FCN voting method. Med. Phys. 2017, 44, 5221–5233. [Google Scholar] [CrossRef] [Scilit]
  22. Milletari, F.; Navab, N.; Ahmadi, S.A. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; IEEE: New York, NY, USA, 2016; pp. 565–571. [Google Scholar] [CrossRef] [Scilit]
  23. Türk, F.; Lüy, M.; Barışçı, N. Kidney and Renal Tumor Segmentation Using a Hybrid V-Net-Based Model. Mathematics 2020, 8, 1772. [Google Scholar] [CrossRef] [Scilit]
  24. Pereira, S.; Pinto, A.; Alves, V.; Silva, C.A. Brain Tumor Segmentation Using Convolutional Neural Networks in MRI Images. IEEE Trans. Med. Imaging 2016, 35, 1240–1251. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Ghaffari, M.; Sowmya, A.; Oliver, R. Automated Brain Tumour Segmentation Using Cascaded 3D Densely-Connected U-Net. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Crimi, A., Bakas, S., Eds.; Springer International Publishing: Cham, Switzerland, 2021; pp. 481–491. [Google Scholar]
  26. Wang, G.; Li, W.; Ourselin, S.; Vercauteren, T. Automatic Brain Tumor Segmentation Using Cascaded Anisotropic Convolutional Neural Networks. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Crimi, A., Bakas, S., Kuijf, H., Menze, B., Reyes, M., Eds.; Springer International Publishing: Cham, Switzerland, 2018; pp. 178–190. [Google Scholar]
  27. Dolz, J.; Ben Ayed, I.; Desrosiers, C. Dense Multi-path U-Net for Ischemic Stroke Lesion Segmentation in Multiple Image Modalities. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Crimi, A., Bakas, S., Kuijf, H., Keyvan, F., Reyes, M., van Walsum, T., Eds.; Springer International Publishing: Cham, Switzerland, 2019; pp. 271–282. [Google Scholar]
  28. Tureckova, A.; Rodríguez-Sánchez, A.J. ISLES Challenge: U-Shaped Convolution Neural Network with Dilated Convolution for 3D Stroke Lesion Segmentation. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Crimi, A., Bakas, S., Kuijf, H., Keyvan, F., Reyes, M., van Walsum, T., Eds.; Springer International Publishing: Cham, Switzerland, 2019; pp. 319–327. [Google Scholar]
  29. Yuan, Y.; Chao, M.; Lo, Y.C. Automatic Skin Lesion Segmentation Using Deep Fully Convolutional Networks With Jaccard Distance. IEEE Trans. Med. Imaging 2017, 36, 1876–1886. [Google Scholar] [CrossRef] [Scilit]
  30. Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. arXiv 2023, arXiv:2304.02643. [Google Scholar] [CrossRef] [Scilit]
  32. Zhang, C.; Han, D.; Qiao, Y.; Kim, J.U.; Bae, S.H.; Lee, S.; Hong, C.S. Faster Segment Anything: Towards Lightweight SAM for Mobile Applications. arXiv 2023, arXiv:2306.14289. [Google Scholar] [CrossRef] [Scilit]
  33. Mattjie, C.; De Moura, L.V.; Ravazio, R.; Kupssinskü, L.; Parraga, O.; Delucis, M.M.; Barros, R.C. Zero-Shot Performance of the Segment Anything Model (SAM) in 2D Medical Imaging: A Comprehensive Evaluation and Practical Guidelines. In Proceedings of the 2023 IEEE 23rd International Conference on Bioinformatics and Bioengineering (BIBE); IEEE: New York, NY, USA, 2023; pp. 108–112. [Google Scholar] [CrossRef] [Scilit]
  34. Mazurowski, M.A.; Dong, H.; Gu, H.; Yang, J.; Konz, N.; Zhang, Y. Segment anything model for medical image analysis: An experimental study. Med. Image Anal. 2023, 89, 102918. [Google Scholar] [CrossRef] [Scilit]
  35. Huang, Y.; Yang, X.; Liu, L.; Zhou, H.; Chang, A.; Zhou, X.; Chen, R.; Yu, J.; Chen, J.; Chen, C.; et al. Segment anything model for medical images? Med. Image Anal. 2024, 92, 103061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Chen, F.; Chen, L.; Han, H.; Zhang, S.; Zhang, D.; Liao, H. The ability of Segmenting Anything Model (SAM) to segment ultrasound images. Biosci. Trends 2023, 17, 2023.01128. [Google Scholar] [CrossRef] [Scilit]
  37. Zhang, J.; Ma, K.; Kapse, S.; Saltz, J.; Vakalopoulou, M.; Prasanna, P.; Samaras, D. SAM-Path: A Segment Anything Model for Semantic Segmentation in Digital Pathology. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2023 Workshops; Celebi, M.E., Salekin, M.S., Kim, H., Albarqouni, S., Barata, C., Halpern, A., Tschandl, P., Combalia, M., Liu, Y., Zamzmi, G., et al., Eds.; Springer International Publishing: Cham, Switzerland, 2023; pp. 161–170. [Google Scholar] [CrossRef] [Scilit]
  38. Dice, L.R. Measures of the amount of ecologic association between species. Ecology 1945, 26, 297–302. [Google Scholar] [CrossRef] [Scilit]
  39. Zhang, S.; Metaxas, D. On the challenges and perspectives of foundation models for medical image analysis. Med. Image Anal. 2024, 91, 102996. [Google Scholar] [CrossRef] [Scilit]
  40. Zhang, Y.; Shen, Z.; Jiao, R. Segment anything model for medical image segmentation: Current applications and future directions. Comput. Biol. Med. 2024, 171, 108238. [Google Scholar] [CrossRef] [Scilit]
  41. Theriault-Lauzier, P.; Cobin, D.; Tastet, O.; Langlais, E.L.; Taji, B.; Kang, G.; Chong, A.Y.; So, D.; Tang, A.; Gichoya, J.W.; et al. A Responsible Framework for Applying Artificial Intelligence on Medical Images and Signals at the Point of Care: The PACS-AI Platform. Can. J. Cardiol. 2024, 40, 1828–1840. [Google Scholar] [CrossRef] [Scilit]
  42. Szilágyi, L.; Kovács, L. Special Issue: Artificial Intelligence Technology in Medical Image Analysis. Appl. Sci. 2024, 14, 2180. [Google Scholar] [CrossRef] [Scilit]
  43. Zettinig, O.; Salehi, M.; Prevost, R.; Wein, W. Recent Advances in Point-of-Care Ultrasound Using the ImFusion Suite for Real-Time Image Analysis. In Proceedings of the Simulation, Image Processing, and Ultrasound Systems for Assisted Diagnosis and Navigation: International Workshops, POCUS 2018, BIVPCS 2018, CuRIOUS 2018, and CPM 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, 16–20 September 2018; Springer: Berlin/Heidelberg, Germany, 2018; pp. 47–55. [Google Scholar] [CrossRef] [Scilit]
  44. Yushkevich, P.A.; Piven, J.; Cody Hazlett, H.; Gimpel Smith, R.; Ho, S.; Gee, J.C.; Gerig, G. User-Guided 3D Active Contour Segmentation of Anatomical Structures: Significantly Improved Efficiency and Reliability. Neuroimage 2006, 31, 1116–1128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Fedorov, A.; Beichel, R.; Kalpathy-Cramer, J.; Finet, J.; Fillion-Robin, J.C.; Pujol, S.; Bauer, C.; Jennings, D.; Fennessy, F.; Sonka, M.; et al. 3D Slicer as an image computing platform for the Quantitative Imaging Network. Magn. Reson. Imaging 2012, 30, 1323–1341. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Anastasopoulos, C.; Reisert, M.; Kellner, E. “Nora Imaging”: A Web-Based Platform for Medical Imaging. Neuropediatrics 2017, 48, S1–S45. [Google Scholar] [CrossRef] [Scilit]
  47. Rayed, M.E.; Islam, S.S.; Niha, S.I.; Jim, J.R.; Kabir, M.M.; Mridha, M. Deep learning for medical image segmentation: State-of-the-art advancements and challenges. Inform. Med. Unlocked 2024, 47, 101504. [Google Scholar] [CrossRef] [Scilit]
  48. Ekman, T.; Barakat, A.; Heiberg, E. Generalizable deep learning framework for 3D medical image segmentation using limited training data. 3D Print. Med. 2025, 11, 9. [Google Scholar] [CrossRef] [Scilit]
  49. Li, Z.; Li, H.; Meng, L. Model Compression for Deep Neural Networks: A Survey. Computers 2023, 12, 60. [Google Scholar] [CrossRef] [Scilit]
  50. Dai, D.; Dong, C.; Yang, X.; Li, Z.; Xu, S. Striking a better balance between segmentation performance and computational costs with a minimalistic network design. Appl. Soft Comput. 2025, 182, 113549. [Google Scholar] [CrossRef] [Scilit]
  51. Zhou, C.; Li, X.; Loy, C.C.; Dai, B. EdgeSAM: Prompt-In-the-Loop Distillation for SAM. Int. J. Comput. Vis. 2025, 133, 8452–8468. [Google Scholar] [CrossRef] [Scilit]
  52. Zhou, C.; Zhu, C.; Xiong, Y.; Suri, S.; Xiao, F.; Wu, L.; Krishnamoorthi, R.; Dai, B.; Loy, C.C.; Chandra, V.; et al. EdgeTAM: On-Device Track Anything Model. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 10–17 June 2025; IEEE: New York, NY, USA, 2025; pp. 13832–13842. [Google Scholar] [CrossRef] [Scilit]
  53. Li, X.; Zhu, W.; Dong, X.; Dumitrascu, O.M.; Wang, Y. EViT-UNET: U-Net Like Efficient Vision Transformer for Medical Image Segmentation on Mobile and Edge Devices. In Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), Houston, TX, USA, 14–17 April 2025; IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  54. Iqbal, S.; Khan, T.M.; Naqvi, S.S.; Naveed, A.; Usman, M.; Khan, H.A.; Razzak, I. LDMRes-Net: A Lightweight Neural Network for Efficient Medical Image Segmentation on IoT and Edge Devices. IEEE J. Biomed. Health Inform. 2024, 28, 3860–3871. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Zagitov, A.; Chebotareva, E.; Toschev, A.; Magid, E. Comparative analysis of neural network models performance on low-power devices for a real-time object detection task. Comput. Opt. 2024, 48, 242–252. [Google Scholar] [CrossRef] [Scilit]
  56. Cantero, D.; Esnaola-Gonzalez, I.; Miguel-Alonso, J.; Jauregi, E. Benchmarking Object Detection Deep Learning Models in Embedded Devices. Sensors 2022, 22, 4205. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Daniel, A.J.; Buchanan, C.E.; Allcock, T.; Scerri, D.; Cox, E.F.; Prestwich, B.L.; Francis, S.T. T2-Weighted Kidney MRI Segmentation [Dataset]. 2021. Available online: https://zenodo.org/records/5153568 (accessed on 19 March 2024).
  58. Xue, W.; Brahm, G.; Pandey, S.; Leung, S.; Li, S. Full left ventricle quantification via deep multitask relationships learning. Med. Image Anal. 2018, 43, 54–65. [Google Scholar] [CrossRef] [Scilit]
  59. García-Herreros, S.; López Gómez, J.J.; Cebria, A.; Izaola, O.; Salvador Coloma, P.; Nozal, S.; Cano, J.; Primo, D.; Godoy, E.J.; de Luis, D. Validation of an Artificial Intelligence-Based Ultrasound Imaging System for Quantifying Muscle Architecture Parameters of the Rectus Femoris in Disease-Related Malnutrition (DRM). Nutrients 2024, 16, 1806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Daniel, A.J.; Buchanan, C.E.; Allcock, T.; Scerri, D.; Cox, E.F.; Prestwich, B.L.; Francis, S.T. Automated renal segmentation in healthy and chronic kidney disease subjects using a convolutional neural network. Magn. Reson. Med. 2021, 86, 1125–1136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Schwaiger, J.P.; Reinstadler, S.J.; Tiller, C.; Holzknecht, M.; Reindl, M.; Mayr, A.; Graziadei, I.; Müller, S.; Metzler, B.; Klug, G. Baseline LV ejection fraction by cardiac magnetic resonance and 2D echocardiography after ST-elevation myocardial infarction—influence of infarct location and prognostic impact. Eur. Radiol. 2020, 30, 663–671. [Google Scholar] [CrossRef] [Scilit]
  62. López-Gómez, J.J.; Primo-Martín, D.; Cebria, A.; Izaola-Jauregui, O.; Godoy, E.J.; Pérez-López, P.; Jiménez Sahagún, R.; Ramos Bachiller, B.; González Gutiérrez, J.; De Luis Román, D.A. Effectiveness of High-Protein Energy-Dense Oral Supplements on Patients with Malnutrition Using Morphofunctional Assessment with AI-Assisted Muscle Ultrasonography: A Real-World One-Arm Study. Nutrients 2024, 16, 3136. [Google Scholar] [CrossRef] [Scilit]
  63. de Luis, D.; Cebria, A.; Primo, D.; Nozal, S.; Izaola, O.; Godoy, E.J.; Lopez-Gomez, J.J. Impact of Hydroxy-Methyl-Butyrate Supplementation on Malnourished Patients Assessed Using AI-Enhanced Ultrasound Imaging. J. Cachexia Sarcopenia Muscle 2025, 16, e13700. [Google Scholar] [CrossRef] [Scilit]
  64. Zhang, Y.; Wang, T.; Xue, L.; Lian, W.; Tao, R. ORSI Salient Object Detection via Progressive Interaction and Saliency-Guided Enhancement. IEEE Geosci. Remote Sens. Lett. 2026, 23, 1–5. [Google Scholar] [CrossRef] [Scilit]
  65. Gu, Y.; Wu, Q.; Tang, H.; Mai, X.; Shu, H.; Li, B.; Chen, Y. LeSAM: Adapt Segment Anything Model for Medical Lesion Segmentation. IEEE J. Biomed. Health Inform. 2024, 28, 6031–6041. [Google Scholar] [CrossRef] [Scilit]
  66. Zhang, Y.; Song, Y.; Liu, J.; Li, M. An automatic laryngoscopic image segmentation system based on SAM prompt engineering: From glottis annotation to vocal fold segmentation. Front. Mol. Biosci. 2025, 12, 1616271. [Google Scholar] [CrossRef] [Scilit]
  67. Tang, S.; Wang, S.; Xiang, G.; Zhao, J.; Wang, Y. TA-MedSAM: Text-augmented improved MedSAM for pulmonary lesion segmentation. Comput. Med. Imaging Graph. 2026, 128, 102698. [Google Scholar] [CrossRef] [Scilit]
  68. Gu, H.; Dong, H.; Yang, J.; Mazurowski, M.A. How to build the best medical image segmentation algorithm using foundation models: A comprehensive empirical study with Segment Anything Model. Mach. Learn. Biomed. Imaging 2025, 3, 88–120. [Google Scholar] [CrossRef] [Scilit]
  69. Asokan, M.; Benjamin, J.G.; Yaqub, M.; Nandakumar, K. A Federated Learning-Friendly Approach for Parameter-Efficient Fine-Tuning of SAM in 3D Segmentation. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2024 Workshops; Celebi, M.E., Reyes, M., Chen, Z., Li, X., Eds.; Springer Nature: Cham, Switzerland, 2025; pp. 226–235. [Google Scholar]
  70. Müller, D.; Soto-Rey, I.; Kramer, F. Towards a guideline for evaluation metrics in medical image segmentation. BMC Res. Notes 2022, 15, 210. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Zhong, Y.; Liu, Y.; Liu, K.; Zhan, T.; Liu, S.; Liang, Y.; Hu, Y.; Li, M.; Lei, G.; Zhou, S.; et al. WC electron microscopy image segmentation based on improved watershed and Hu-moment edge matching algorithms. Comput. Mater. Sci. 2025, 246, 113401. [Google Scholar] [CrossRef] [Scilit]
  72. Anghelache Nastase, I.N.; Moldovanu, S.; Moraru, L. Image Moment-Based Features for Mass Detection in Breast US Images via Machine Learning and Neural Network Classification Models. Inventions 2022, 7, 42. [Google Scholar] [CrossRef] [Scilit]
  73. Damian, F.A.; Moldovanu, S.; Dey, N.; Ashour, A.S.; Moraru, L. Feature Selection of Non-Dermoscopic Skin Lesion Images for Nevus and Melanoma Classification. Computation 2020, 8, 41. [Google Scholar] [CrossRef] [Scilit]
  74. Kim, T.; On, S.; Gwon, J.G.; Kim, N. Computed tomography-based automated measurement of abdominal aortic aneurysm using semantic segmentation with active learning. Sci. Rep. 2024, 14, 8924. [Google Scholar] [CrossRef] [Scilit]
  75. Wu, B.; Zhang, F.; Xu, L.; Shen, S.; Shao, P.; Sun, M.; Liu, P.; Yao, P.; Xu, R.X. Modality preserving U-Net for segmentation of multimodal medical images. Quant. Imaging Med. Surg. 2023, 13, 5242–5257. [Google Scholar] [CrossRef] [Scilit]
  76. Linardos, A.; Pati, S.; Baid, U.; Edwards, B.; Foley, P.; Ta, K.; Chung, V.; Sheller, M.; Khan, M.I.; Jafaritadi, M.; et al. The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods. Mach. Learn. Biomed. Imaging 2025, 3, 757–774. [Google Scholar] [CrossRef] [Scilit]
  77. Peiris, H.; Chen, Z.; Egan, G.; Harandi, M. Reciprocal Adversarial Learning for Brain Tumor Segmentation: A Solution to BraTS Challenge 2021 Segmentation Task. In Proceedings of the Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Crimi, A., Bakas, S., Eds.; Springer International Publishing: Cham, Switzerland, 2022; pp. 171–181. [Google Scholar]
  78. Deng, G.; Zou, K.; Ren, K.; Wang, M.; Yuan, X.; Ying, S.; Fu, H. SAM-U: Multi-box Prompts Triggered Uncertainty Estimation for Reliable SAM in Medical Image. In Proceedings of the Medical Image Computing and Computer Assisted Intervention—MICCAI 2023 Workshops; Woo, J., Hering, A., Silva, W., Li, X., Fu, H., Liu, X., Xing, F., Purushotham, S., Mathai, T.S., Mukherjee, P., et al., Eds.; Springer International Publishing: Cham, Switzerland, 2023; pp. 368–377. [Google Scholar]
  79. Zhang, Y.; Hu, S.; Xue, L.; Ren, S.; Hu, Z.; Cheng, Y.; Qi, Y. Enhancing the Reliability of Auto-Prompting SAM for Medical Image Segmentation with Uncertainty Estimation and Rectification. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Honolulu, HI, USA, 19–20 October 2025; IEEE: New York, NY, USA, 2025; pp. 1293–1302. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Mask decoder finetuning pipeline.
Figure 1. Mask decoder finetuning pipeline.
Bdcc 10 00142 g001
Figure 2. Representation for the accuracy of the foundation models on (a) the kidney dataset; (b) LVM dataset; and (c) RF Muscle dataset.
Figure 2. Representation for the accuracy of the foundation models on (a) the kidney dataset; (b) LVM dataset; and (c) RF Muscle dataset.
Bdcc 10 00142 g002
Figure 3. Representation for the inference speed of the foundation models on (a) the kidney dataset; (b) the LVM dataset; and (c) the RF Muscle dataset.
Figure 3. Representation for the inference speed of the foundation models on (a) the kidney dataset; (b) the LVM dataset; and (c) the RF Muscle dataset.
Bdcc 10 00142 g003
Figure 4. A visualization of the tradeoff of the foundational models in terms of accuracy (Hausdorff distance), inference time, and model size.
Figure 4. A visualization of the tradeoff of the foundational models in terms of accuracy (Hausdorff distance), inference time, and model size.
Bdcc 10 00142 g004
Figure 5. Comparison of the accuracy of finetuned MobileSAM and nnU-Net models on (a) LVM dataset and (b) RF muscle dataset. Inference speed of finetuned MobileSAM and nnU-Net on (c) LVM dataset and (d) RF muscle dataset on all devices is also shown.
Figure 5. Comparison of the accuracy of finetuned MobileSAM and nnU-Net models on (a) LVM dataset and (b) RF muscle dataset. Inference speed of finetuned MobileSAM and nnU-Net on (c) LVM dataset and (d) RF muscle dataset on all devices is also shown.
Bdcc 10 00142 g005
Figure 6. Visualization of finetuned MobileSAM colour-coded segmentation selected samples from the 75th (leftmost column, letters (a,d,g)), 50th (middle column, letters (b,e,h)), and 25th (rightmost column, letters (c,f,i)) percentiles cases of the HD shown in Figure 2. Datasets are Kidney (uppermost row, letters (ac)), LVM (middle row, letters (df)), and RF muscle (lowermost row, letters (gi)). Yellow stands for True Positives, Blue for False Negatives (i.e., under-segmentation), and Purple for False Positives (i.e., over-segmentation).
Figure 6. Visualization of finetuned MobileSAM colour-coded segmentation selected samples from the 75th (leftmost column, letters (a,d,g)), 50th (middle column, letters (b,e,h)), and 25th (rightmost column, letters (c,f,i)) percentiles cases of the HD shown in Figure 2. Datasets are Kidney (uppermost row, letters (ac)), LVM (middle row, letters (df)), and RF muscle (lowermost row, letters (gi)). Yellow stands for True Positives, Blue for False Negatives (i.e., under-segmentation), and Purple for False Positives (i.e., over-segmentation).
Bdcc 10 00142 g006aBdcc 10 00142 g006b
Table 1. Dataset train, validation, and test splits.
Table 1. Dataset train, validation, and test splits.
DatasetContentTrain (%)Val (%)Test (%)
Kidney60 vols (9–13 slices/vol)602020
Left ventricle myocardium56 patients (20 frames/patient)602020
Rectus femoris muscle269 images801010
Table 2. Model aliases for the plots, together with model sizes.
Table 2. Model aliases for the plots, together with model sizes.
ModelAliasModel Size
Fine-tuned SAM vit-hf_h2.4 GB
Fine-tuned SAM vit-lf_l1.2 GB
Fine-tuned SAM vit-bf_b358 MB
Fine-tuned MedSAMf_Md358 MB
Fine-tuned MobileSAMf_Mb39 MB
Original MedSAMMdS358 MB
Original MobileSAMMbS39 MB
Original SAM vit-hSAM2.4 GB
Trained nnU-NetnnU1.6 GB
Table 3. Training times per epoch on all devices, FLOPS, and number of parameters for finetuned MobileSAM and nnU-Net.
Table 3. Training times per epoch on all devices, FLOPS, and number of parameters for finetuned MobileSAM and nnU-Net.
ModelGPU (s)Server (s)Desktop (s)Laptop (s)FLOPS (G)Parameters (M)
Finetuned MobileSAM2.1612.3039.0466.3040.9410.13
nnU-Net92.6911,173.00130.68140.02
Table 4. Model ranking on low-cost device deployment constraints.
Table 4. Model ranking on low-cost device deployment constraints.
ConstraintMobileSAMnnU-NetRemarks
Accuracy21Both model accuracies are acceptable; only slight manual corrections needed when segmentation is not fully successful
Memory12
Task expansion12
Real-time inference12
Data localisationYesYes
InteractiveYesNoMobileSAM supports prompting to manually correct segmentations when needed
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Nnadozie, E.C.; Merino-Caviedes, S.; de Luis-Román, D.A.; Martín-Fernández, M.; Alberola-López, C. Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings. Big Data Cogn. Comput. 2026, 10, 142. https://doi.org/10.3390/bdcc10050142

AMA Style

Nnadozie EC, Merino-Caviedes S, de Luis-Román DA, Martín-Fernández M, Alberola-López C. Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings. Big Data and Cognitive Computing. 2026; 10(5):142. https://doi.org/10.3390/bdcc10050142

Chicago/Turabian Style

Nnadozie, Emmanuel Chibuikem, Susana Merino-Caviedes, Daniel A. de Luis-Román, Marcos Martín-Fernández, and Carlos Alberola-López. 2026. "Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings" Big Data and Cognitive Computing 10, no. 5: 142. https://doi.org/10.3390/bdcc10050142

APA Style

Nnadozie, E. C., Merino-Caviedes, S., de Luis-Román, D. A., Martín-Fernández, M., & Alberola-López, C. (2026). Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings. Big Data and Cognitive Computing, 10(5), 142. https://doi.org/10.3390/bdcc10050142

Article Metrics

Back to TopTop