Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (82)

Search Parameters:
Keywords = VGG16-UNet segmentation model

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 3179 KB  
Article
Forest Road Extraction from High-Resolution Remote Sensing Imagery Based on an Improved U-Net Model
by Hongrong Wang, Haoquan Chen, Feifan Yang, Dong Chen, Zhiqiang Min, Hanmin Sheng, Kai Chen, Hang Geng, Chengde Wang and Fang Song
Sustainability 2026, 18(17), 9029; https://doi.org/10.3390/su18179029 - 2 Sep 2026
Viewed by 353
Abstract
To address the challenges of vegetation interference, background confusion, and road fragmentation caused by the narrow and elongated structures of forest roads in complex remote sensing imagery, this study proposed an improved U-Net-based model, namely HAA-UNet, for automatic forest road extraction. The proposed [...] Read more.
To address the challenges of vegetation interference, background confusion, and road fragmentation caused by the narrow and elongated structures of forest roads in complex remote sensing imagery, this study proposed an improved U-Net-based model, namely HAA-UNet, for automatic forest road extraction. The proposed model integrates a VGG16 encoder, an Atrous Spatial Pyramid Pooling (ASPP) module, and a Hybrid Dilated Convolution (HDCC) module to enhance feature representation and spatial detail reconstruction of road targets under complex forest environments. Experimental results demonstrated that HAA-UNet achieved superior segmentation performance on the forest road dataset of Xichang City, Sichuan Province, with Precision, Recall, F1-score, and mIoU values of 87.24%, 87.83%, 87.53%, and 79.85%, respectively, outperforming all comparison models. These results indicate that the proposed method effectively improves road continuity and boundary delineation in complex forest scenes. Furthermore, the extracted road information was integrated into the Forest Fire Risk Index (FFRI) assessment framework, demonstrating that accurate road data can improve the spatial characterization of fire risk and provide reliable data support for forest fire risk assessment and forest resource management. Full article
Show Figures

Figure 1

22 pages, 3041 KB  
Article
A Two-Stage Method for Detecting and Assessing the Severity of Diseases and Pests on Lotus Leaves in Complex Aquatic Environments
by Yifei Miao, Zhiqi Cai, Siqiao Tan, Bo Li, Dazhi Liu and Donghui Li
Agronomy 2026, 16(17), 1636; https://doi.org/10.3390/agronomy16171636 - 27 Aug 2026
Viewed by 296
Abstract
Addressing the challenges posed by the small scale, diverse morphology, and severe occlusion of pest and disease targets on lotus leaves in complex aquatic environments—as well as the difficulty of existing methods in automatically quantifying disease severity—this paper proposes an approach for the [...] Read more.
Addressing the challenges posed by the small scale, diverse morphology, and severe occlusion of pest and disease targets on lotus leaves in complex aquatic environments—as well as the difficulty of existing methods in automatically quantifying disease severity—this paper proposes an approach for the identification of lotus leaf pests and diseases and quantitative grading of leaf spot disease severity. First, by integrating high-altitude canopy imagery captured by unmanned aerial vehicles (UAVs) with high-definition ground-level data, a multi-perspective dataset comprising object detection bounding box annotations and pixel-level segmentation annotations is constructed. Second, the YOLOv12n-DFFN object detection model is proposed; this model enhances interaction between deep and shallow features through a dynamic feature feedback mechanism, thereby improving the ability to localize disease targets against complex aquatic backgrounds. Finally, taking typical leaf spot disease as the subject, YOLOv12n-DFFN is used to detect and extract diseased leaf regions, while a VGG-UNet semantic segmentation model is employed to achieve pixel-level fine-grained segmentation of leaf and lesion areas. By calculating the ratio of lesion area to total leaf area, automatic quantitative grading of disease severity is realized. Experimental results show that YOLOv12n-DFFN achieved a precision of 93.58%, representing an improvement of 4.37 percentage points over the YOLOv12n baseline, with an mAP50 of 83.87%; the segmentation model attained an average intersection-over-union of 90.47%, and the overall accuracy of the two-stage framework for leaf spot disease severity grading reached 96.0%. Through a “detection first, segmentation second” two-stage strategy, this framework enables the intelligent identification and severity quantification of lotus leaf diseases in complex aquatic environments, providing an effective approach for the intelligent monitoring and precision management of aquatic crop diseases. Full article
(This article belongs to the Section Pest and Disease Management)
Show Figures

Figure 1

32 pages, 75104 KB  
Article
A Feature-Optimized Deep Learning Framework for Mapping and Spatial Characterization of Tea Plantations in Complex Mountain Landscapes
by Ruyi Wang, Jixian Zhang, Xiaoping Lu, Qi Kang, Bowen Chi, Junfeng Li, Yahang Li and Zhengfang Lou
Remote Sens. 2026, 18(9), 1281; https://doi.org/10.3390/rs18091281 - 23 Apr 2026
Viewed by 495
Abstract
The unchecked expansion of tea plantations onto steep, forest-adjacent slopes in subtropical mountains engenders a conflict between agricultural productivity and ecosystem integrity, particularly by exacerbating habitat fragmentation and soil erosion. While precise monitoring is essential to navigate this trade-off for sustainable management, accurate [...] Read more.
The unchecked expansion of tea plantations onto steep, forest-adjacent slopes in subtropical mountains engenders a conflict between agricultural productivity and ecosystem integrity, particularly by exacerbating habitat fragmentation and soil erosion. While precise monitoring is essential to navigate this trade-off for sustainable management, accurate inventorying remains a challenge due to the plantations’ strong phenological variability, heterogeneous canopy structures, and high spectral confusion with surrounding vegetation. This study proposes a feature-optimized deep learning framework for mapping and characterizing tea plantations in complex landscapes, using Xinyang City, China, as a study area. The framework integrates multi-temporal Sentinel-1/2 observations with a sequential Jeffries-Matusita (JM)-Pearson feature filtering strategy. This approach effectively condenses a 132-variable high-dimensional pool (including optical spectra, vegetation indices, textures, and SAR polarimetry) into a compact 28-feature subset (a 78.8% reduction), preserving critical phenological and structural cues while minimizing redundancy. These optimized predictors drive a hybrid VGG16–UNet++ segmentation network, which couples transfer-learning-based semantic encoding with detail-preserving dense skip fusion. Extensive experiments across 18 model–feature configurations demonstrate that the optimal setting achieves an Overall Accuracy of 97.82%, an F1-score of 0.9093, and a mean IoU of 0.7968. Notably, the method significantly reduces misclassification in rugged, cloud-prone terrain, yielding a User’s Accuracy of 91.14% for tea. Based on the generated wall-to-wall map, we derived two decision-support indicators: multi-threshold steep-slope exposure and a normalized tea–forest interface density. This framework provides actionable, high-precision spatial products to support slope-based zoning, ecological restoration, and sustainable management in fragile mountain agroforestry systems. Full article
Show Figures

Figure 1

35 pages, 5739 KB  
Article
Multi-Scale Atrous Feature Fusion Based on a VGG19-UNet Encoder for Brain Tumor Segmentation
by Shoffan Saifullah and Rafał Dreżewski
Appl. Sci. 2026, 16(8), 3971; https://doi.org/10.3390/app16083971 - 19 Apr 2026
Viewed by 587
Abstract
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) remains challenging due to heterogeneous tumor morphology, intensity variability, and multi-scale structural complexity. This study proposes a DeepLabV3+-based segmentation framework integrating a VGG19-UNet encoder, Atrous Spatial Pyramid Pooling (ASPP), and low-level feature refinement to [...] Read more.
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) remains challenging due to heterogeneous tumor morphology, intensity variability, and multi-scale structural complexity. This study proposes a DeepLabV3+-based segmentation framework integrating a VGG19-UNet encoder, Atrous Spatial Pyramid Pooling (ASPP), and low-level feature refinement to simultaneously capture hierarchical semantics and boundary-sensitive spatial details. The architecture enhances receptive field coverage without additional downsampling while preserving fine-grained contour information during reconstruction. Extensive evaluation was conducted on the Figshare Brain Tumor Segmentation (FBTS) dataset and the BraTS 2021 and BraTS 2018 benchmarks, focusing on Whole Tumor segmentation across multiple MRI modalities and tumor grades. Under five-fold cross-validation, the proposed model achieved a mean Dice Similarity Coefficient of 0.9717 and Jaccard Index of 0.9456 on FBTS, with stable and competitive performance across FLAIR, T1, T2, and T1CE modalities in both HGG and LGG cases. Boundary-level analysis further confirmed controlled Hausdorff Distance and low Average Symmetric Surface Distance. Statistical validation and ablation analysis demonstrate consistent improvements over baseline U-Net configurations. The proposed framework provides a robust and computationally efficient solution for automated brain tumor segmentation across heterogeneous datasets. Full article
(This article belongs to the Special Issue Research on Artificial Intelligence in Healthcare)
Show Figures

Figure 1

26 pages, 4917 KB  
Article
A Comprehensive Clinical Decision Support System for the Early Diagnosis of Axial Spondyloarthritis: Multi-Sequence MRI, Clinical Risk Integration, and Explainable Segmentation
by Fatih Tarakci, Ilker Ali Ozkan, Musa Dogan, Halil Ozer, Dilek Tezcan and Sema Yilmaz
Diagnostics 2026, 16(7), 1037; https://doi.org/10.3390/diagnostics16071037 - 30 Mar 2026
Viewed by 978
Abstract
Background/Objectives: This study aims to develop a comprehensive Clinical Decision Support System (CDSS) that integrates multi-sequence sacroiliac joint (SIJ) MRIs with rheumatological, clinical, and laboratory findings into the decision-making process for the early diagnosis of axial spondyloarthritis (axSpA), incorporating segmentation-supported explainability. Methods: Multi-sequence [...] Read more.
Background/Objectives: This study aims to develop a comprehensive Clinical Decision Support System (CDSS) that integrates multi-sequence sacroiliac joint (SIJ) MRIs with rheumatological, clinical, and laboratory findings into the decision-making process for the early diagnosis of axial spondyloarthritis (axSpA), incorporating segmentation-supported explainability. Methods: Multi-sequence SIJ MRI data (T1-WI, T2-WI, STIR, and PD-WI) were analysed from 367 participants (n = 193 axSpA; n = 174 non-axSpA controls). Sequence-based classification was performed using VGG16, ResNet50, DenseNet121, and InceptionV3 models; additionally, a lightweight and parameter-efficient SacroNet architecture was developed. Slice-level probability scores were converted to patient-level scores using the Dynamic Top-K Averaging method. Image-based scores were combined with a logistic regression-based clinical risk score using weighted linear integration (0.60 image/0.40 clinical) and a conservative threshold (τ = 0.70). Grad-CAM was applied for visual interpretability. Furthermore, to support the diagnostic outcomes with precise spatial data, active inflammation in STIR and T2-WI sequences was segmented. For this purpose, the MDC-UNet model was employed and compared with baseline U-Net derivatives. Results: Sequence-specific analysis showed VGG16 performing best on T1-WI (AUC = 0.920; Accuracy = 0.878) and DenseNet121 on STIR (AUC = 0.793; Accuracy = 0.771). The SacroNet architecture provided competitive classification performance at the patient level despite its low number of parameters (~110 K). Furthermore, MDC-UNet successfully segmented active inflammation, yielding Dice scores of 0.752 (HD95: 19.25) for STIR and 0.682 (HD95: 26.21) for T2-WI. Conclusions: The findings demonstrate that patient-level decision integration based on multi-sequence MRI, when used in conjunction with clinical risk scoring and segmentation-assisted interpretability, can provide a feasible and interpretable DSS framework for the early diagnosis of axSpA. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Figure 1

22 pages, 8609 KB  
Article
Integrating SimAM Attention and S-DRU Feature Reconstruction for Sentinel-2 Imagery-Based Soybean Planting Area Extraction
by Haotong Wu, Xinwen Wan, Rong Qian, Chao Ruan, Jinling Zhao and Chuanjian Wang
Agriculture 2026, 16(6), 693; https://doi.org/10.3390/agriculture16060693 - 19 Mar 2026
Cited by 2 | Viewed by 606
Abstract
Accurate and stable acquisition of the spatial distribution of soybean planting areas is essential for supporting precision agricultural monitoring and ensuring food security. However, crop remote-sensing mapping for specific regions still faces critical data bottlenecks: high-precision, large-scale pixel-level annotation is costly, resulting in [...] Read more.
Accurate and stable acquisition of the spatial distribution of soybean planting areas is essential for supporting precision agricultural monitoring and ensuring food security. However, crop remote-sensing mapping for specific regions still faces critical data bottlenecks: high-precision, large-scale pixel-level annotation is costly, resulting in scarce available labeled samples that make it difficult to construct large-scale training datasets. Although parameter-intensive models such as FCN and SegNet can achieve sufficient end-to-end training on large-scale public remote sensing datasets like LoveDA, when directly applied to the data-limited dataset in this study area, the models are prone to overfitting, leading to a significant decline in generalization ability. To address these issues, this study proposes a lightweight U-shaped semantic segmentation model, SimSDRU-Net. The model utilizes a pre-trained VGG-16 backbone to extract shallow texture and deep semantic features. The pre-trained weights mitigate the impact of overfitting in data-limited settings. In the decoding stage, a parameter-free lightweight SimAM attention module enhances effective soybean features and suppresses soil background redundancy, while an embedded S-DRU unit fuses multi-scale features for deep complementary reconstruction to improve edge detail capture. A label dataset was constructed using Sentinel-2 images as the data source and Menard County (USA) as the study area. The USDA CDL was used as a foundation for the dataset, with Google high-resolution images serving as visual interpretation aids. In the context of the experiment, Deeplabv3+ and U-Net++ were compared with U-Net under identical conditions. The results demonstrated that SimSDRU-Net exhibited optimal performance, with MIoU of 89.03%, MPA of 93.81%, and OA of 95.96%. Specifically, SimSDRU-Net uses the SimAM attention module to generate spatial attention weights by analyzing feature statistical differences through an energy function, so as to adaptively enhance soybean texture features. Meanwhile, the S-DRU unit groups, dynamically weights, and cross-branch reconstructs multi-scale convolutional features to preserve fine boundary details and achieve accurate segmentation of soybean plots. The present study demonstrates that SimSDRU-Net integrates lightweight design and high precision in data-limited scenarios, thereby providing effective technical support for the rapid extraction of soybean planting areas in North America. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

32 pages, 14017 KB  
Article
Optimized Image Segmentation Model for Pellet Microstructure Incorporating KL Divergence Constraints
by Yuwen Ai, Xia Li, Aimin Yang, Yunjie Bai and Xuezhi Wu
Mathematics 2026, 14(3), 574; https://doi.org/10.3390/math14030574 - 5 Feb 2026
Viewed by 627
Abstract
Accurate segmentation of pellet microstructure images is crucial for evaluating their metallurgical performance and optimizing production processes. To address the challenges posed by complex structures, blurred boundaries, and fine-grained textures of hematite and magnetite in pellet micrographs, this study proposes a hybrid intelligently [...] Read more.
Accurate segmentation of pellet microstructure images is crucial for evaluating their metallurgical performance and optimizing production processes. To address the challenges posed by complex structures, blurred boundaries, and fine-grained textures of hematite and magnetite in pellet micrographs, this study proposes a hybrid intelligently optimized VGG16-U-Net semantic segmentation model. The model incorporates an improved SPC-SA channel self-attention mechanism in the encoder to enhance deep feature representation, while a simplified SAN and SAW module is integrated into the decoder to strengthen its response to key mineral regions. Additionally, a hybrid loss strategy is employed with KL regularization for training optimization. Experimental results show that the model achieves an mIoU of 85.58%, an mPA of 91.54%, and an overall accuracy of 93.58%. Compared with the baseline models, the proposed method achieves improved performance to some extent. Full article
(This article belongs to the Special Issue Mathematical Methods for Image Processing and Computer Vision)
Show Figures

Figure 1

18 pages, 14590 KB  
Article
VTC-Net: A Semantic Segmentation Network for Ore Particles Integrating Transformer and Convolutional Block Attention Module (CBAM)
by Yijing Wu, Weinong Liang, Jiandong Fang, Chunxia Zhou and Xiaolu Sun
Sensors 2026, 26(3), 787; https://doi.org/10.3390/s26030787 - 24 Jan 2026
Cited by 4 | Viewed by 725
Abstract
In mineral processing, visual-based online particle size analysis systems depend on high-precision image segmentation to accurately quantify ore particle size distribution, thereby optimizing crushing and sorting operations. However, due to multi-scale variations, severe adhesion, and occlusion within ore particle clusters, existing segmentation models [...] Read more.
In mineral processing, visual-based online particle size analysis systems depend on high-precision image segmentation to accurately quantify ore particle size distribution, thereby optimizing crushing and sorting operations. However, due to multi-scale variations, severe adhesion, and occlusion within ore particle clusters, existing segmentation models often exhibit undersegmentation and misclassification, leading to blurred boundaries and limited generalization. To address these challenges, this paper proposes a novel semantic segmentation model named VTC-Net. The model employs VGG16 as the backbone encoder, integrates Transformer modules in deeper layers to capture global contextual dependencies, and incorporates a Convolutional Block Attention Module (CBAM) at the fourth stage to enhance focus on critical regions such as adhesion edges. BatchNorm layers are used to stabilize training. Experiments on ore image datasets show that VTC-Net outperforms mainstream models such as UNet and DeepLabV3 in key metrics, including MIoU (89.90%) and pixel accuracy (96.80%). Ablation studies confirm the effectiveness and complementary role of each module. Visual analysis further demonstrates that the model identifies ore contours and adhesion areas more accurately, significantly improving segmentation robustness and precision under complex operational conditions. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

22 pages, 5552 KB  
Article
MSA-UNet: Multiscale Feature Aggregation with Attentive Skip Connections for Precise Building Extraction
by Guobiao Yao, Yan Chen, Wenxiao Sun, Zeyu Zhang, Yifei Tang and Jingxue Bi
ISPRS Int. J. Geo-Inf. 2025, 14(12), 497; https://doi.org/10.3390/ijgi14120497 - 17 Dec 2025
Cited by 1 | Viewed by 1127
Abstract
An accurate and reliable extraction of building structures from high-resolution (HR) remote sensing images is an important research topic in 3D cartography and smart city construction. However, despite the strong overall performance of recent deep learning models, limitations remain in handling significant variations [...] Read more.
An accurate and reliable extraction of building structures from high-resolution (HR) remote sensing images is an important research topic in 3D cartography and smart city construction. However, despite the strong overall performance of recent deep learning models, limitations remain in handling significant variations in building scales and complex architectural forms, which may lead to inaccurate boundaries or difficulties in extracting small or irregular structures. Therefore, the present study proposes MSA-UNet, a reliable semantic segmentation framework that leverages multiscale feature aggregation and attentive skip connections for an accurate extraction of building footprints. This framework is constructed based on the U-Net architecture, incorporating VGG16 as a replacement for the original encoder structure, which enhances its ability to capture low-discriminative features. To further improve the representation of image buildings with different scales and shapes, a serial coarse-to-fine feature aggregation mechanism was used. Additionally, a novel skip connection was built between the encoder and decoder layers to enable adaptive weights. Furthermore, a dual-attention mechanism, implemented through the convolutional block attention module, was integrated to enhance the focus of the network on building regions. Extensive experiments conducted on the WHU and Inria building datasets validated the effectiveness of MSA-UNet. On the WHU dataset, the model demonstrated a state-of-the-art performance with a mean Intersection over Union (mIoU) of 94.26%, accuracy of 98.32%, F1-score of 96.57%, and mean Pixel accuracy (mPA) of 96.85%, corresponding to gains of 1.41% in mIoU over the baseline U-Net. On the more challenging Inria dataset, MSA-UNet achieved an mIoU of 85.92%, indicating a consistent improvement of up to 1.9% over the baseline U-Net. These results confirmed that MSA-UNet markedly improved the accuracy and boundary integrity of building extraction from HR data, outperforming existing classic models in terms of segmentation quality and robustness. Full article
(This article belongs to the Special Issue Spatial Data Science and Knowledge Discovery)
Show Figures

Figure 1

20 pages, 4879 KB  
Article
A Multi-Phenotype Acquisition System for Pleurotus eryngii Based on RGB and Depth Imaging
by Yueyue Cai, Zhijun Wang, Ziqin Liao, Yujie Li, Weijie Shi, Peijie Huang, Bingzhi Chen, Jie Pang, Xiangzeng Kong and Xuan Wei
Agriculture 2025, 15(24), 2566; https://doi.org/10.3390/agriculture15242566 - 11 Dec 2025
Cited by 1 | Viewed by 830
Abstract
High-throughput phenotypic acquisition and analysis allow us to accurately quantify trait expressions, which is essential for developing intelligent breeding strategies. However, there is still much potential to explore in the field of high-throughput phenotyping for edible fungi. In this study, we developed a [...] Read more.
High-throughput phenotypic acquisition and analysis allow us to accurately quantify trait expressions, which is essential for developing intelligent breeding strategies. However, there is still much potential to explore in the field of high-throughput phenotyping for edible fungi. In this study, we developed a portable multi-phenotypic acquisition system for Pleurotus eryngii using RGB and RGB-D cameras. We developed an innovative Unet-based semantic segmentation model by integrating the ASPP structure with the VGG16 architecture. This allows for precise segmentation of the cap, gills and stem of the fruiting body. By leveraging depth images from RGB-D cameras, we can effectively collect phenotypic information about Pleurotus eryngii. By combining K-means clustering with Lab color space thresholds, we are able to achieve more precise automatic classification of Pleurotus eryngii cap colors. Moreover, AlexNet is utilized to classify the shapes of the fruiting bodies. The Aspp-VGGUnet network demonstrates remarkable performance with a mean Intersection over Union (mIoU) of 96.47% and a mean pixel accuracy (mPA) of 98.53%. These results reflect respective improvements of 3.03% and 2.23% compared to the standard Unet model, respectively. The average error in size phenotype measurement is just 0.15 ± 0.03 cm. The accuracy for cap color classification reaches 91.04%, while fruiting body shape classification achieves 97.90%. The proposed multi-phenotype acquisition system reduces the measurement time per sample from an average of 76 s (manual method) to about 2 s, substantially increasing data acquisition throughput and providing robust support for scalable phenotyping workflows in breeding research. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Graphical abstract

25 pages, 7527 KB  
Article
A Multifocal RSSeg Approach for Skeletal Age Estimation in an Indian Medicolegal Perspective
by Priyanka Manchegowda, Manohar Nageshmurthy, Suresha Raju and Dayananda Rudrappa
Algorithms 2025, 18(12), 765; https://doi.org/10.3390/a18120765 - 4 Dec 2025
Cited by 1 | Viewed by 1012
Abstract
Estimating bone age is essential for accurate diagnoses, appropriate care based on biological age, and fairness in legal matters. In the Indian medicolegal context, determining age through a clinical approach involves analyzing multiple joints; however, the traditional method can be tedious and subjective, [...] Read more.
Estimating bone age is essential for accurate diagnoses, appropriate care based on biological age, and fairness in legal matters. In the Indian medicolegal context, determining age through a clinical approach involves analyzing multiple joints; however, the traditional method can be tedious and subjective, relying heavily on human expertise, which may lead to biased decisions in age-related legal disputes. Moreover, commonly used radiographs often exhibit pixel-level variations due to heterogeneous contrast, which complicate segmentation tasks and lead to inconsistencies and reduced model performance. The study presents a multifocal region-based symbolic segmentation technique to automatically retain the soft-tissue region that harbors a growth pattern of an ossification center. Experimental results demonstrate an 84.5% Jaccard similarity, an 81.4% Dice coefficient, an 88.3% precision, a 90.0% recall, and a 91.5% pixel accuracy on a novel multifocal dataset of Indian inhabitants. The proposed segmentation technique outperforms U-Net, Attention U-Net, TransU-Net, DeepLabV3+, Adaptive Otsu, and Watershed segmentation in terms of accuracy, indicating strong generalizability across joints and improving reliability. Compared with 86.4% without segmentation, the proposed integration of segmentation with VGG16 classification increases the overall accuracy to 93.8%, demonstrating that target-focused-region processing reduces unnecessary computations and improves feature discrimination without sacrificing accuracy. Full article
(This article belongs to the Special Issue Machine Learning in Medical Signal and Image Processing (4th Edition))
Show Figures

Figure 1

21 pages, 8098 KB  
Article
Multi-Sensor AI-Based Urban Tree Crown Segmentation from High-Resolution Satellite Imagery for Smart Environmental Monitoring
by Amirmohammad Sharifi, Reza Shah-Hosseini, Danesh Shokri and Saeid Homayouni
Smart Cities 2025, 8(6), 187; https://doi.org/10.3390/smartcities8060187 - 6 Nov 2025
Cited by 2 | Viewed by 2628
Abstract
Urban tree detection is fundamental to effective forestry management, biodiversity preservation, and environmental monitoring—key components of sustainable smart city development. This study introduces a deep learning framework for urban tree crown segmentation that exclusively leverages high-resolution satellite imagery from GeoEye-1, WorldView-2, and WorldView-3, [...] Read more.
Urban tree detection is fundamental to effective forestry management, biodiversity preservation, and environmental monitoring—key components of sustainable smart city development. This study introduces a deep learning framework for urban tree crown segmentation that exclusively leverages high-resolution satellite imagery from GeoEye-1, WorldView-2, and WorldView-3, thereby eliminating the need for additional data sources such as LiDAR or UAV imagery. The proposed framework employs a Residual U-Net architecture augmented with Attention Gates (AGs) to address major challenges, including class imbalance, overlapping crowns, and spectral interference from complex urban structures, using a custom composite loss function. The main contribution of this work is to integrate data from three distinct satellite sensors with varying spatial and spectral characteristics into a single processing pipeline, demonstrating that such well-established architectures can yield reliable, high-accuracy results across heterogeneous resolutions and imaging conditions. A further advancement of this study is the development of a hybrid ground-truth generation strategy that integrates NDVI-based watershed segmentation, manual annotation, and the Segment Anything Model (SAM), thereby reducing annotation effort while enhancing mask fidelity. In addition, by training on 4-band RGBN imagery from multiple satellite sensors, the model exhibits generalization capabilities across diverse urban environments. Despite being trained on a relatively small dataset comprising only 1200 image patches, the framework achieves state-of-the-art performance (F1-score: 0.9121; IoU: 0.8384; precision: 0.9321; recall: 0.8930). These results stem from the integration of the Residual U-Net with Attention Gates, which enhance feature representation and suppress noise from urban backgrounds, as well as from hybrid ground-truth generation and the combined BCE–Dice loss function, which effectively mitigates class imbalance. Collectively, these design choices enable robust model generalization and clear performance superiority over baseline networks such as DeepLab v3 and U-Net with VGG19. Fully automated and computationally efficient, the proposed approach delivers cost-effective, accurate segmentation using satellite data alone, rendering it particularly suitable for scalable, operational smart city applications and environmental monitoring initiatives. Full article
Show Figures

Figure 1

16 pages, 2692 KB  
Article
Improved UNet-Based Detection of 3D Cotton Cup Indentations and Analysis of Automatic Cutting Accuracy
by Lin Liu, Xizhao Li, Hongze Lv, Jianhuang Wang, Fucai Lai, Fangwei Zhao and Xibing Li
Processes 2025, 13(10), 3144; https://doi.org/10.3390/pr13103144 - 30 Sep 2025
Viewed by 704
Abstract
With the advancement of intelligent technology and the rise in labor costs, manual identification and cutting of 3D cotton cup indentations can no longer meet modern demands. The increasing variety and shape of 3D cotton cups due to personalized requirements make the use [...] Read more.
With the advancement of intelligent technology and the rise in labor costs, manual identification and cutting of 3D cotton cup indentations can no longer meet modern demands. The increasing variety and shape of 3D cotton cups due to personalized requirements make the use of fixed molds for cutting inefficient, leading to a large number of molds and high costs. Therefore, this paper proposes a UNet-based indentation segmentation algorithm to automatically extract 3D cotton cup indentation data. By incorporating the VGG16 network and Leaky-ReLU activation function into the UNet model, the method improves the model’s generalization capability, convergence speed, detection speed, and reduces the risk of overfitting. Additionally, attention mechanisms and an Atrous Spatial Pyramid Pooling (ASPP) module are introduced to enhance feature extraction, improving the network’s spatial feature extraction ability. Experiments conducted on a self-made 3D cotton cup dataset demonstrate a precision of 99.53%, a recall of 99.69%, a mIoU of 99.18%, and an mPA of 99.73%, meeting practical application requirements. The extracted 3D cotton cup indentation contour data is automatically input into an intelligent CNC cutting machine to cut 3D cotton cup. The cutting results of 400 data points show an 0.20 mm ± 0.42 mm error, meeting the cutting accuracy requirements for flexible material 3D cotton cups. This study may serve as a reference for machine vision, image segmentation, improvements to deep learning architectures, and automated cutting machinery for flexible materials such as fabrics. Full article
(This article belongs to the Section Automation Control Systems)
Show Figures

Figure 1

23 pages, 347 KB  
Article
Comparative Analysis of Foundational, Advanced, and Traditional Deep Learning Models for Hyperpolarized Gas MRI Lung Segmentation: Robust Performance in Data-Constrained Scenarios
by Ramtin Babaeipour, Matthew S. Fox, Grace Parraga and Alexei Ouriadov
Bioengineering 2025, 12(10), 1062; https://doi.org/10.3390/bioengineering12101062 - 30 Sep 2025
Cited by 2 | Viewed by 1486
Abstract
This study investigates the comparative performance of foundational models, advanced large-kernel architectures, and traditional deep learning approaches for hyperpolarized gas MRI segmentation across progressive data reduction scenarios. Chronic obstructive pulmonary disease (COPD) remains a leading global health concern, and advanced imaging techniques are [...] Read more.
This study investigates the comparative performance of foundational models, advanced large-kernel architectures, and traditional deep learning approaches for hyperpolarized gas MRI segmentation across progressive data reduction scenarios. Chronic obstructive pulmonary disease (COPD) remains a leading global health concern, and advanced imaging techniques are crucial for its diagnosis and management. Hyperpolarized gas MRI, utilizing helium-3 (3He) and xenon-129 (129Xe), offers a non-invasive way to assess lung function. We evaluated foundational models (Segment Anything Model and MedSAM), advanced architectures (UniRepLKNet and TransXNet), and traditional deep learning models (UNet with VGG19 backbone, Feature Pyramid Network with MIT-B5 backbone, and DeepLabV3 with ResNet152 backbone) using four data availability scenarios: 100%, 50%, 25%, and 10% of the full training dataset (1640 2D MRI slices from 205 participants). The results demonstrate that foundational and advanced models achieve statistically equivalent performance across all data scenarios (p > 0.01), while both significantly outperform traditional architectures under data constraints (p < 0.001). Under extreme data scarcity (10% training data), foundational and advanced models maintained DSC values above 0.86, while traditional models experienced catastrophic performance collapse. This work highlights the critical advantage of architectures with large effective receptive fields in medical imaging applications where data collection is challenging, demonstrating their potential to democratize advanced medical imaging analysis in resource-limited settings. Full article
(This article belongs to the Special Issue Artificial Intelligence-Based Medical Imaging Processing)
Show Figures

Figure 1

18 pages, 1597 KB  
Article
A Comparative Analysis of SegFormer, FabE-Net and VGG-UNet Models for the Segmentation of Neural Structures on Histological Sections
by Igor Makarov, Elena Koshevaya, Alina Pechenina, Galina Boyko, Anna Starshinova, Dmitry Kudlay, Taiana Makarova and Lubov Mitrofanova
Diagnostics 2025, 15(18), 2408; https://doi.org/10.3390/diagnostics15182408 - 22 Sep 2025
Cited by 3 | Viewed by 2296
Abstract
Background: Segmenting nerve fibres in histological images is a tricky job because of how much the tissue looks can change. Modern neural network architectures, including U-Net and transformers, demonstrate varying degrees of effectiveness in this area. The aim of this study is to [...] Read more.
Background: Segmenting nerve fibres in histological images is a tricky job because of how much the tissue looks can change. Modern neural network architectures, including U-Net and transformers, demonstrate varying degrees of effectiveness in this area. The aim of this study is to conduct a comparative analysis of the SegFormer, VGG-UNet, and FabE-Net models in terms of segmentation quality and speed. Methods: The training sample consisted of more than 75,000 pairs of images of different tissues (original slice and corresponding mask), scaled from 1024 × 1024 to 224 × 224 pixels to optimise computations. Three neural network architectures were used: the classic VGG-UNet, FabE-Net with attention and global context perception blocks, and the SegFormer transformer model. For an objective assessment of the quality of the models, expert validation was carried out with the participation of four independent pathologists, who evaluated the quality of segmentation according to specified criteria. Quality metrics (precision, recall, F1-score, accuracy) were calculated as averages based on the assessments of all experts, which made it possible to take into account variability in interpretation and increase the reliability of the results. Results: SegFormer achieved stable stabilisation of the loss function faster than the other models—by the 20–30th epoch, compared to 45–60 epochs for VGG-UNet and FabE-Net. Despite taking longer to train per epoch, SegFormer produced the best segmentation quality, with the following metrics: precision 0.84, recall 0.99, F1-score 0.91 and accuracy 0.89. It also annotated a complete histological section in the fastest time. Visual analysis revealed that, compared to other models, which tended to produce incomplete or excessive segmentation, SegFormer more accurately and completely highlights nerve structures. Conclusions: Using attention mechanisms in SegFormer compensates for morphological variability in tissues, resulting in faster and higher-quality segmentation. Image scaling does not impair training quality while significantly accelerating computational processes. These results confirm the potential of SegFormer for practical use in digital pathology, while also highlighting the need for high-precision, immunohistochemistry-informed labelling to improve segmentation accuracy. Full article
(This article belongs to the Special Issue Pathology and Diagnosis of Neurological Disorders, 2nd Edition)
Show Figures

Figure 1

Back to TopTop