Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (273)

Search Parameters:
Keywords = sparse and dense features

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 1661 KB  
Article
DANet: Joint Density- and Semantics-Adaptive Convolution for 3D Point-Cloud Semantic Segmentation
by Weijian Hu, Shuning Wang, Lingfang Li, Jikai Zhang and Ke Han
Sensors 2026, 26(14), 4561; https://doi.org/10.3390/s26144561 - 18 Jul 2026
Viewed by 208
Abstract
Semantic segmentation of 3D point clouds remains difficult when LiDAR or depth-camera data are sampled unevenly. This paper presents DANet, a 3D semantic segmentation framework built on joint density- and semantics-adaptive convolution. Its core operator, Density-Adaptive Radius Convolution (DAR-Conv), predicts point-wise neighborhood radii [...] Read more.
Semantic segmentation of 3D point clouds remains difficult when LiDAR or depth-camera data are sampled unevenly. This paper presents DANet, a 3D semantic segmentation framework built on joint density- and semantics-adaptive convolution. Its core operator, Density-Adaptive Radius Convolution (DAR-Conv), predicts point-wise neighborhood radii before feature aggregation by combining density-driven initialization with semantics-aware modulation. In this way, dense regions can use compact receptive fields, whereas sparse or semantically complex regions can draw on broader contextual support. DANet also includes a Gated Adaptive Cross-Layer Fusion (GACF) module, which aligns encoder–decoder features and performs gated fusion with residual refinement. Experiments on S3DIS and NPM3D show that DANet obtains the highest reported mean accuracy (mAcc) among the compared methods on S3DIS, and high mean Intersection over Union (mIoU) and overall accuracy (OA) on NPM3D, supporting the usefulness of density- and semantics-aware receptive-field adaptation. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

31 pages, 81549 KB  
Article
Carbohydrate Availability Modulates Biomass, Architecture, and Viability of Mono- and Polymicrobial ESKAPE Biofilms
by Majd M. Alsaleh, Dana A. Alqudah, Bassam I. El-Eswed and Mamoon M. D. Al-Rshaidat
Appl. Biosci. 2026, 5(3), 60; https://doi.org/10.3390/applbiosci5030060 - 14 Jul 2026
Viewed by 132
Abstract
Biofilm formation of ESKAPE pathogens creates a major problem in healthcare-associated infections, and nutrient supply significantly contributes to biofilm structure and bacterial survival. An understanding of how various nutritional environments influence polymicrobial biofilm dynamics is important for the development of targeted antimicrobials. In [...] Read more.
Biofilm formation of ESKAPE pathogens creates a major problem in healthcare-associated infections, and nutrient supply significantly contributes to biofilm structure and bacterial survival. An understanding of how various nutritional environments influence polymicrobial biofilm dynamics is important for the development of targeted antimicrobials. In the present work, biofilm formation of ESKAPE pathogens was tested under three different nutritional environments—nutrient broth with no supplementation (NB; minimal), nutrient broth with supplementation of 2% glucose + 2% sucrose (NB+G+S; moderate), and Brain Heart Infusion supplemented with 2% glucose + 2% sucrose (BHI-G+S; Rich). Furthermore, we studied the effect of bacterial mixtures on biofilm features in each medium. Standard ATCC strains of Staphylococcus aureus (S), Enterococcus faecalis (E), Klebsiella pneumoniae (K), and Pseudomonas aeruginosa (P) were grown alone or in combination as mono-, dual-, tri-, and tetra-species. Biofilm formation was determined by crystal violet spectrophotometric assay (n = nine biological replicates, with three technical replicates each). LIVE/DEAD staining and confocal laser scanning microscopy (CLSM) were used for samples with absorbance values > 0.5, including 3-D architectural analysis of viability using ImageJ software. Biofilm height was quantified by using the ZEN depth code tool. Statistical analysis included one-way ANOVA with Tukey’s post hoc test (p < 0.05), and unpaired t-tests were performed when applicable to compare variables. Biofilm formation was drastically increased in rich medium compared to moderate and minimal NB. The tetra-species cocktail SKEP reached the highest biofilm dry weight in BHI+G+S and greater elevation/viability compared to NB+G+S, whereas P. aeruginosa monospecies exhibited an opposite trend (greater elevation in NB+G+S than BHI+G+S). Interestingly, the E. faecalis and P. aeruginosa group exhibited strong biofilm characteristics in a range of nutritional environments. Confocal analysis showed structurally distinct biofilms: dense, vertical biofilms in BHI+G+S and sparse, horizontal patterns in NB+G+S. Nutrient composition significantly impacts polymicrobial biofilm formation, with nutrient-rich conditions (BHI+G+S) supporting robust viable three-dimensional structures, particularly within complex bacterial communities. These observations are of potential significance with respect to biofilm formation in hospital settings, such as in wounds, and may be used in strategies against the targeted intervention of polymicrobial infections. Full article
Show Figures

Graphical abstract

18 pages, 914 KB  
Article
Research on SCADA Data Preprocessing Method for Wind Turbines Based on Variable Grid Optimization (VGO-K-Means)
by Huilin Li and Junqing Li
Electronics 2026, 15(14), 3078; https://doi.org/10.3390/electronics15143078 - 13 Jul 2026
Viewed by 200
Abstract
To address the high dimensionality, redundancy, and noise interference present in wind turbine Supervisory Control and Data Acquisition (SCADA) data, as well as the limitations of conventional K-means algorithms—including excessive reliance on manual parameter tuning and weak anti-noise performance—this paper proposes a Variable [...] Read more.
To address the high dimensionality, redundancy, and noise interference present in wind turbine Supervisory Control and Data Acquisition (SCADA) data, as well as the limitations of conventional K-means algorithms—including excessive reliance on manual parameter tuning and weak anti-noise performance—this paper proposes a Variable Grid Optimized K-means (VGO-K-means) preprocessing algorithm. Z-score standardization is adopted to unify feature magnitudes. Meanwhile, an adaptive grid density calculation strategy is developed to dynamically adjust grid resolution according to the value range of each dimension, enabling accurate characterization of the spatial distribution of monitoring samples. Furthermore, a multi-index voting mechanism integrating the silhouette coefficient, Davies–Bouldin index, and inertia metric is established to adaptively determine the optimal cluster number without manual intervention. Utilizing the density discrepancy between dense normal samples and sparse outliers, the proposed method identifies abnormal samples through a clustering density threshold. Validated on real wind farm SCADA data containing 892 manually labeled abnormal samples, the VGO-K-means algorithm achieves a precision of 96.8% and an F1-score (the harmonic mean of precision and recall) of 0.89. Under identical test conditions, it outperforms traditional K-means, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), and fixed-grid K-means methods. The entire workflow consumes only 3.9 s, achieving an excellent balance between detection accuracy and computational cost. The proposed framework provides a reliable and practical preprocessing solution for wind turbine condition monitoring data. Full article
(This article belongs to the Section Power Electronics)
Show Figures

Figure 1

27 pages, 883 KB  
Article
Reconfigurable Transmission Design for PASS-MIMO via Waveguide Indexing
by Yaxian Wang, Songjie Yang and Juhong Peng
Sensors 2026, 26(14), 4407; https://doi.org/10.3390/s26144407 - 11 Jul 2026
Viewed by 263
Abstract
To address the stringent requirements of 6G industrial Internet of Things (IoT) and ultra-dense networks on spectral efficiency, hardware cost, and transmission reliability, this paper investigates waveguide index modulation based on the pinching-antenna system (PASS), a promising flexible multiple-input multiple-output (MIMO) architecture featuring [...] Read more.
To address the stringent requirements of 6G industrial Internet of Things (IoT) and ultra-dense networks on spectral efficiency, hardware cost, and transmission reliability, this paper investigates waveguide index modulation based on the pinching-antenna system (PASS), a promising flexible multiple-input multiple-output (MIMO) architecture featuring large-scale reconfigurability and robust line-of-sight (LoS) link establishment. A two-stage sparse transmission framework is proposed, where a Simulated Annealing-based Constrained Discrete Optimization (SA-CDO) algorithm is first employed to optimize pinching-antenna (PA) positions and construct a near-orthogonal equivalent channel dictionary for inter-waveguide interference suppression. Subsequently, an Orthogonal Least Squares-based Constellation-Constrained (OLS-CC) detector is developed to jointly recover active waveguide indices and modulation symbols with low computational complexity. Monte Carlo simulations demonstrate that the proposed scheme consistently outperforms conventional antenna index modulation under both LoS and Rician fading channels across the entire SNR range. The SA-CDO optimization significantly reduces the bit error rate (BER), while the OLS-CC detector further improves sparse recovery accuracy and reduces the detection complexity from exponential to polynomial order. These results provide valuable insights for the design of highly reliable 6G IoT communication systems. Full article
(This article belongs to the Special Issue MIMO Systems for Future Wireless Communications)
Show Figures

Figure 1

24 pages, 11360 KB  
Article
Influences of Pearlite Interlamellar Spacing on Wear and Rolling Contact Fatigue Behaviors of Pearlitic Rails on Field Tracks
by Junjie Fei, Hongfang Qi, Bei Yuan, Minbiao Wan and Linlang Zhang
Lubricants 2026, 14(7), 267; https://doi.org/10.3390/lubricants14070267 - 10 Jul 2026
Viewed by 269
Abstract
As a core load-bearing component for railway vehicles, rails are largely responsible for the safety and stability of train operation, and their service performance is inherently governed by material microstructure. In this study, rails with varied pearlite interlamellar spacing were prepared and laid [...] Read more.
As a core load-bearing component for railway vehicles, rails are largely responsible for the safety and stability of train operation, and their service performance is inherently governed by material microstructure. In this study, rails with varied pearlite interlamellar spacing were prepared and laid on field tracks for 8 months of service testing to investigate the influence of pearlite interlamellar spacing on rail wear and rolling contact fatigue (RCF). The results indicate that decreasing pearlite interlamellar spacing facilitated tread work hardening and reduced cumulative wear loss of rails. At the early service stage, rails with coarse pearlite lamellae exhibited earlier RCF crack initiation and longer crack morphologies, while rails featuring finer pearlite lamellae exhibited the latest-occurring crack initiation. With prolonged service duration, wear loss rose continuously, and the tread hardening rate first increased sharply and then tended to gradually become stable. Obvious differences in damage evolution were observed for rails with different pearlite interlamellar spacing. Coarse-lamellar rail suffered sparse short cracks dominated by wear; fine-lamellar rail developed dense fast-growing cracks controlled by RCF; and medium-lamellar rail achieved a relatively good balance between wear and RCF. A competitive relationship exists between wear and RCF during rail service. Reasonable regulation of pearlite interlamellar spacing facilitates a balanced evolution of wear and RCF, which provides a feasible microstructural optimization strategy for improving the service performance and service life of pearlitic rails. Full article
Show Figures

Figure 1

22 pages, 67005 KB  
Article
DEAF-Net: Dual-Domain Enhanced Adaptive Fusion Network for UAV Visible–Infrared Object Detection
by Qian Weng, Yu Zhang, Xiansheng Huang, Liming Deng and Jiawen Lin
Remote Sens. 2026, 18(13), 2241; https://doi.org/10.3390/rs18132241 - 7 Jul 2026
Viewed by 342
Abstract
In Unmanned Aerial Vehicle (UAV) object detection tasks, complex lighting conditions and variable weather render robust all-weather perception challenging when relying solely on the visible modality. Although infrared modalities can provide complementary information, the reliability of individual modalities is highly scene-dependent. Existing multimodal [...] Read more.
In Unmanned Aerial Vehicle (UAV) object detection tasks, complex lighting conditions and variable weather render robust all-weather perception challenging when relying solely on the visible modality. Although infrared modalities can provide complementary information, the reliability of individual modalities is highly scene-dependent. Existing multimodal detection methods typically adopt static fusion strategies, which ignore spatial heterogeneity of modal reliability and under-explore spatial-frequency collaborative representation, thus limiting detection robustness in dynamic environments. To address these issues, this paper proposes a Dual-domain Enhanced Adaptive Fusion Network (DEAF-Net), with two core innovative modules to tackle the above challenges. First, the Dual Domain Progressive Refinement (DDPR) module mitigates feature degradation caused by poor imaging conditions via the joint design of frequency-domain learnable filtering and scale-aware contextual refinement in the spatial domain, effectively suppressing noise, enhancing textures, and yielding a purified feature basis for fusion. Second, the Consistency–Discrepancy Guided Fusion (CDGF) strategy leverages the selective scanning mechanism of VMamba to model consistent and differential patterns across modalities, dynamically generates local modal contribution maps for adaptive fusion, and integrates global scene prior via entropy weights for calibration. Extensive experiments on the DroneVehicle and VEDAI datasets show that DEAF-Net outperforms mainstream multimodal detection methods, achieving mAP@0.5 scores of 81.9% and 76.2%, respectively, while delivering improved robustness in low-light, dense fog, and sparse-category scenarios. Full article
(This article belongs to the Special Issue Intelligent Processing of Multimodal Remote Sensing Data)
Show Figures

Figure 1

26 pages, 9175 KB  
Article
RT-DETR-DCEA: A Lightweight Citrus Defective Fruit Detection Algorithm for Complex Orchard Environments
by Jihui Qiao, Yuchen Sun, Binyuan Zhong, Lun Wang, Siyu Li, Hang Liu, Youqing Chen and Tong Li
Plants 2026, 15(13), 2077; https://doi.org/10.3390/plants15132077 - 3 Jul 2026
Viewed by 221
Abstract
Given the issues in natural orchard environments, such as large-scale variations of defective citrus fruits, weak texture boundaries, strong illumination changes, branch and leaf occlusion, and significant background interference, this paper constructs a lightweight detection model, RT-DETR-DCEA, based on RT-DETR-R18. This model is [...] Read more.
Given the issues in natural orchard environments, such as large-scale variations of defective citrus fruits, weak texture boundaries, strong illumination changes, branch and leaf occlusion, and significant background interference, this paper constructs a lightweight detection model, RT-DETR-DCEA, based on RT-DETR-R18. This model is improved through four aspects: “fine-grained defective feature extraction—multi-scale feature fusion—up-sampling detail recovery—global feature interaction for noise suppression”. First, a Dynamic Hybrid Convolution Module (DIMB) is introduced into the backbone network, drawing on the ideas of Inception-style multi-branch depthwise convolution and MetaFormer residual mixing. It extracts local textures of various forms through square convolution, horizontal strip convolution, and vertical strip convolution, and utilizes dynamic branch weights to enhance the model’s adaptability to irregular defects such as lesions, mildew, and external damage. Second, a Content-Guided Attention Feature Fusion Network (CGAFN) is designed in the neck network, which achieves adaptive fusion of low-level detail features and high-level semantic features through channel attention, spatial attention, and pixel-level fusion weights. Next, a lightweight upsampling enhancement module called EUCB-SC is constructed, which introduces channel rearrangement and Shift spatial offset into the efficient upsampling convolutional structure to enhance the local spatial interaction capability of upsampled features with low parameter overhead. Finally, adaptive sparse self-attention is introduced into the AIFI module to form AIFI-ASSA, which suppresses irrelevant background interactions through a sparse attention branch and retains necessary contextual information through a dense attention branch. The experimental results demonstrate that on a dataset containing four categories of citrus images—healthy, diseased, moldy, and severely externally damaged—RT-DETR-DCEA achieves 92.1% Precision, 86.1% Recall, and 91.8% mAP@50, with a parameter count of 1.477 × 107 and an inference speed of 81 FPS. Compared with the original RT-DETR-R18 and various YOLO series models, this method strikes a favorable balance among detection accuracy, recall capability, and model lightweightness. This paper also discusses limitations such as data scale, ratio of private data, single training result, and insufficient validation on edge devices, providing a basis for subsequent cross-regional data validation and real-world deployment testing. Full article
(This article belongs to the Special Issue AI-Driven Machine Vision Technologies in Plant Science)
Show Figures

Figure 1

24 pages, 34784 KB  
Article
Occluder-Mask-Constrained 3D Reconstruction from Tower-Crane Construction Site Imagery
by Qirun He, Rong Zhang, Changjiang Yin, Qin Ye and Shaoming Zhang
Electronics 2026, 15(13), 2883; https://doi.org/10.3390/electronics15132883 - 1 Jul 2026
Viewed by 248
Abstract
3D reconstruction of construction scenes is an important enabling technology for digital and intelligent construction project management. Recurring foreground occluders and dynamic disturbances in tower-crane imagery can destabilize image registration and introduce spurious depth responses. This paper proposes an occluder-mask-constrained 3D reconstruction framework [...] Read more.
3D reconstruction of construction scenes is an important enabling technology for digital and intelligent construction project management. Recurring foreground occluders and dynamic disturbances in tower-crane imagery can destabilize image registration and introduce spurious depth responses. This paper proposes an occluder-mask-constrained 3D reconstruction framework driven by multi-view geometric anomalies. Adjacent-view geometric outliers are spatially aggregated to generate foreground prompt points, which are converted into occluder masks using Segment Anything Model 2 (SAM2). The masks are propagated as unified pixel-validity constraints through sparse feature filtering, Adaptive Patch Deformation Multi-View Stereo (APD-MVS) matching-cost evaluation, support-region selection, and depth-map fusion. Experiments on three real construction-site datasets show increased sparse-registration completeness in the tested sequences and fewer visually identifiable occluder-induced artifacts in dense point clouds. A representative 308-image sequence was further evaluated against no-mask reconstruction, You Only Look Once version 8 (YOLOv8) bounding-box removal, manually prompted Segment Anything Model 2.1 (SAM2.1), a Segment Anything Model 3 (SAM3) text-prompt baseline, and Visibility-Aware Multi-View Stereo Network (Vis-MVSNet). The evaluation combines sparse-reconstruction metrics, pixel-level mask-quality metrics from a manually annotated validation subset, module-wise runtime accounting, controlled ablations, and aligned dense-point-cloud visualization. These results show improved sparse-stage registration completeness and visible artifact suppression. Because high-precision 3D reference point clouds are unavailable, the dense results are interpreted as visual evidence of artifact suppression rather than as proof of improved absolute dense-reconstruction accuracy. Full article
(This article belongs to the Special Issue Advances in Object Tracking and Localization)
Show Figures

Figure 1

18 pages, 3210 KB  
Article
Multimodal Feature-Level Fusion CBAM U-Net for Static Plantar Pressure Prediction Using Plantar Geometry and Sparse Anatomical Landmarks
by Chongguang Wang, Kerrie Evans, Dean Hartley, Scott Morrison, Stuart McDonald, Martin Veidt and Gui Wang
Sensors 2026, 26(13), 4143; https://doi.org/10.3390/s26134143 - 1 Jul 2026
Viewed by 340
Abstract
Accurate plantar pressure distribution is important for biomechanics, gait analysis, rehabilitation, and diabetic foot assessment. However, wearable plantar pressure systems are often limited by sparse sensor layouts due to hardware complexity, power consumption, and user comfort constraints. This study proposes a multimodal deep [...] Read more.
Accurate plantar pressure distribution is important for biomechanics, gait analysis, rehabilitation, and diabetic foot assessment. However, wearable plantar pressure systems are often limited by sparse sensor layouts due to hardware complexity, power consumption, and user comfort constraints. This study proposes a multimodal deep learning framework for static plantar pressure prediction using plantar geometry information and sparse landmark constraints. A convolutional block attention module U-Net architecture was developed to integrate plantar geometry and sparse landmark modalities through dual-encoder feature fusion with attention refinement. Different network architectures, fusion strategies, and landmark densities were systematically evaluated using a controlled-variable experimental design. Results demonstrated that feature-level fusion consistently outperformed data-level fusion and unimodal configurations across all landmark densities. The proposed model achieved the best performance with a normalized root mean square error of 0.087 using 16 landmarks, and the same model maintained a normalized root mean square error of 0.138 using only two landmarks, indicating promising reconstruction performance even under highly sparse sensing conditions. Marginal contribution and synergy analyses further showed that feature-level fusion more effectively captured complementary interactions between plantar geometry and sparse anatomical guidance, particularly under sparse landmark conditions. These findings suggest that multimodal feature-level fusion provides an effective strategy for sparse-to-dense plantar pressure reconstruction and may support the development of low-cost intelligent insole systems for biomechanical monitoring and clinical applications. Full article
Show Figures

Figure 1

17 pages, 6175 KB  
Article
Flexible Light Field Reconstruction: Enabling Arbitrary Sampling and Angular Resolution
by Xia Liu, Junzhen Ye, Zhangmin Wu and Qiang Fu
Electronics 2026, 15(13), 2763; https://doi.org/10.3390/electronics15132763 - 23 Jun 2026
Viewed by 181
Abstract
Compared with hardware-dependent methods, light field (LF) reconstruction algorithms enable a more economical and convenient acquisition of densely sampled LF (DSLF). Existing learning-based LF reconstruction methods suffer from limited flexibility, as they rely on fixed sampling patterns and predefined angular resolutions. In this [...] Read more.
Compared with hardware-dependent methods, light field (LF) reconstruction algorithms enable a more economical and convenient acquisition of densely sampled LF (DSLF). Existing learning-based LF reconstruction methods suffer from limited flexibility, as they rely on fixed sampling patterns and predefined angular resolutions. In this paper, we propose a flexible deep learning framework, which can reconstruct DSLF with arbitrary angular resolution from randomly distributed sparse input views of an arbitrary quantity. The proposed framework consists of two core stages, namely the SAI Synthesis and the LF Refinement. The SAI Synthesis adopts Plane Sweep Volume (PSV) to cope with randomly sampled input views, and leverages the Multi-Scale Attention (MSA) module to compute per-view weights for adaptive feature fusion and support arbitrary numbers of input views. The LF Refinement stage integrates intermediate results and fully exploits LF parallax structures to further improve reconstruction quality. Experimental results demonstrate that our method achieves superior flexibility and reconstruction quality, and outperforms most state-of-the-art LF reconstruction methods. Full article
(This article belongs to the Special Issue Computer Vision and Image Processing in Machine Learning)
Show Figures

Figure 1

22 pages, 4519 KB  
Article
Multi-Level Attention Dueling Double Deep Q-Network for Local Path Planning
by Hepengfei Wang, Jie Huang, Nan Wang and Huajie Hong
Appl. Sci. 2026, 16(12), 6235; https://doi.org/10.3390/app16126235 - 21 Jun 2026
Viewed by 368
Abstract
Deep reinforcement learning (DRL) has shown considerable potential in local path planning for autonomous robots. However, existing DRL methods still suffer from limited training efficiency, poor generalization, and weak sim-to-real transferability in complex environments. To address these issues, this paper proposes a Multi-Level [...] Read more.
Deep reinforcement learning (DRL) has shown considerable potential in local path planning for autonomous robots. However, existing DRL methods still suffer from limited training efficiency, poor generalization, and weak sim-to-real transferability in complex environments. To address these issues, this paper proposes a Multi-Level Attention Dueling Double Deep Q-Network (MLA-D3QN) framework, which progressively enhances feature extraction, spatial perception, and modality fusion through three attention levels: rule-based attention for obstacle contour extraction, implicit neural multi-scale spatial attention for environment perception, and bidirectional cross-attention for multi-modal feature alignment. Simulation results show that MLA-D3QN outperforms baseline and comparison methods in terms of convergence speed and average reward. Real-world experiments are conducted on a Scout mini platform with 50 trials in simple task scenarios (sparse obstacles, short distance) and 50 trials in complex task scenarios (dense obstacles, long distance). The proposed method achieves success rates of 98% in simple tasks and 94% in complex tasks. Compared to CNN-D3QN and D3QN, MLA-D3QN improves success rates by 10 percentage points (vs. CNN-D3QN) and 38 percentage points (vs. D3QN) in simple tasks, and by 34 percentage points (vs. CNN-D3QN) and 84 percentage points (vs. D3QN) in complex tasks. Path costs are reduced by 24.0% (vs. CNN-D3QN) and 59.9% (vs. D3QN). These results validate the effectiveness of MLA-D3QN in improving generalization and sim-to-real transferability for local path planning in complex environments. Full article
Show Figures

Figure 1

28 pages, 21429 KB  
Article
EDM-Net: A Multi-Scale Network for Object Detection in Remote Sensing Images
by Shuai Liang, Xiao Wang, Jialong Sun, Hui Liu and Huilei Yang
Sensors 2026, 26(12), 3927; https://doi.org/10.3390/s26123927 - 20 Jun 2026
Viewed by 443
Abstract
Remote sensing object detection remains challenging because objects often appear with large scale variation, dense spatial layouts, and strong interference from complex geographical backgrounds. To address these coupled difficulties, we propose EDM-Net, an end-to-end multi-scale detector that organizes feature processing into three coordinated [...] Read more.
Remote sensing object detection remains challenging because objects often appear with large scale variation, dense spatial layouts, and strong interference from complex geographical backgrounds. To address these coupled difficulties, we propose EDM-Net, an end-to-end multi-scale detector that organizes feature processing into three coordinated stages: adaptive extraction, intra-scale interaction, and cross-scale fusion. First, an efficient sparse mixture-of-experts (ES-MoE) module is embedded in the backbone to allocate scale-specific convolutional experts according to scene-level feature responses, providing a more adaptive feature basis than a single static extraction path. Second, a dynamic mixing intra-scale feature interaction (DMIFI) module is introduced into the Transformer encoder. This module combines global self-attention with dynamic spatial mixing, thereby preserving long-range context while reintroducing local two-dimensional inductive bias for dense and small objects. Third, a multi-scale synergistic attention fusion (MSAF) module aligns adjacent feature levels through parallel local and global attention branches and structural re-parameterization, reducing semantic dilution during feature aggregation. Comprehensive experiments on three large-scale remote sensing benchmark datasets, DIOR, NWPU VHR-10, and RSOD, demonstrate that EDM-Net consistently improves over the re-trained RT-DETR-R18 baseline under the same experimental protocol, attaining mAP50 scores of 83.7%, 95.6%, and 95.8% respectively. Additional ablation and scale-specific analyses indicate that the three modules contribute complementary gains, especially for small and densely distributed objects. These results suggest that coordinated extraction, interaction, and fusion can improve remote sensing object detection under complex scale and background conditions. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

17 pages, 2753 KB  
Article
KoSim-GL: A Large-Scale Simulation-Based Dataset for UAV Cross-View Geo-Localization in Korean Urban Environments
by Heejin Ahn, Changhwan Lee, Sangwook Lee, HyeonJoong Wi, Insung Jang and Dong-Geol Choi
Electronics 2026, 15(12), 2720; https://doi.org/10.3390/electronics15122720 - 19 Jun 2026
Viewed by 328
Abstract
We propose KoSim-GL, a large-scale vision-based geo-localization dataset for drone positioning in GPS-denied environments. Geo-localization estimates a drone’s location by matching drone-view imagery against a geo-referenced satellite image database, offering a reliable alternative to GPS under conditions such as signal jamming, spoofing, or [...] Read more.
We propose KoSim-GL, a large-scale vision-based geo-localization dataset for drone positioning in GPS-denied environments. Geo-localization estimates a drone’s location by matching drone-view imagery against a geo-referenced satellite image database, offering a reliable alternative to GPS under conditions such as signal jamming, spoofing, or degradation in dense urban canyons. Although this task is challenging due to the domain gap between drone-view and satellite-view imagery, existing benchmarks are built predominantly around urban environments in the United States and China, leaving South Korea largely unrepresented, despite its distinctive landscape in which mountainous terrain coexists with dense high-rise districts and low-rise residential neighborhoods. To address this gap, we introduce KoSim-GL, constructed from drone-view images captured via an AirSim- and ROS-based flight simulator and satellite images collected through the Google Maps Tile API, covering the urban area of Daejeon, South Korea. Its key feature is a multi-view configuration that simultaneously captures five views, one nadir and four oblique, at each flight position across altitudes from 100 m to 600 m, enabling robust localization even in feature-sparse environments where nadir-only matching is prone to fail. In total, KoSim-GL comprises 2,450,315 drone images and 1704 satellite images. We further provide systematic comparisons against five existing benchmarks and baseline evaluations of ten representative geo-localization models under single- and multi-view settings. Experimental results show that the multi-view configuration substantially improves localization performance; for example, FSRA improves Recall@1 from 44.08% (single-view) to 65.37% (multi-view), a gain of 21.29 percentage points. The dataset is publicly available. Full article
(This article belongs to the Section Computer Science & Engineering)
Show Figures

Figure 1

21 pages, 1161 KB  
Article
SSMSNet: Scribble-Supervised Myocardial Scar Segmentation in Late Gadolinium Enhancement Images
by Xuewen Liao, Kangwen Yang, Xingtao Lin, Lin Pan, Yazhou Lin, Mingjing Yang and Jiancheng Zhang
Diagnostics 2026, 16(12), 1895; https://doi.org/10.3390/diagnostics16121895 - 18 Jun 2026
Viewed by 299
Abstract
Background: Myocardial scar segmentation from late gadolinium enhancement (LGE) cardiac magnetic resonance (CMR) images plays an important role in cardiac disease assessment and prognosis evaluation. However, accurate scar annotation is labor-intensive and requires substantial clinical expertise because scar regions are typically small, [...] Read more.
Background: Myocardial scar segmentation from late gadolinium enhancement (LGE) cardiac magnetic resonance (CMR) images plays an important role in cardiac disease assessment and prognosis evaluation. However, accurate scar annotation is labor-intensive and requires substantial clinical expertise because scar regions are typically small, irregularly shaped, and characterized by ambiguous boundaries. Although scribble supervision provides a more practical alternative to dense annotation by substantially reducing labeling costs, the extreme sparsity of scribbles and the high similarity between scar tissue and surrounding myocardium make accurate weakly supervised segmentation challenging. Methods: To address these challenges, we propose SSMSNet, a novel scribble-supervised framework for myocardial scar segmentation. Specifically, a weakly supervised anatomical segmentation network is first employed to provide reliable myocardial structural priors and suppress irrelevant background interference. Subsequently, a local distance prior map is dynamically generated from scribble annotations, and a corresponding loss is introduced to enhance structural awareness and improve training stability. Meanwhile, by leveraging the spatial correlation between the myocardium and scar regions, teacher–student consistency supervision progressively recovers more complete scar structures from sparse annotations. Furthermore, a detail-aware feature enhancement module strengthens low-level representations through contextual interactions and attention mechanisms, improving the perception of scars with ambiguous boundaries. Results: Extensive experiments conducted on two public cardiac pathology datasets demonstrate that the proposed framework consistently outperforms state-of-the-art scribble-supervised methods and achieves competitive performance compared with fully supervised methods. Conclusions: The proposed SSMSNet effectively alleviates the limitations imposed by scribble annotations by integrating anatomical guidance, local distance priors, and consistency learning. These findings suggest that the framework provides an effective and annotation-efficient solution for myocardial scar segmentation in LGE CMR images. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Figure 1

23 pages, 2122 KB  
Article
DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion
by Xinyi Feng, Shaochen Jiang, Liejun Wang and Beibei Gao
Sensors 2026, 26(12), 3864; https://doi.org/10.3390/s26123864 - 17 Jun 2026
Viewed by 341
Abstract
High-resolution remote sensing image segmentation is important for urban mapping but remains challenging because of spectral ambiguity, large scale variations, fragmented elongated structures, and background interference. This study aims to improve semantic segmentation in complex aerial scenes by combining local feature extraction, selective [...] Read more.
High-resolution remote sensing image segmentation is important for urban mapping but remains challenging because of spectral ambiguity, large scale variations, fragmented elongated structures, and background interference. This study aims to improve semantic segmentation in complex aerial scenes by combining local feature extraction, selective multi-scale fusion, and global sequence modeling. We propose DSD-Mamba, an asymmetric dual-stream architecture with a ResNet-18 encoder. The Dense-Sparse Pyramid Fusion Module aligns multi-level features and applies dual Top-k selective value aggregation for cross-scale response filtering and background-response suppression. This Top-k operation is used as a feature-selection mechanism and is not intended to reduce the theoretical memory footprint of dense attention. Scale-Aware Strip Attention refines skip connections through horizontal and vertical dependency modeling, and the Dual-Stream Context Decoder combines a Mamba-based global branch with a CNN-based local branch during upsampling. Experiments were conducted on UAVid, ISPRS Vaihingen, and ISPRS Potsdam under a single-model inference protocol without test-time augmentation. DSD-Mamba achieved mIoU scores of 73.4%, 85.2%, and 87.2%, respectively. Ablation experiments on Vaihingen showed that DSPFM, SASA, and DSCD improved performance over the baseline when evaluated in this setting, with the full model reaching the highest mIoU. The method improves segmentation accuracy under the tested protocols, although its higher FLOPs indicate an accuracy-oriented rather than lightweight design. Full article
Show Figures

Figure 1

Back to TopTop