Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (194)

Search Parameters:
Keywords = prior-conditioned feature modulation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 33421 KB  
Article
Climate-Aware Self-Retrospective Representation Learning for Spatio-Temporal Epidemic Forecasting
by Qi Yuan, Han Shu, Yizhi Pan, Tianshuo Li, Hangyi Shen, Weiqi Jiang, Zidan Zhu, Pengpeng Zhang, Ningli Xi, Junyi Xin, Kai Li and Guanqun Sun
Trop. Med. Infect. Dis. 2026, 11(9), 240; https://doi.org/10.3390/tropicalmed11090240 - 24 Aug 2026
Abstract
Spatio-temporal epidemic forecasting aims to predict future outbreak trajectories across interconnected regions from historical epidemiological observations and meteorological covariates. However, existing approaches often fail to preserve historically salient epidemic states or to fully exploit delayed and region-varying meteorological associations, leading to unstable temporal [...] Read more.
Spatio-temporal epidemic forecasting aims to predict future outbreak trajectories across interconnected regions from historical epidemiological observations and meteorological covariates. However, existing approaches often fail to preserve historically salient epidemic states or to fully exploit delayed and region-varying meteorological associations, leading to unstable temporal representations and insufficient meteorological-context-aware spatio-temporal context for prediction at later forecast horizons. In this paper, we propose CASRL, a Climate-Aware Self-Retrospective Representation Learning network for stable and meteorological-context-aware spatio-temporal epidemic forecasting. CASRL first employs a Self-Retrospective Epidemic Encoder (SREE) to retrospectively aggregate historically salient epidemic states through query-guided weighting and adaptive gating, thereby preserving informative historical epidemic states within the look-back window. It then introduces a Climate-Adaptive Graph Message Passing (CAGMP) module that breaks away from traditional passive feature concatenation. Instead, it constructs a separate meteorological-view predictive graph conditioned on the static spatial prior and adaptively fuses it with the incidence-associated topology to model complex cross-regional predictive associations. By integrating self-retrospective epidemic representations with meteorological-view spatio-temporal interactions, CASRL produces forecasts with improved predictive stability at later forecast horizons. Extensive experiments on two public influenza benchmarks show that CASRL is competitive at shorter forecast horizons and provides clearer advantages at later forecast horizons, particularly in phase-alignment-related evaluation and 15-week-ahead forecasting. Full article
Show Figures

Figure 1

45 pages, 5137 KB  
Article
FO-FCGFNet: A Fractional-Order Image Processing and Fractal Complexity-Guided Intelligent Estimation Method for Fault Diagnosis in Oil-Immersed Transformer Complex Systems
by Xin Zhang, Yuanda Song and Chunpeng Xu
Fractal Fract. 2026, 10(8), 587; https://doi.org/10.3390/fractalfract10080587 - 21 Aug 2026
Viewed by 109
Abstract
A fractional-order image enhancement and fractal complexity-guided fusion method was developed to improve weak fault representation and introduce complexity-aware priors into dissolved gas analysis (DGA)-based diagnosis of oil-immersed power transformers. Gas-ratio features derived from five characteristic gases were combined into an extended DGA [...] Read more.
A fractional-order image enhancement and fractal complexity-guided fusion method was developed to improve weak fault representation and introduce complexity-aware priors into dissolved gas analysis (DGA)-based diagnosis of oil-immersed power transformers. Gas-ratio features derived from five characteristic gases were combined into an extended DGA feature sequence and converted into two-dimensional representations using the Markov transition field (MTF), recurrence plot (RP), and Gramian angular field (GAF). A fractional-order difference operator then strengthened texture details and local variations, while fractal complexity features quantified structural irregularities across fault conditions. Based on these features, a fractal complexity-guided multi-image attention fusion module was designed to adaptively integrate the three image representations. An improved RIME optimization algorithm was further employed to jointly optimize the fractional order, imaging parameters, and network hyperparameters. On the public DGA dataset, the proposed model achieved precision, recall, accuracy, and F1-score values of 97.68%, 97.51%, 97.82%, and 97.71%, respectively. External validation on a self-collected DGA dataset further demonstrated its robust cross-condition generalization capability. Full article
Show Figures

Figure 1

23 pages, 11731 KB  
Article
A Physics-Guided Raw-Dominant Gated Fusion Method for Fine-Grained Bearing Fault Diagnosis
by Chuanbo Wu, Guoao Jiao, Yongdi Zhang, Kangning Jin, Zihang Zhang and Zeming Li
Appl. Sci. 2026, 16(16), 8337; https://doi.org/10.3390/app16168337 - 21 Aug 2026
Viewed by 154
Abstract
Fine-grained bearing condition diagnosis is challenging because different bearing states within the same fault location often exhibit similar fault-characteristic-frequency responses, making them difficult to distinguish using envelope-spectrum information alone. To address this problem, a physics-guided raw-dominant gated fusion network (PG-RDGFN) is proposed for [...] Read more.
Fine-grained bearing condition diagnosis is challenging because different bearing states within the same fault location often exhibit similar fault-characteristic-frequency responses, making them difficult to distinguish using envelope-spectrum information alone. To address this problem, a physics-guided raw-dominant gated fusion network (PG-RDGFN) is proposed for fine-grained bearing fault diagnosis. In the proposed framework, the raw vibration signal is retained as the dominant information source, while the envelope spectrum provides complementary fault-modulation evidence. A physics-aware descriptor derived from bearing characteristic-frequency responses is incorporated into a sample-wise gating mechanism to adaptively regulate the contribution of the envelope-spectrum features. Distinct from conventional direct multi-branch fusion or physics-informed schemes that mainly use physical knowledge as an auxiliary input or regularization constraint, PG-RDGFN uses mechanism-derived physical confidence to regulate how much auxiliary envelope-spectrum evidence participates in the fusion, rather than directly using the physical prior as a fine-grained classification cue. Meanwhile, a physical-consistency loss constrains the learned gate using a scalar physical-confidence target, and a raw-branch auxiliary loss preserves the discriminative capability of the dominant raw representation. Experiments are conducted on an eight-class diagnosis task constructed from the Paderborn University bearing dataset. Compared with SVM, MLP, 1D-CNN, CNN-LSTM, and TCN, the PG-RDGFN achieves the highest accuracy of 98.87%. Ablation and gate-consistency analyses further verify the effectiveness of the envelope branch, gated fusion, physics guidance, and auxiliary supervision. These results demonstrate that PG-RDGFN provides an accurate and physically interpretable solution for fine-grained bearing condition diagnosis. Full article
Show Figures

Figure 1

22 pages, 2602 KB  
Article
A Balanced Spectral–Spatial Cross-Fusion Network for Hyperspectral Anomaly Detection
by Yuquan Gan, Mengjiao Wang, Lei Zhang, Weidong Zhang, Tao Yu and Hongwei Wang
Remote Sens. 2026, 18(16), 2820; https://doi.org/10.3390/rs18162820 - 20 Aug 2026
Viewed by 170
Abstract
Hyperspectral anomaly detection aims to find abnormal targets without prior information. However, the high dimensionality of hyperspectral data, complicated spatial structures, and varying object scales make it challenging to jointly utilize spectral and spatial information. Therefore, anomalies may be confused with background regions. [...] Read more.
Hyperspectral anomaly detection aims to find abnormal targets without prior information. However, the high dimensionality of hyperspectral data, complicated spatial structures, and varying object scales make it challenging to jointly utilize spectral and spatial information. Therefore, anomalies may be confused with background regions. A Balanced Spectral–Spatial Cross-Fusion Network (BSCF-Net) is proposed for hyperspectral anomaly detection. The network uses a multi-branch encoder, where spatial branches capture features with different receptive fields and spectral branches extract spectral patterns through one-dimensional convolutions and channel attention. The Bidirectional Spectral–Spatial Cross-Attention (BSCA) mechanism enables information exchange between spectral and spatial features. The Multi-Scale Gated Refiner (MSGR) module is used to refine the fused features. With an autoencoder reconstruction framework, BSCF-Net identifies anomalies according to reconstruction errors and reduces background interference. Experimental results on five public hyperspectral datasets demonstrate the effectiveness of BSCF-Net, achieving competitive AUC performance under diverse background conditions. Full article
(This article belongs to the Section Earth Observation Data)
Show Figures

Figure 1

22 pages, 65601 KB  
Article
Dual-Domain Illumination Prior for Low-Light Remote Sensing Image Enhancement
by Chao Wang, Zhe Pan, Liangtian He, Jun Liu, Lin Mei, Rongsheng Lin, Hongming Chen and Chuansheng Yang
Remote Sens. 2026, 18(16), 2817; https://doi.org/10.3390/rs18162817 - 20 Aug 2026
Viewed by 158
Abstract
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively [...] Read more.
Low-light conditions degrade remote sensing imagery by reducing contrast, distorting color, and obscuring fine terrain structures and small objects critical for Earth observation. Accurate illumination adjustment under spatially varying scene content remains challenging for existing enhancement methods, and many prior-guided approaches operate exclusively in either the spatial domain or the frequency domain. In this work, we propose a Dual-Domain Illumination Prior (DDIP), a trainable dual-domain illumination-prior module that is jointly optimized with each host backbone and exploits frequency-domain and spatial-domain illumination statistics. DDIP comprises three components: a Frequency-Domain Illumination Distribution Prior (FIDP) that performs per-color-channel amplitude calibration in Fourier space to improve global brightness; a Spatial-Domain Illumination Distribution Prior (SIDP), adapted from IDP-Net, that performs multi-scale sub-region statistical correction for local illumination adjustment; and a Selective Core Feature Fusion (SCFF) module that adaptively combines the frequency-domain output, the spatial-domain output, and the original input through an attention-based gating mechanism with dual pooling. DDIP is integrated with each host backbone while leaving its main restoration blocks unchanged. In the controlled reconstruction comparisons on iSAID-dark and the evaluated general low-light benchmarks, equipping the tested backbone networks with DDIP improves PSNR and SSIM over their corresponding baselines. Complementary LPIPS and CIELAB lightness measurements characterize perceptual similarity and lightness behavior, while a fixed-detector object-detection evaluation on the tested high-resolution iSAID-dark scenes examines the effect of the enhancement pipelines under the reported synthetic low-light conditions. The ablation studies further examine the contribution of the module components within the reported experimental settings. Full article
Show Figures

Figure 1

60 pages, 11445 KB  
Article
A Mamba-Driven Spatiotemporal Graph Neural Network for Fault Location in Low-Observability Active Distribution Networks
by Zhengying Hou, Jilong Ma and Xuguang Hu
Machines 2026, 14(8), 948; https://doi.org/10.3390/machines14080948 - 19 Aug 2026
Viewed by 163
Abstract
Accurate fault location in low-observability active distribution networks is hindered by uncertain inter-node relationships, underutilized early transients, and insufficient global context. To address these challenges, this paper proposes an adaptive Mamba-driven spatiotemporal graph neural network (AM-STGNN). It provides a unified task-driven spatiotemporal representation [...] Read more.
Accurate fault location in low-observability active distribution networks is hindered by uncertain inter-node relationships, underutilized early transients, and insufficient global context. To address these challenges, this paper proposes an adaptive Mamba-driven spatiotemporal graph neural network (AM-STGNN). It provides a unified task-driven spatiotemporal representation framework that progressively integrates fault-propagation modeling, global dependency modeling, transient-state learning, and topology-aware discriminative enhancement. Specifically, a prior-guided adaptive implicit topology is first learned to characterize task-dependent electrical coupling relationships among sparse observation nodes. Based on the resulting topology, topology-conditioned multi-order feature propagation and a dual-axis linear-attention module based on the spatiotemporal graph transformer (STGformer) are employed to capture local and global spatiotemporal dependencies. The resulting global spatiotemporal representation is subsequently processed by a Mamba selective state-space encoder to model input-dependent temporal evolution and emphasize informative fault transients. Finally, element-wise gated fusion, topology-aware differential output, and a margin constraint are employed to integrate the STGformer and Mamba representations and enhance the separability of adjacent faulted line sections with similar response characteristics. Extensive experiments demonstrate the effectiveness of AM-STGNN, while additional evaluations confirm its applicability to larger-scale networks, strongly phase-unbalanced conditions, and field-measured operating backgrounds. Robustness tests under individual and multi-level joint disturbances further demonstrate the practical relevance of the proposed architecture. Compared with the baseline models, AM-STGNN achieves consistent improvements in the macro-averaged F1 score (Macro-F1), exact accuracy, and one-hop accuracy under the clean IEEE 123-node condition. More importantly, it maintains clear performance advantages under identical mild, moderate, and severe joint disturbances, demonstrating improved robustness and practical relevance under simulated non-ideal operating conditions. Full article
Show Figures

Figure 1

28 pages, 3635 KB  
Article
VCDH-YOLO: Viewpoint-Conditioned Dual-Head Detection for Mixed-View Crack Inspection Across Drone and Ground Platforms
by Fangyi Lu, Yifan Hu, Yutong Guo and Zhenglong Ding
Remote Sens. 2026, 18(16), 2791; https://doi.org/10.3390/rs18162791 - 18 Aug 2026
Viewed by 205
Abstract
Mixed-viewpoint pavement crack detection remains challenging because aerial (Drone) and ground-level (Ground) images exhibit substantially different feature distributions, while repeated downsampling inevitably weakens the representation of fine crack structures in UAV (unmanned aerial vehicles) imagery. To address these issues, this study proposes VCDH-Net, [...] Read more.
Mixed-viewpoint pavement crack detection remains challenging because aerial (Drone) and ground-level (Ground) images exhibit substantially different feature distributions, while repeated downsampling inevitably weakens the representation of fine crack structures in UAV (unmanned aerial vehicles) imagery. To address these issues, this study proposes VCDH-Net, a lightweight mixed-viewpoint crack detection framework built upon YOLOv11n. The framework introduces a Viewpoint-Conditioned Dual-Head Detection (VCDH) architecture that dynamically routes features to viewpoint-specific detection heads through a lightweight viewpoint classifier, enabling specialized optimization while maintaining a shared feature extraction backbone. On this basis, a Lightweight Structure Enhancement (LSE) module is incorporated into mid-level feature layers to reinforce directional crack structures by exploiting local contrast and geometric priors. Furthermore, a Pyramid Detail Refinement (PDR) module is developed for the Drone branch to recover fine-grained spatial information of ultra-small cracks through a lightweight upsample–refine–downsample residual pathway. To provide a more comprehensive evaluation of mixed-viewpoint detection performance, a cross-view assessment framework is further established by introducing three complementary metrics, namely Cross-View Gap (CV-Gap), Worst-View Score (VWS), and Cross-View Balance (CVB). Experiments conducted on a dual-viewpoint pavement crack dataset collected from roads in and around Nanjing, China demonstrate that the proposed method achieves an mAP50-95 of 57.88%, improving the baseline YOLOv11n by 1.55 percentage points. Meanwhile, the Drone-view mAP50-95 increases from 46.98% to 48.83%, and the proposed framework attains the highest CVB score of 0.4965, indicating improved cross-viewpoint detection consistency under the tested conditions. These results, validated on a single-region dataset, demonstrate that VCDH-Net effectively alleviates viewpoint-induced feature discrepancies on the tested data while enhancing the representation of fine crack structures. Generalization to other geographic regions requires further validation on multi-viewpoint datasets not yet publicly available. Full article
(This article belongs to the Special Issue Object Detection and Tracking in Satellite Imagery and Video)
Show Figures

Figure 1

30 pages, 1840 KB  
Article
Weak Ridge-Flow Prior-Guided Fingerprint Reconstruction Under Severe Degradation
by Haiyong Xie, Lin Wang, Yonghao Dai and Yunqian Cheng
Computers 2026, 15(8), 527; https://doi.org/10.3390/computers15080527 - 14 Aug 2026
Viewed by 172
Abstract
Fingerprint enhancement plays an important role in recovering identity-related ridge structures from degraded fingerprints. However, existing methods primarily focus on local texture restoration and may struggle to preserve ridge continuity and structural consistency under severe degradation conditions, including ridge fragmentation, diffusion blur, and [...] Read more.
Fingerprint enhancement plays an important role in recovering identity-related ridge structures from degraded fingerprints. However, existing methods primarily focus on local texture restoration and may struggle to preserve ridge continuity and structural consistency under severe degradation conditions, including ridge fragmentation, diffusion blur, and partial information loss. In this paper, we observe that degraded fingerprints may retain incomplete ridge-flow information that can provide useful structural guidance for fingerprint reconstruction. Based on this observation, we propose a conditional generative adversarial network guided by a weak ridge-flow prior (WRP-cGAN) for degraded fingerprint enhancement. The proposed method treats the estimated ridge-flow information as a weak structural prior rather than an exact structural constraint and introduces prior-conditioned feature modulation to adaptively incorporate structural cues during reconstruction. The framework is jointly optimized using adversarial, image-space reconstruction, ridge-flow orientation-consistency, and gradient-consistency losses to improve ridge continuity, structural coherence, and local detail preservation. On the NIST SD301-derived test set, the proposed method increases the median NFIQ2 score from 9 to 44, improves the minutiae-restoration F1-score from 0.2507 to 0.5426, and increases the SourceAFIS Rank-1 identification rate from 39% to 86%. An additional qualitative evaluation on FVC2004 DB1 provides preliminary evidence of cross-dataset transferability without fine-tuning. These results suggest that weak ridge-flow priors provide useful structural guidance for degraded fingerprint reconstruction and improve recognition-oriented fingerprint quality under the degradation conditions considered in this study. Full article
(This article belongs to the Section AI-Driven Innovations)
Show Figures

Figure 1

22 pages, 1095 KB  
Article
Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech
by Rakhat Meiramov, Tomiris Rakhimzhanova, Adil Taibassarov, Zhanat Makhataeva and Huseyin Atakan Varol
Mach. Learn. Knowl. Extr. 2026, 8(8), 246; https://doi.org/10.3390/make8080246 - 14 Aug 2026
Viewed by 296
Abstract
Target-speaker automatic speech recognition (TS-ASR) enables transcription of a specific speaker in multi-talker environments, yet remains largely unexplored for multilingual, low-resource languages. Existing TS-ASR systems predominantly target monolingual English using diarization-based or speaker-embedding approaches, leaving a critical gap for languages such as Kazakh, [...] Read more.
Target-speaker automatic speech recognition (TS-ASR) enables transcription of a specific speaker in multi-talker environments, yet remains largely unexplored for multilingual, low-resource languages. Existing TS-ASR systems predominantly target monolingual English using diarization-based or speaker-embedding approaches, leaving a critical gap for languages such as Kazakh, where code-switching with Russian and English is commonplace. We propose Persona-ASR, a modular two-stage architecture. The first stage is an explicit target-presence gate that verifies whether the enrolled speaker appears in the mixture and emits a <no_target> token to suppress transcription when the speaker is absent, directly addressing the acoustic-hallucination failure mode of prior systems. The second stage performs enrollment-conditioned recognition: a 192-dimensional ECAPA-TDNN speaker embedding modulates a WavLM-Base-Plus encoder through feature-wise linear modulation (FiLM), while language-specific CTC heads enable joint Kazakh and English decoding without forcing Latin and Cyrillic symbols to compete in a single output space. To evaluate the system, we introduce KazMix3, a Kazakh overlap dataset for TS-ASR training, and PersonaMix, a controlled bilingual benchmark spanning same- and cross-language enrollment across varying interferer counts (1–3) and signal-to-noise ratios (3 to +3 dB). Persona-ASR outperforms a strong off-the-shelf cascade baseline by 13.3 WER points on English and 24.6 on Kazakh, and matches a published monolingual English baseline. On PersonaMix, speaker conditioning reduces relative word error rate by 40.7% on English and 59.3% on Kazakh mixtures over an unconditioned variant of the same model, and cross-language enrollment (unseen during training) remains effective, increasing average raw WER by only 4.1 points (English) and 2.2 points (Kazakh) relative to same-language enrollment. To our knowledge, Persona-ASR is the first TS-ASR system for the Kazakh language, establishing a foundation for multilingual personalized ASR in low-resource settings. Full article
Show Figures

Figure 1

40 pages, 3662 KB  
Article
A Semantic-Conditional GAN Framework for Structure-Preserving and Controllable Interior Style Generation
by Tianxi Lu, Chang Wen and Siti Sarah Binti Herman
Appl. Sci. 2026, 16(16), 7892; https://doi.org/10.3390/app16167892 - 7 Aug 2026
Viewed by 202
Abstract
Interior style generation requires expressive visual transformation while preserving spatial layout, object boundaries, and semantic relationships. Existing generative and style-transfer methods have improved indoor scene synthesis, but they often suffer from boundary drift, furniture deformation, texture leakage, and unstable style representation when stronger [...] Read more.
Interior style generation requires expressive visual transformation while preserving spatial layout, object boundaries, and semantic relationships. Existing generative and style-transfer methods have improved indoor scene synthesis, but they often suffer from boundary drift, furniture deformation, texture leakage, and unstable style representation when stronger style signals are introduced. To address this problem, this study proposes a semantic-conditional generative adversarial network framework for structure-preserving and controllable interior style generation. The framework constructs semantic masks, boundary maps, and semantic embeddings from indoor images and injects these semantic priors into the generator feature stream through a multi-scale conditional mechanism. A style encoding network is further introduced to represent interior style characteristics and to modulate image-level visual appearance features, including texture patterns, color tones, material-like surface appearance, and illumination-related visual cues, in a controllable manner. The model was trained and evaluated using 5000 selected indoor images from SUN RGB-D and ADE20K, together with a self-constructed style reference set of 600 images covering modern minimalist, Nordic, industrial, and neoclassical interiors. Across baseline comparisons and ablation analyses, the proposed framework achieved a semantic region consistency score of 0.91, maintained higher semantic boundary consistency under increasing style intensity, obtained style consistency scores ranging from 0.88 to 0.90, and achieved an LPIPS score of 0.128. A subjective evaluation with 20 participants also showed higher ratings for realism, style expression, and structural consistency. These results provide quantitative and perceptual evidence that semantic-conditional injection can improve the balance between spatial-semantic preservation and controllable image-level style expression in AI-assisted interior visualization. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

39 pages, 13901 KB  
Article
Traffic-Prior-Guided State-Aware Framework for Robust Urban Traffic Anomaly Detection
by Lingguang Wang, Changbo Kang, Yanchen Qiu, Yixuan Shang, Xiaomeng Wang and Qifeng Yu
Urban Sci. 2026, 10(8), 457; https://doi.org/10.3390/urbansci10080457 - 7 Aug 2026
Viewed by 271
Abstract
Urban traffic systems are increasingly vulnerable to non-recurrent congestion and abnormal traffic fluctuations, posing significant challenges to intelligent traffic management and resilient transportation operations. Existing traffic anomaly detection methods often struggle to simultaneously characterize heterogeneous anomaly patterns under dynamically evolving traffic states, while [...] Read more.
Urban traffic systems are increasingly vulnerable to non-recurrent congestion and abnormal traffic fluctuations, posing significant challenges to intelligent traffic management and resilient transportation operations. Existing traffic anomaly detection methods often struggle to simultaneously characterize heterogeneous anomaly patterns under dynamically evolving traffic states, while severe class imbalance and limited data plausibility further constrain detection reliability. To address these challenges, this study proposes a traffic-prior-guided state-aware framework for robust urban traffic anomaly detection. A Multi-Scale Natural Neighborhood (MS-NaN) module transforms one-dimensional traffic flow sequences into a nine-dimensional representation integrating sequence dynamics, multiscale statistical deviations, and spatiotemporal phase characteristics, thereby embedding traffic state priors into the detection process. Building upon these representations, the Dual-Branch Context-Gated Network (DB-CGNet) separately captures instantaneous traffic disruptions and trend-evolving congestion patterns. An adaptive context-aware gated fusion mechanism then combines the branch features to enhance robustness under complex and non-stationary traffic conditions. To improve evaluation realism, high-fidelity baseline traffic data are generated through B-spline smoothing and first-order autoregressive residual modeling, and anomaly patterns are constructed under Highway Capacity Manual (HCM)-constrained capacity reduction mechanisms. Experiments conducted on a 91-day urban expressway dataset demonstrate that the proposed method achieves the best overall performance among eight benchmark models under a 72 min observation window, attaining an F1-score of 0.7757 and an area under the receiver operating characteristic curve (AUC) of 0.9112. Ablation studies further reveal the critical role of traffic prior features in detecting short-duration evolving anomalies. The proposed framework provides a robust and interpretable solution for intelligent urban traffic monitoring, anomaly warning, and resilient traffic operation management. Full article
(This article belongs to the Section Urban Mobility and Transportation)
Show Figures

Figure 1

33 pages, 23794 KB  
Article
Navigation Line Extraction Method for Alfalfa Crops Based on RACG-RandLA Point Cloud Segmentation Model
by Kehua Dang, Jiachen Cao, Pengjie Pan, Zijie Niu, Zehan Lu, Dongyan Zhang and Yongjie Cui
Agriculture 2026, 16(15), 1683; https://doi.org/10.3390/agriculture16151683 - 5 Aug 2026
Viewed by 344
Abstract
Early-stage alfalfa navigation faces challenges like low plants, narrow rows, and weed interference, causing camera–LiDAR colored point clouds to suffer from sparsity, discontinuous boundaries, and varying illumination. Standard point-level semantic segmentation struggles to support stable crop row allocation and navigation line fitting under [...] Read more.
Early-stage alfalfa navigation faces challenges like low plants, narrow rows, and weed interference, causing camera–LiDAR colored point clouds to suffer from sparsity, discontinuous boundaries, and varying illumination. Standard point-level semantic segmentation struggles to support stable crop row allocation and navigation line fitting under these conditions. To address this, we propose RACG-RandLA, a multi-output row-aware color-geometric point cloud segmentation model. Pseudo-labels (crop, background, ‘ignore’) are generated using color and spatial priors, alongside transverse offset and point-level confidence labels for crop points. Built on RandLA-Net, the multi-task network simultaneously outputs crop semantics, transverse offsets, and confidences. It features a color-geometry residual fusion module that adaptively integrates 3D geometric and RGB/ExG features via zero-initialized scaling to handle complex lighting and missing data. Additionally, a late Row-cued Local Feature Aggregation (LFA) module embeds longitudinal continuity and transverse offset constraints into deep layers, effectively mitigating cross-row feature aliasing. During inference, a multi-output pipeline utilizes these predictions for reliable crop point filtering, center refinement, and navigation line fitting. Experiments show that the xyzrgb_exg input achieves an optimal balance between accuracy and conciseness. RACG-RandLA achieves a Test IoU of 0.8590, outperforming PointNet variants, and secures the highest Row Count Accuracy of 0.8761. Furthermore, it reduces lateral navigation jitter to 0.0241 m while maintaining an 85.04% success rate. Ultimately, the proposed method demonstrates a superior balance of semantic accuracy, structural consistency, and navigation stability, providing a robust perception solution for agricultural robots. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

17 pages, 2033 KB  
Article
Multi-Axle Reference and Temporal-Consistency Deep SVDD for EMU Traction Motor Bearing Anomaly Detection Using Field Vibration Data
by Qi Wu, Xiaomin Zhu, Zhikai Jia and Zhongkai Wang
Sensors 2026, 26(15), 4891; https://doi.org/10.3390/s26154891 - 3 Aug 2026
Viewed by 243
Abstract
Field vibration monitoring of EMU traction motor bearings is commonly constrained by weak or relative labels, fluctuations in operating conditions, and limited abnormal samples. Under these conditions, learning a normal boundary from a single bearing position may be unstable, and isolated score spikes [...] Read more.
Field vibration monitoring of EMU traction motor bearings is commonly constrained by weak or relative labels, fluctuations in operating conditions, and limited abnormal samples. Under these conditions, learning a normal boundary from a single bearing position may be unstable, and isolated score spikes may lead to unreliable alarms. To address these issues, this study proposes a multi-axle reference and temporal-consistency-enhanced Deep SVDD framework, termed MA-TC-Deep SVDD, for field anomaly detection of EMU traction motor bearings. Unlike closed-set fault diagnosis that requires known fault labels, the proposed framework focuses on identifying deviations from the stable operating regime. First, a compact 10-dimensional time-frequency representation is constructed from valid vibration segments. Second, stable samples from the target bearing position and screened stable samples from other monitored positions on the same EMU are organized as a multi-axle reference set for one-class normal-boundary learning. Third, feature recalibration, temporal-consistency regularization, reference-score standardization, causal smoothing, and consecutive-alarm judgment are incorporated to improve robustness against field disturbances. The anomaly-prior-guided health-state interpretation module is retained only as post hoc evidence for describing severity evolution and does not feed back into the anomaly detection threshold. Field data collected from an in-service EMU over D1–D5 are used for validation. The results show that bearing position 1 has low anomaly scores on D1–D2, exhibits transitional deviation on D3, and shows persistent state deviation on D4–D5, while the other monitored positions remain comparatively stable. Under the current weak-label evaluation protocol, MA-TC-Deep SVDD achieves higher average anomaly detection performance than the compared baseline methods, with AUC = 0.909, AP = 0.872, Precision = 0.887, Recall = 0.802, F1 = 0.843, and FAR = 0.047. These results indicate that the proposed framework can provide field anomaly-warning and severity-oriented interpretation under weak-label monitoring conditions. However, it should not be interpreted as a replacement for disassembly-confirmed fault-type diagnosis or remaining useful life prediction. Full article
(This article belongs to the Section Fault Diagnosis & Sensors)
Show Figures

Figure 1

27 pages, 8174 KB  
Article
A Multi-Feature Early Fusion Network with Domain-Specific Contrastive Representation and Attention Mechanism for Side-Scan Sonar Target Detection
by Junhui Zhu, Houpu Li, Shaofeng Bian, Xueshen Li, Lei Liu, Guojun Zhai and Ye Peng
Appl. Sci. 2026, 16(15), 7683; https://doi.org/10.3390/app16157683 - 2 Aug 2026
Viewed by 220
Abstract
Accurate target detection in side-scan sonar imagery is important for marine resource investigation, underwater infrastructure inspection, and maritime security. However, Side-scan sonar images are often affected by low contrast, acoustic speckle, weak target boundaries, and cluttered seabed backgrounds, which make target detection particularly [...] Read more.
Accurate target detection in side-scan sonar imagery is important for marine resource investigation, underwater infrastructure inspection, and maritime security. However, Side-scan sonar images are often affected by low contrast, acoustic speckle, weak target boundaries, and cluttered seabed backgrounds, which make target detection particularly challenging, especially under limited training data conditions. To improve the representation of sonar-specific structures, this study proposes a multi-feature early fusion detection network, referred to as MFEF-Det, for side-scan sonar target detection. The method combines multiple handcrafted features with data-driven representations to provide complementary information. In particular, a directional contrast core response feature (DCCR) is introduced to better emphasize the echo–shadow structure commonly observed in side-scan sonar imagery. An adaptive fusion strategy is then adopted to combine multiple feature maps before feeding them into a detection network, and an attention refinement module is further employed for complex scenes to improve the discrimination between target-related regions and cluttered backgrounds. Experiments were conducted on two publicly available sonar datasets. Experiments on the KLSG and SSS-Bottom datasets demonstrate that MFEF-Det achieves 0.933 ± 0.016 mAP@0.5 and 0.866 ± 0.052 mAP@0.5, respectively. The results indicate that the proposed feature representation can improve detection performance in both relatively clean and more challenging noisy scenes. These findings suggest that incorporating sonar-specific priors can be beneficial for side-scan sonar detection in marine survey and maritime-security applications. Full article
Show Figures

Figure 1

24 pages, 47100 KB  
Article
VSTF-Net: A Vision-Semantic and Target-Aware Fusion Framework for Ship Detection in Complex Maritime Sensing Scenarios
by Yinqing Peng, Xiulin Qiu, Yuhao Wu, Yuwang Yang, Yu Wang and Yuxin Wei
Remote Sens. 2026, 18(15), 2517; https://doi.org/10.3390/rs18152517 - 2 Aug 2026
Viewed by 343
Abstract
Ship detection in complex maritime and remote sensing scenes is of great importance for Earth observation, maritime surveillance, and intelligent ocean monitoring. However, existing methods often struggle with severe background interference, adverse weather conditions, and insufficient structural feature representation of ship targets. To [...] Read more.
Ship detection in complex maritime and remote sensing scenes is of great importance for Earth observation, maritime surveillance, and intelligent ocean monitoring. However, existing methods often struggle with severe background interference, adverse weather conditions, and insufficient structural feature representation of ship targets. To address these challenges, we propose VSTF-Net, a Vision-Semantic and Target-aware Fusion Framework for Ship Detection in Complex Maritime Sensing Scenarios. Specifically, to enhance high-level semantic representation under complex maritime environments, a Visual-Semantic Environment Modulation (VSEM) module is designed to introduce semantic priors extracted from a CLIP-pretrained visual encoder for adaptive feature modulation. To capture the elongated structural and directional characteristics commonly exhibited by ship targets in maritime and remote sensing scenes, a Multi-branch Asymmetric Dilated (MAD) module is developed to strengthen structural and directional feature modeling through asymmetric and dilated convolutions. Furthermore, to suppress high-frequency background interference in complex water environments, a Frequency-Adaptive Denoising (FAD) module is integrated to adaptively recalibrate frequency-domain features and improve the discriminability between ship targets and surrounding backgrounds. Experimental results show that the proposed method achieves 89.2% mAP@0.5 and 56.2% mAP@0.5:0.95 on the enhanced SeaShips dataset, with improvements of 1.9% and 2.4% over the baseline model, respectively. Further experiments on the HRSC2016 and MEIWVD datasets demonstrate the effectiveness and robustness of the proposed method across both remote sensing and maritime scenes. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

Back to TopTop