Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (522)

Search Parameters:
Keywords = awareness cues

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
34 pages, 4627 KB  
Article
Pyramid Target Perception Network with Efficient Context Modeling and Multi-Scale Cross-Attention for Infrared Small Target Detection
by Xinlu Zong, Zhenke Wang, Quan Wen and Hui Xu
Electronics 2026, 15(17), 3840; https://doi.org/10.3390/electronics15173840 - 26 Aug 2026
Abstract
Infrared small target detection (IRSTD) is a challenging task in intelligent infrared sensing and electronic imaging systems, because dim targets often occupy only a few pixels and are easily disturbed by clutter, noise, and low-contrast background structures. A practical detector should preserve pixel-level [...] Read more.
Infrared small target detection (IRSTD) is a challenging task in intelligent infrared sensing and electronic imaging systems, because dim targets often occupy only a few pixels and are easily disturbed by clutter, noise, and low-contrast background structures. A practical detector should preserve pixel-level target cues while suppressing target-like false responses. This paper proposes a Pyramid Target Perception Network (PTPN) for single-frame pixel-level IRSTD. The network integrates three complementary components: an Efficient Context Modeling (ECM) encoder employing 7 × 7 depthwise separable convolution for lightweight contextual feature extraction, a multi-scale target cross-attention (MTCA) module for hierarchical feature interaction, and a small-target feature pyramid network (STFPN) for target-preserving multi-scale aggregation. In addition, a physics-constrained loss (PCL) is introduced during training to regularize predictions according to infrared imaging characteristics, including point spread consistency, target-region relative intensity consistency, and signal-to-noise-ratio-aware separability. Experiments on IRSTD-1k, NUAA-SIRST, and NUDT-SIRST demonstrate that PTPN achieves IoU scores of 71.87%, 79.56%, and 86.47%, respectively, with 4.55M parameters, 4.96G FLOPs at an input resolution of 256 × 256, and an inference speed of 45.0 FPS. Although PTPN achieves competitive overall performance, it does not attain the highest IoU on NUDT-SIRST, indicating that pixel-level target-region estimation under complex scenes remains an area for further improvement. Overall, PTPN provides an effective balance between target localization, false-alarm suppression, and computational efficiency, supporting its potential application in AI-driven infrared image processing and intelligent electronic sensing systems. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

22 pages, 4118 KB  
Article
Edge-Geometry-Guided Deformable Detection for Sub-Millimeter Defects in Underwater Nuclear Component Inspection
by Jinkun Li, Lingyu Sun, Minglu Zhang, Chao Ma and Xinbao Li
Big Data Cogn. Comput. 2026, 10(9), 287; https://doi.org/10.3390/bdcc10090287 - 26 Aug 2026
Abstract
Accurate detection of sub-millimeter defects in reactor core-plate cotter-pin holes is essential for nuclear safety. However, underwater inspection images often suffer from low signal-to-noise ratios, weak boundary responses, and pseudo-edge interference, resulting in unstable localization of defects. Existing deformable and attention-based detectors remain [...] Read more.
Accurate detection of sub-millimeter defects in reactor core-plate cotter-pin holes is essential for nuclear safety. However, underwater inspection images often suffer from low signal-to-noise ratios, weak boundary responses, and pseudo-edge interference, resulting in unstable localization of defects. Existing deformable and attention-based detectors remain vulnerable to sampling drift and semantic–boundary inconsistency under such conditions. To address these challenges, an Edge-Geometry-Guided Deformable Detection Network (EGD-Net) is proposed for underwater defect detection. EGD-Net introduces an edge-geometry-constrained deformable sampling mechanism that embeds edge-confidence priors into deformable convolution to improve boundary-aware feature sampling. A cross-level semantic–geometric alignment strategy is designed to enhance the interaction between defect semantics and geometric boundary cues, while a top-down feedback recalibration mechanism improves multi-scale response consistency for weak defects. Experiments on the Core-Plate Pin-Hole Defect (CPHD) dataset demonstrate that EGD-Net achieves the highest AP@[0.5:0.95] on both datasets while maintaining competitive or superior Precision, Recall, and F1-score while reducing engineering center error under a fixed operating point. Performance across the two complementary domains suggests its robustness to variations between coupon images and practical underwater inspection scenes. These results indicate that EGD-Net provides a reliable solution for boundary-sensitive localization of underwater sub-millimeter defects in nuclear inspection. Full article
Show Figures

Figure 1

32 pages, 3582 KB  
Article
BSCNet: Boundary- and Scale-Consistent Mean Teacher for Semi-Supervised Building Change Detection in High-Resolution Remote Sensing Images
by Sujin Cai, Taizhi Lv, Xing Li, Chengyi Shi, Caifeng Wu, Xin Li, Linyang Li and Zhen Jia
Symmetry 2026, 18(9), 1428; https://doi.org/10.3390/sym18091428 - 26 Aug 2026
Abstract
Pixel-level annotation of bi-temporal high-resolution imagery is costly because annotators must distinguish genuine changes from pseudo-changes caused by illumination, seasonality, shadows, and residual misregistration. From a temporal-symmetry perspective, unchanged regions approximately preserve cross-temporal semantic correspondence, whereas genuine building changes introduce localized symmetry breaking [...] Read more.
Pixel-level annotation of bi-temporal high-resolution imagery is costly because annotators must distinguish genuine changes from pseudo-changes caused by illumination, seasonality, shadows, and residual misregistration. From a temporal-symmetry perspective, unchanged regions approximately preserve cross-temporal semantic correspondence, whereas genuine building changes introduce localized symmetry breaking between the two acquisition times. This paper presents BSCNet, a semi-supervised framework for binary building change detection that jointly models boundary-sensitive differences and scene-dependent scale preferences. A shared-weight MixTransformer extracts multi-level bi-temporal features. The Edge-Aware Optimization Module suppresses spatially invariant channel responses, enhances residual spatial cues, and predicts a Sobel-supervised edge map. The Parallel Selective Context Module aggregates depthwise-separable branches with different receptive fields and produces an image-level scale distribution. The Multi-scale Edge-Consistent Mean Teacher framework aligns the final prediction, intermediate edge representation, and scale-selection distribution between an exponential-moving-average teacher and the student. Experiments on WHU-CD and LEVIR-CD under 5%, 10%, and 20% labeled-data settings show consistent improvements over RCL, C2F-SemiCD, and CutMix-CD. With 5% labeled data, BSCNet achieves F1/IoU scores of 88.57%/79.49% on WHU-CD and 88.88%/79.98% on LEVIR-CD. An additional UAV-CD evaluation examines transfer to 0.06 m low-altitude UAV imagery containing both building and land changes; under 5% supervision, BSCNet obtains an F1/IoU of 68.07%/51.60%. Progressive ablations confirm complementary gains from the boundary, scale, and consistency components. Full article
(This article belongs to the Special Issue Symmetry/Asymmetry in Digital Image Processing)
Show Figures

Figure 1

37 pages, 3015 KB  
Article
Deepfake Detection via Frequency-Aware Vision Transformer and Bidirectional Cross-Attention Fusion with Post-Processing Robustness
by Wasin Alkishri, Shahid Kamal and Jabar Yousif
Information 2026, 17(9), 819; https://doi.org/10.3390/info17090819 - 26 Aug 2026
Abstract
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or [...] Read more.
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or high-level semantic representations of Vision Transformer; both have major drawbacks in effectively leveraging multi-domain forensic cues. This paper presents FAViT (Frequency-Aware Vision Transformer), a hybrid architecture capable of jointly utilizing spatial- and frequency-domain forensic information by the means of a bidirectional cross-attention fusion scheme. We use an 11-channel forensic tensor in each face image (including per-channel Fast Fourier Transform (FFT) magnitude maps, Discrete Wavelet Transform (DWT) sub-bands, channel noise residual maps, Sobel gradient magnitude and channels of Error Level Analysis (ELA)). A Frequency Branch CNN processes this multi-domain tensor and the original RGB image is encoded with a pretrained ViT-B/16 spatial branch. The two streams are combined through the bidirectional cross-attention which allows the model to localize both spatial and spectral manipulation artifacts. We also present an adversarial cleaning simulation pipeline which partitions the training process with five post-processing attack methods, namely GFPGAN neural face restoration, learned autoencoder cleaning, etc., to increase resistance to real-world forensic defenses. Tests of FaceForensics++ C23 (7926 images, consisting of four manipulation types) show that FAViT attains F1-score of 86.22, AUC-ROC of 94.26 and accuracy of 85.55 on the held-out test set. The strength analysis of 21 attack conditions shows that the max degradation in AUC is 30.3, with specific strengths in GFPGAN restoration (AUC = 98.51). Robustness is evaluated based on 21 post-processing attack cases that include JPEG compression, Gaussian blurring, down-sampling, and GFDGAN neural-based restoration; it should be noted that robustness against gradient-based adaptive attacks requires additional attention. Testing on the CIFAKE and Celeb-DF v2 datasets reveals some limitations of domain generalization. Full article
(This article belongs to the Special Issue Artificial Intelligence for Signal, Image and Video Processing)
Show Figures

Graphical abstract

29 pages, 16248 KB  
Article
Multimodal Instance Segmentation of Building Façade Damage with Frequency-Decoupled Feature Fusion
by Shenglin Xu, Zhengyuan Chen, Yibo Wang, Yanran Shi, Hao Lu, Haowei Gu and Ziqiang Sun
Buildings 2026, 16(17), 3396; https://doi.org/10.3390/buildings16173396 - 25 Aug 2026
Abstract
Building façade damage segmentation remains challenging because cracks, peeling, hollow-like areas, stains, and erosion often exhibit weak contrast, irregular boundaries, discontinuous structures, and substantial scale variations under complex surface textures and illumination conditions. To address these challenges, this study proposes MMDF-Net, a multimodal [...] Read more.
Building façade damage segmentation remains challenging because cracks, peeling, hollow-like areas, stains, and erosion often exhibit weak contrast, irregular boundaries, discontinuous structures, and substantial scale variations under complex surface textures and illumination conditions. To address these challenges, this study proposes MMDF-Net, a multimodal instance segmentation network for paired RGB–IR façade damage images. The network adopts dual RGB and infrared branches and performs multi-scale cross-modal interaction across P2–P5 levels. A Frequency-Decoupled Cross-Modal Bridge is designed to separately model low-frequency material responses and high-frequency local damage details, thereby enabling gated bidirectional exchange between the two modalities. A Crack Topology-Aware Encoder further enhances the continuity of thin, curved, branched, and locally discontinuous cracks through directional strip convolution and bending-aware modeling. In addition, a Multi-Scale Damage Decoding Pyramid integrates high-resolution boundary details, mid-level structural cues, and high-level semantic context to support stable mask prediction across different damage scales. Experiments were conducted on a custom-built RGB–IR building façade damage dataset containing 1500 paired image samples across five damage categories. MMDF-Net achieved 82.1% mask precision, 78.6% mask recall, 72.3% mask mAP50, and 65.6% mask mAP50–95, with 7.8 M parameters and 23.5 GFLOPs, outperforming representative RGB and RGB–IR segmentation baselines. These results indicate that task-oriented multimodal feature organization can improve façade damage segmentation without relying on excessive model expansion. Full article
Show Figures

Figure 1

19 pages, 822 KB  
Article
Digital Data Engagement and Health Behavior Diversity Among Smartwatch Users
by Fang-Wu Tung and Liang-Ming Jia
World 2026, 7(9), 143; https://doi.org/10.3390/world7090143 - 24 Aug 2026
Viewed by 165
Abstract
Smartwatches generate continuous personal health data, yet their value for self-care depends on users’ ability to engage with these data meaningfully. Using a cross-sectional survey of 838 adult smartwatch users in Taiwan, this study examined digital data engagement (DDE) as a data-specific construct [...] Read more.
Smartwatches generate continuous personal health data, yet their value for self-care depends on users’ ability to engage with these data meaningfully. Using a cross-sectional survey of 838 adult smartwatch users in Taiwan, this study examined digital data engagement (DDE) as a data-specific construct linking smartwatch use to everyday self-care. DDE captured monitoring, interpretation, verification, and goal-oriented use of smartwatch-generated data, while health behavior diversity (HBD) represented the breadth of routine self-care practices across preventive, mind–body, physical activity/exercise, and restorative or expressive domains. The DDE-centered model showed that DDE was positively associated with health motivation, body awareness, and HBD. Health motivation showed the strongest association with HBD, whereas body awareness was not directly associated with HBD, suggesting that bodily cue awareness alone may not broaden self-care repertoires. HBD analysis further showed that smartwatch-enabled self-care was not merely exercise-oriented; preventive self-care had the highest item-adjusted coverage. Gender differences appeared mainly in domain composition, whereas health concern profiles were associated with both repertoire breadth and domain composition. These findings contribute a data-engagement perspective to smartwatch research and introduce HBD as a multidomain lens for evaluating everyday digital self-care beyond device access, tracking frequency, or single behavior change. Full article
(This article belongs to the Section Health, Population, and Crisis Systems)
Show Figures

Figure 1

18 pages, 400 KB  
Article
Effects of a Reflection-Integrated CICARE Program on Empathy and Clinical Judgment in Nursing Students: A Quasi-Experimental Study
by Qingqing Feng, Min Liu, Tang Tang and Lulu Zhang
Nurs. Rep. 2026, 16(9), 296; https://doi.org/10.3390/nursrep16090296 - 24 Aug 2026
Viewed by 78
Abstract
Background/Objectives: Communication training in undergraduate nursing skills courses is often overshadowed by procedural skill acquisition, leaving limited opportunities for students to develop patient-centered communication skills and engage in reflective learning. This study evaluated a reflection-integrated CICARE communication program embedded in a Fundamentals of [...] Read more.
Background/Objectives: Communication training in undergraduate nursing skills courses is often overshadowed by procedural skill acquisition, leaving limited opportunities for students to develop patient-centered communication skills and engage in reflective learning. This study evaluated a reflection-integrated CICARE communication program embedded in a Fundamentals of Nursing skills course and explored students’ learning experiences. Methods: A quasi-experimental study with an embedded qualitative component was conducted among second-year undergraduate nursing students in China. The quantitative component used a non-equivalent control group design, while reflective journals from the Intervention Group provided qualitative data on students’ learning experiences. Two intact classes were assigned to the Intervention Group (n = 65), and two intact classes were assigned to the Control Group (n = 66). The Intervention Group received an 18-week CICARE-based communication program integrated into skills training, including scenario-based role-play, emotional cue recognition, demonstration videos, feedback, and reflective journals. The Control Group received conventional skills training. Outcomes included communication competence, empathy, professional identity, clinical judgment, and course performance. Analysis of covariance was used to compare post-intervention outcomes after adjustment for baseline scores. Reflective journals from the Intervention Group were analyzed thematically. Results: After adjustment for baseline scores, the Intervention Group showed significantly higher communication competence than the Control Group (adjusted mean difference = 2.699, 95% CI: 0.108–5.291, p = 0.041). Empathy was also higher in the Intervention Group (adjusted mean difference = 3.796, 95% CI: 0.538–7.053, p = 0.023). The between-group difference in professional identity was not statistically significant (p = 0.073). The Intervention Group achieved higher total clinical judgment scores than the Control Group (31.52 ± 5.50 vs. 27.20 ± 6.45, p < 0.001), with significant differences in noticing, interpreting, and reflecting. Course performance was also higher in the Intervention Group. In their reflective journals, students described greater patient-centered awareness, attention to patients’ emotional needs, professional reflection, and motivation for self-directed learning. Conclusions: Integrating CICARE-based communication training with reflective learning may strengthen communication competence, empathy, clinical judgment, and course performance in undergraduate nursing skills education. Further studies with longer follow-up are needed to determine whether the observed differences persist over time and whether professional identity and action-oriented clinical responses improve with longer educational exposure. Registration: This study was not registered. Full article
(This article belongs to the Special Issue Advancing Nursing Practice Through Innovative Education)
Show Figures

Figure 1

26 pages, 3199 KB  
Article
MCSwin-YOLOv8: Multi-Scale Feature Learning for Maritime Ship Detection
by Yuqing Ren, Guohao Wen and Yingbang Huang
Appl. Sci. 2026, 16(17), 8421; https://doi.org/10.3390/app16178421 - 24 Aug 2026
Viewed by 165
Abstract
Maritime ship detection remains challenging because of large scale variations, high inter-class visual similarity, weak target boundaries, and complex maritime backgrounds. This study proposes MCSwin-YOLOv8, an enhanced YOLOv8-based detector that combines three complementary architectural designs. First, a re-parameterizable multi-scale convolutional backbone, named RepMCSwin, [...] Read more.
Maritime ship detection remains challenging because of large scale variations, high inter-class visual similarity, weak target boundaries, and complex maritime backgrounds. This study proposes MCSwin-YOLOv8, an enhanced YOLOv8-based detector that combines three complementary architectural designs. First, a re-parameterizable multi-scale convolutional backbone, named RepMCSwin, is introduced to extract scale-aware semantic information and fine-grained boundary cues. Unlike the standard Swin Transformer, the MCSwin block does not use window-based self-attention but adopts cascaded multi-scale convolutions and residual feature transformation. Second, a Multi-Feature Parallel Convolutional Block Attention Module (MFPCBAM) is developed to compute channel and spatial attention in parallel, thereby preserving weak ship features while suppressing irrelevant background responses. Third, a Modified Generalized Feature Pyramid Network (MGFPN) is constructed to improve cross-level feature interaction and retain high-resolution spatial information through an additional 160 × 160 prediction branch. Experiments were conducted on the public SeaShips dataset and a private infrared maritime ship dataset. MCSwin-YOLOv8 achieved an F1-score of 94.4%, mAP@0.5 of 97.4%, and mAP@0.5:0.95 of 75.1% on SeaShips. On the infrared dataset, the corresponding results were 91.2%, 94.1%, and 66.9%, respectively. Compared with the YOLOv8 baseline, mAP@0.5 increased by 1.4 and 3.0 percentage points on the two datasets. These accuracy gains were accompanied by an increase in model complexity from 11.12 M to 19.29 M parameters and from 28.5 G to 56.7 G FLOPs, indicating an accuracy–complexity trade-off that requires further runtime evaluation. Full article
(This article belongs to the Section Marine Science and Engineering)
Show Figures

Figure 1

30 pages, 3682 KB  
Article
Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity
by Taher M. Ghazal, Fareeha Anwar, Sumaia Mohammed Al-Ghuribi, Amjed A. Ahmed, Ali Hamzah Najim, Omar Almomani, Prabu Pachiyannan and Hesham A. Sakr
Math. Comput. Appl. 2026, 31(5), 168; https://doi.org/10.3390/mca31050168 - 23 Aug 2026
Viewed by 150
Abstract
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security [...] Read more.
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security alerts, increasing susceptibility to social engineering, identity theft, and data breaches. This paper presents SECURE-A2RC, a rubric-constrained, Arabic-aware large language model framework designed to deliver scalable, interpretable, and culturally relevant cybersecurity education. The framework comprises two coupled components. The first, the Arabic-Aware Secure Communication Encoder (A-SCE), employs an instruction-tuned LLM to produce multidimensional encodings that capture three learner competencies: security intent comprehension; linguistic deception cue recognition encompassing urgency, authority impersonation, and incentive framing; and action-critical translation fidelity across Arabic dialectal registers and Arabic–English bilingual contexts. The second, the Rubric-Constrained Adaptive Feedback Generator (RCAFG), translates A-SCE encodings into personalized, expert-aligned instructional feedback and proficiency-calibrated adaptive tasks, ensuring pedagogical consistency, security correctness, and dialect awareness throughout the learning cycle. The framework is evaluated on three domain-relevant corpora: the English–Arabic Parallel Phishing Email Corpus, the Open MalSec dataset, and the Arabic Spam and Ham Tweets dataset. SECURE-A2RC achieves a 31% improvement in phishing identification accuracy and a 26% reduction in action-critical translation errors compared to conventional awareness materials. A comparative evaluation against SERENA, a Multi-Agent LLM, and the Arabic Multitask Learning Model confirms consistent superiority across detection accuracy, F1-score, dialectal robustness, and educational effectiveness metrics, affirming rubric-constrained LLM integration as a viable approach to equitable multilingual cybersecurity education. Full article
Show Figures

Figure 1

33 pages, 9024 KB  
Article
Motion-Guided Dynamic-Graph Construction with Kinematic-Aware Transformer for Skeleton Action Recognition
by Kabul Khudaybergenov and Avazjon Marakhimov
Appl. Sci. 2026, 16(17), 8382; https://doi.org/10.3390/app16178382 - 23 Aug 2026
Viewed by 188
Abstract
Skeleton-based action recognition has attracted considerable research interest because skeleton data are inherently robust to illumination changes, viewpoint variation, background clutter, and camera motion. Nevertheless, extracting informative representations from skeleton sequences remains a challenging problem, as it requires capturing both the spatial co-occurrence [...] Read more.
Skeleton-based action recognition has attracted considerable research interest because skeleton data are inherently robust to illumination changes, viewpoint variation, background clutter, and camera motion. Nevertheless, extracting informative representations from skeleton sequences remains a challenging problem, as it requires capturing both the spatial co-occurrence patterns among body joints and the fine-grained kinematic cues that distinguish different actions. In this paper, we propose a single-stream architecture that constructs an action-specific skeleton graph directly from motion and processes it with a kinematic-aware Transformer. Rather than relying on a fixed skeleton topology, a motion-guided dynamic-graph construction module infers a per-frame adjacency matrix from short-term motion cues through a differentiable edge predictor and Gumbel-Softmax sparsification, allowing the model to discover action-driven connections between distant joints that lack direct bone connectivity (e.g., coordinated hand motion during clapping). Each joint is described by kinematic node features that combine its 3D position, instantaneous velocity, and limb-angle encodings within a single descriptor, so that both motion dynamics and higher-order limb configurations are available to the spatial encoder from the outset. A graph-attention network (GAT) encodes the spatial configuration of every frame over the learned graph, and the resulting sequence of frame descriptors is processed by a Transformer encoder that models long-range temporal dependencies; a learnable classification token aggregates the sequence, and a multi-layer perceptron (MLP) produces the final action classification. The entire model is trained end-to-end from action labels alone. We conduct a comprehensive ablation study and evaluate the proposed method on the large-scale NTU RGB+D 60 and NTU RGB+D 120 benchmarks, where the results demonstrate that our approach achieves competitive performance compared to state-of-the-art architectures. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

33 pages, 4671 KB  
Article
Saliency-Guided RT-DETR for Multi-Class Detection in Processing Tomato Sorting
by Xingyu Jiang, Yingjie Zhang, Ximei Wei, Xia Peng and Weitao Chen
Agriculture 2026, 16(17), 1805; https://doi.org/10.3390/agriculture16171805 - 22 Aug 2026
Viewed by 173
Abstract
This study considers multi-class visual detection for processing tomato sorting under controlled laboratory conditions. In densely arranged images, occlusion and visual similarity can weaken the local boundary and texture cues needed to distinguish ripe tomatoes, unripe tomatoes, defective tomatoes, and soil clods. We [...] Read more.
This study considers multi-class visual detection for processing tomato sorting under controlled laboratory conditions. In densely arranged images, occlusion and visual similarity can weaken the local boundary and texture cues needed to distinguish ripe tomatoes, unripe tomatoes, defective tomatoes, and soil clods. We propose SG-RTDETR, an RT-DETRv2 adaptation that combines detail-preserving downsampling, saliency-guided token encoding with residual spatial refill, context-aware feature organization, and adaptive cross-scale fusion. A four-class dataset was constructed from 1000 multi-object images and 400 single-object images used for supplementary representation learning. Across three independent training runs under controlled laboratory evaluation, SG-RTDETR achieved 87.8±0.4% mAP50:95, 92.4±0.4% mAP50, and 95.5±0.3% mAR50:95 (mean ± sample standard deviation). Relative to RT-DETRv2, the mean mAP50:95 increased by 3.4 percentage points, while total FLOPs remained at 61.17 G and forward-pass inference throughput decreased from 110.5 to 103.6 FPS. These results indicate an accuracy-oriented trade-off under the laboratory evaluation protocol; validation under realistic postharvest sorting conditions remains necessary. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

20 pages, 1191 KB  
Article
ChoreDiffusion: Beat-Aware Diffusion for Music-to-Dance Generation
by Yufei Gao, Qian Wu, Shuliang Zhu, Keren He, Wei Weng and Jinjia Zhou
Information 2026, 17(8), 808; https://doi.org/10.3390/info17080808 - 21 Aug 2026
Viewed by 182
Abstract
Music-to-dance generation requires precisely aligning movement dynamics with musical rhythm, yet existing methods rely on shallow conditioning or auxiliary beat-alignment objectives that fail to establish stable beat–motion correspondences. We present ChoreDiffusion, a diffusion-based framework that integrates explicit beat guidance directly into the denoising [...] Read more.
Music-to-dance generation requires precisely aligning movement dynamics with musical rhythm, yet existing methods rely on shallow conditioning or auxiliary beat-alignment objectives that fail to establish stable beat–motion correspondences. We present ChoreDiffusion, a diffusion-based framework that integrates explicit beat guidance directly into the denoising process. Central to our approach is a beat-enhanced cross-modal attention mechanism that injects beat-salience cues at every refinement step, promoting fine-grained synchronization beyond the reach of conventional conditioning pipelines. To support multiple dance styles within a unified model, we incorporate lightweight low-rank adaptation (LoRA) modules that encode style-specific motion signatures with only a small set of additional parameters per style, and a three-stage progressive curriculum stabilizes the joint learning of rhythmic alignment and stylistic expressivity. Experiments on two public multi-style dance benchmarks (AIST++ and FineDance) show that ChoreDiffusion achieves the lowest FID values among the compared generation methods on both benchmarks, while maintaining competitive rhythm alignment and multi-style controllability. These results indicate that embedding beat-aware guidance during generation, rather than applying it afterwards, is an effective route toward human-like musicality in music-driven choreography. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

23 pages, 11731 KB  
Article
A Physics-Guided Raw-Dominant Gated Fusion Method for Fine-Grained Bearing Fault Diagnosis
by Chuanbo Wu, Guoao Jiao, Yongdi Zhang, Kangning Jin, Zihang Zhang and Zeming Li
Appl. Sci. 2026, 16(16), 8337; https://doi.org/10.3390/app16168337 - 21 Aug 2026
Viewed by 205
Abstract
Fine-grained bearing condition diagnosis is challenging because different bearing states within the same fault location often exhibit similar fault-characteristic-frequency responses, making them difficult to distinguish using envelope-spectrum information alone. To address this problem, a physics-guided raw-dominant gated fusion network (PG-RDGFN) is proposed for [...] Read more.
Fine-grained bearing condition diagnosis is challenging because different bearing states within the same fault location often exhibit similar fault-characteristic-frequency responses, making them difficult to distinguish using envelope-spectrum information alone. To address this problem, a physics-guided raw-dominant gated fusion network (PG-RDGFN) is proposed for fine-grained bearing fault diagnosis. In the proposed framework, the raw vibration signal is retained as the dominant information source, while the envelope spectrum provides complementary fault-modulation evidence. A physics-aware descriptor derived from bearing characteristic-frequency responses is incorporated into a sample-wise gating mechanism to adaptively regulate the contribution of the envelope-spectrum features. Distinct from conventional direct multi-branch fusion or physics-informed schemes that mainly use physical knowledge as an auxiliary input or regularization constraint, PG-RDGFN uses mechanism-derived physical confidence to regulate how much auxiliary envelope-spectrum evidence participates in the fusion, rather than directly using the physical prior as a fine-grained classification cue. Meanwhile, a physical-consistency loss constrains the learned gate using a scalar physical-confidence target, and a raw-branch auxiliary loss preserves the discriminative capability of the dominant raw representation. Experiments are conducted on an eight-class diagnosis task constructed from the Paderborn University bearing dataset. Compared with SVM, MLP, 1D-CNN, CNN-LSTM, and TCN, the PG-RDGFN achieves the highest accuracy of 98.87%. Ablation and gate-consistency analyses further verify the effectiveness of the envelope branch, gated fusion, physics guidance, and auxiliary supervision. These results demonstrate that PG-RDGFN provides an accurate and physically interpretable solution for fine-grained bearing condition diagnosis. Full article
Show Figures

Figure 1

28 pages, 12985 KB  
Article
GCF-Net: Stage-Aligned Optical–Elevation Fusion for Aerial Remote Sensing Semantic Segmentation
by Yifan Yu, Song Deng, Yang Yang and Fan Min
Remote Sens. 2026, 18(16), 2827; https://doi.org/10.3390/rs18162827 - 20 Aug 2026
Viewed by 199
Abstract
Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit [...] Read more.
Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit cross-modal structural inconsistency, frequency–spatial response imbalance, and decoder-stage structural attenuation. To address these challenges, we propose GCF-Net, a stage-aligned optical–elevation fusion network that matches different cross-modal processing objectives to the evolving representation states of the encoder–decoder pipeline. A Structure-Guided Cross-Modal Correction Module first performs structure-conditioned correction of modality-specific features before fusion. A Frequency–Spatial Cross-Modal Fusion Module then constructs joint representations through bounded cross-modal magnitude conditioning, frequency-to-spatial reconstruction, and spatial recalibration. During decoding, a Geometry-Aware Cross-Scale Refinement Module reintroduces elevation-derived structural guidance into multiscale fused features. Experiments on ISPRS Vaihingen, ISPRS Potsdam, and MMHunan yield mIoU scores of 72.37%, 75.38%, and 52.23%, respectively, achieving the highest mIoU among the evaluated unimodal, multimodal, and SAM-based methods under the unified protocol. Ablation and replacement experiments verify the complementary roles of the three stage-specific components, while sensitivity, elevation perturbation, and complexity analyses indicate architectural flexibility, tolerance to moderate elevation degradation, and a balanced accuracy–efficiency trade-off. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

33 pages, 3739 KB  
Article
SEM-PDPL: Semantic Exposure Graphs for Privacy-Law-Informed Risk Assessment of Public Social-Media Data
by Heba Ismail
Information 2026, 17(8), 803; https://doi.org/10.3390/info17080803 - 20 Aug 2026
Viewed by 218
Abstract
Public social-media content often contains self-disclosed personal attributes that appear low-risk in isolation but become privacy-relevant when linked across posts, platform accounts, or user-level traces. Existing research has advanced privacy-sensitive content detection, de-anonymization analysis, social-media research ethics, and privacy-compliance workflows; however, limited work [...] Read more.
Public social-media content often contains self-disclosed personal attributes that appear low-risk in isolation but become privacy-relevant when linked across posts, platform accounts, or user-level traces. Existing research has advanced privacy-sensitive content detection, de-anonymization analysis, social-media research ethics, and privacy-compliance workflows; however, limited work operationalizes how personal-data disclosures combine structurally and how these structures can be translated into auditable governance actions. This paper proposes SEM-PDPL, a computational, privacy-law-informed risk-assessment framework for modeling public social-media exposure as semantic exposure graphs and mapping graph patterns to controls aligned with the United Arab Emirates Personal Data Protection Law (PDPL) and compatible with GDPR principles. SEM-PDPL combines governance scoping; a PDPL-informed disclosure taxonomy; hybrid extraction using rule-based methods; named-entity recognition; fine-tuned BERT; and schema-constrained large language model annotation, followed by graph construction at post, platform, corpus, and persona levels. The framework is evaluated on a synthetic multi-platform corpus of 1095 posts generated for 150 personas across 290 platform accounts. Results show that, within this controlled synthetic corpus, fine-tuned BERT provides the strongest extraction performance among six evaluated methods, achieving a macro-F1 of 0.975. Graph analysis shows that exposure density increases with aggregation, rising from 0.275 at post level to 1.000 at corpus level, and from 0.859 at platform level to 0.967 at persona level. Across all graph resolutions, quasi-identifiers emerge as the dominant weighted-degree and betweenness node, indicating that ordinary location, employer, school, and demographic cues often function as bridges connecting sensitive categories such as health and biometric data to identifying information. These findings indicate that, within this controlled corpus, privacy risk in public social-media data is not only attribute-based but also structurally graph-shaped. SEM-PDPL contributes an explainable and reproducible framework for identifying exposure hubs, sensitive bridges, and aggregation risks before applying masking, minimization, exclusion, retention, or review controls. The framework does not automate legal compliance; rather, it provides evidence-based decision support for privacy-aware social-media analytics. Full article
(This article belongs to the Special Issue Semantic Networks for Social Media and Policy Insights)
Show Figures

Figure 1

Back to TopTop