Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (784)

Search Parameters:
Keywords = scene-referred

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
33 pages, 10867 KB  
Article
Object-Centric 2D-to-3D Pipeline for Interior-Design Visualization: Reference-Free Asset Evaluation and a Structured3D Scene-Level Benchmark
by Dan Toderici, Tiberiu-Gabriel Rodanciuc, George-Alexandru Micu, Răzvan Rughiniș, Sergiu-Rareș Lupșa and Dinu Țurcanu
Electronics 2026, 15(15), 3295; https://doi.org/10.3390/electronics15153295 - 26 Jul 2026
Abstract
This study presents a modular AI-assisted workflow for converting single 2D interior images into textured 3D assets and for evaluating those assets when ground-truth 3D meshes are unavailable. The proposed pipeline combines object detection, instance isolation, monocular-depth estimation, image-to-3D generation, texture synthesis, mesh [...] Read more.
This study presents a modular AI-assisted workflow for converting single 2D interior images into textured 3D assets and for evaluating those assets when ground-truth 3D meshes are unavailable. The proposed pipeline combines object detection, instance isolation, monocular-depth estimation, image-to-3D generation, texture synthesis, mesh export, and cloud-based execution to support early-stage interior-design and real-estate visualization tasks. A reference-free validation protocol is introduced, based on rendered multi-view comparisons, silhouette Intersection-over-Union, automated captioning, and multimodal embedding similarity, and is complemented by a composite validation framework that benchmarks reconstructed scenes against 200 panoramic indoor scenes from the Structured3D dataset using Hungarian-matched placement, size, recall, and relative-distance metrics. The workflow was implemented and tested using contemporary computer-vision and generative 3D components, with Hunyuan3D 2.0 used as the main reconstruction model. Proof-of-concept experiments on a representative corpus of 178 synthetically generated single-object images spanning a range of interior furniture categories show comparable silhouette IoU for textured and non-textured outputs and indicate that texture-preserving renderings improve visual and semantic similarity scores across CLIP-based evaluations. The 200-scene dataset evaluation reveals stable spatial localization (placement error ≈ 1.18 m, relative-distance error ≈ 0.54 m) alongside systematic over-prediction and size-calibration errors. Beyond the applied pipeline, the study contributes a reference-free, ground-truth-free protocol for 3D-asset evaluation and a first quantified account of where object-centric single-image reconstruction is reliable—spatial placement—and where it is not—object scale and spurious detection—at interior-scene scale. The results demonstrate the feasibility of integrating perception, 3D reconstruction, semantic assessment, and scalable deployment into a single applied pipeline, while remaining proof-of-concept and requiring extension to larger object and scene corpora, baselines, real-photograph evaluation, and human-centered assessment before broad claims about general interior-scene reconstruction can be made. Full article
(This article belongs to the Special Issue Advances in 3D Computer Vision and 3D Data Processing)
Show Figures

Figure 1

21 pages, 1336 KB  
Article
Geometry-Guided Diffusion SAR Point Cloud Denoising
by Chengwei Zhang, Tao Jiang, Xinhao Xu, Wenjie Li, Fubo Zhang and Longyong Chen
Remote Sens. 2026, 18(15), 2458; https://doi.org/10.3390/rs18152458 - 26 Jul 2026
Abstract
Three-dimensional synthetic aperture radar (SAR) point clouds provide valuable geometric observations of urban scenes, but they often suffer from severe noise and layer-like artifacts caused by the low signal-to-noise ratio and tomographic imaging mechanism. These degradations make SAR point cloud denoising significantly more [...] Read more.
Three-dimensional synthetic aperture radar (SAR) point clouds provide valuable geometric observations of urban scenes, but they often suffer from severe noise and layer-like artifacts caused by the low signal-to-noise ratio and tomographic imaging mechanism. These degradations make SAR point cloud denoising significantly more challenging than conventional LiDAR point cloud denoising. In this paper, we propose a Geometry-guided Diffusion SAR Point Cloud Denoising (GDSD) framework to recover geometrically coherent building surfaces from noisy SAR point clouds.The key idea is to exploit relatively clean LiDAR point clouds as geometry priors while avoiding the need for paired SAR–LiDAR supervision or clean SAR ground truth. Specifically, we introduce a Forward Gaussian Noising Process to disrupt the intrinsic layer-like artifacts of SAR point clouds and reduce the input-level discrepancy between SAR and LiDAR domains. We further design a geometry prototype-based alignment module that projects SAR and LiDAR bottleneck features into a shared LiDAR-dominated latent space, enabling geometry-aware conditional reverse diffusion. A DiT-3D-based denoising network is then trained with LiDAR-domain diffusion supervision and applied to SAR point clouds using the aligned SAR geometry condition. To evaluate the proposed method, we construct a SAR point cloud denoising benchmark based on the MV3DSAR dataset with CAD-derived reference surfaces. Experimental results show that GDSD significantly improves the quality of noisy SAR point clouds and clearly outperforms the previous conventional LiDAR point cloud denoising baseline, producing more continuous and geometrically coherent SAR building point clouds. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

31 pages, 10622 KB  
Article
UAS-Validated Comparison of Sentinel-2 Shoreline Extraction Techniques for Large-Lake Coastal Mapping
by Mohamed M. Elmeligy, Ahmed El-Rabbany, Saad Mesbah Abdelrahman, Mohamed Mohasseb, Mahmoud A. Hassaan and Hamed Majidiyan
Technologies 2026, 14(8), 459; https://doi.org/10.3390/technologies14080459 (registering DOI) - 25 Jul 2026
Abstract
Reliable assessment of shorelines extracted from medium-resolution satellite imagery requires independent high-resolution reference data and statistical methods that account for spatial dependence. This study compared three conventional analyst-assisted shoreline-extraction workflows—histogram thresholding, band ratio, and the Normalised Difference Water Index (NDWI)—at Coronation Park, Lake [...] Read more.
Reliable assessment of shorelines extracted from medium-resolution satellite imagery requires independent high-resolution reference data and statistical methods that account for spatial dependence. This study compared three conventional analyst-assisted shoreline-extraction workflows—histogram thresholding, band ratio, and the Normalised Difference Water Index (NDWI)—at Coronation Park, Lake Ontario, Canada, using Sentinel-2 Level-2A imagery. A manually digitised shoreline derived from a UAV-based orthomosaic acquired approximately 27 h before the Sentinel-2 scene served as the independent reference. The UAV-based reference and each Sentinel-2-derived shoreline were divided into 31 ordered segments. For each Sentinel-2-derived segment midpoint, the shortest planar Euclidean distance to the nearest UAV-based reference midpoint was calculated and used to derive mean absolute error (MAE) and root mean square error (RMSE). Residual spatial autocorrelation was assessed using Moran’s I with 9999 permutations. Because the paired differences departed from normality, the Friedman test was treated as the primary overall comparison, while contiguous spatial-block permutation tests across block sizes of two to eight shoreline locations assessed robustness to local spatial dependence. NDWI achieved the highest positional agreement (MAE = 5.645 m; RMSE = 6.429 m), followed by band ratio (MAE = 14.303 m; RMSE = 14.797 m) and histogram thresholding (MAE = 26.167 m; RMSE = 26.910 m). Significant positive residual spatial autocorrelation was identified for all three methods (Moran’s I = 0.587–0.832, all p < 0.001). The Friedman test confirmed a significant extraction-method effect, χ2(2) = 49.226, p < 0.001, Kendall’s W = 0.794, and the effect remained significant across all tested spatial-block sizes, with empirical p-values ranging from 0.000007 to 0.004630. Among the three conventional methods tested at this large-lake site, NDWI provided the highest positional agreement and therefore offers a defensible baseline for evaluating future Sentinel-2 image-enhancement approaches. Full article
Show Figures

Figure 1

20 pages, 5403 KB  
Article
TCM-CR: Multi-Temporal SAR–Optical Cloud Removal with a Reference Image and Gated Bounded Residual
by Xianjian Shi, Jiefang Zheng, Lilong Liu, Lv Zhou and Xin Bao
Remote Sens. 2026, 18(15), 2443; https://doi.org/10.3390/rs18152443 - 23 Jul 2026
Viewed by 181
Abstract
Cloud removal is an indispensable preprocessing step in optical remote sensing. Reconstructing cloud-free imagery by combining multi-temporal optical observations with cloud-penetrating synthetic aperture radar (SAR) has become a mainstream approach. However, the existing studies mostly adopt simple composites, such as per-pixel least-cloudy selection [...] Read more.
Cloud removal is an indispensable preprocessing step in optical remote sensing. Reconstructing cloud-free imagery by combining multi-temporal optical observations with cloud-penetrating synthetic aperture radar (SAR) has become a mainstream approach. However, the existing studies mostly adopt simple composites, such as per-pixel least-cloudy selection or the temporal median, as baselines, and average accuracy metrics over entire scenes; together, these two practices may overstate the true gains of deep-learning methods. This paper proposes a temporal cross-modal cloud removal method (TCM-CR). In a multi-temporal sequence, the acquisition with the lowest cloud fraction retains true surface reflectance at its cloud-free pixels and is itself a high-accuracy baseline. TCM-CR exploits this baseline in two ways. First, on clear and light inputs, cloud-free pixels are taken unchanged from the reference image, so the true reflectance is preserved without loss, independent of training. Second, only cloud-covered pixels receive a bounded correction, in which SAR supplies the surface structure beneath clouds and multi-temporal observations are integrated along time while suppressing heavily clouded acquisitions. Experiments on the SEN12MS-CR-TS dataset show that TCM-CR maintains accuracy on par with the reference image on clear and light samples and improves the peak signal-to-noise ratio on heavy samples by 7.93 dB. In a cross-region experiment where one region is excluded from training entirely and used only for testing, heavy samples still improve by 7.27 dB. Full article
(This article belongs to the Special Issue Advances in Multi-Source Remote Sensing Data Fusion and Analysis)
Show Figures

Figure 1

31 pages, 15220 KB  
Article
SRGFormer: Semantic Role-Guided Graph Reasoning for Referring Remote Sensing Image Segmentation
by Libang Liu, Jianxiang Li, Yaqin Li, Cao Yuan, Lili Fan, Xinyu Xiong and Wei Huang
Sensors 2026, 26(14), 4657; https://doi.org/10.3390/s26144657 - 22 Jul 2026
Viewed by 157
Abstract
Referring remote sensing image segmentation (RRSIS) aims to segment a target instance from remote sensing imagery according to a natural-language expression. It provides a flexible way to retrieve and localize specific objects in remote sensing scenes, benefiting intelligent Earth observation applications. Although existing [...] Read more.
Referring remote sensing image segmentation (RRSIS) aims to segment a target instance from remote sensing imagery according to a natural-language expression. It provides a flexible way to retrieve and localize specific objects in remote sensing scenes, benefiting intelligent Earth observation applications. Although existing methods have achieved promising progress by strengthening vision–language alignment, most of them still represent the expression as a holistic language feature and rely on convolution-dominated decoding for mask prediction. Such a paradigm tends to entangle target category, inter-object relation, and spatial position cues, making it difficult to distinguish the intended instance from multiple same-class distractors in complex remote sensing scenes. To address this limitation, we propose SRGFormer, a graph reasoning framework for RRSIS. Specifically, a semantic role decomposition (SRD) module decomposes the referring expression into target, relation, and position semantics, providing explicit linguistic priors for instance-level localization. Guided by the decomposed relation semantics, a semantic-relational graph transformer (SRGT) performs relation-aware graph reasoning over fused multi-scale visual features, enabling long-range dependency modeling among spatially distributed candidate instances. Furthermore, a progressive mask refinement (PMR) module continuously injects the decomposed semantic priors into semantic modulation, query initialization, and iterative mask decoding, thereby alleviating semantic fading during mask generation. Extensive experiments demonstrate that SRGFormer achieves substantial improvements on RefSegRS, attaining 66.08% mIoU and 76.93% oIoU (surpassing the prior state of the art by 3.96% and 2.83%, respectively) along with a notable 15.95% gain in Pr@0.7. Experiments on the additional RRSIS-D benchmark further demonstrate the general applicability of our approach, where SRGFormer maintains competitive performance (65.87% mIoU and 24.61% Pr@0.9) against existing methods. These results demonstrate that the proposed framework improves target localization and fine-grained mask prediction in complex remote sensing scenes. Full article
Show Figures

Figure 1

23 pages, 19255 KB  
Article
CLIFF: A Multi-Modal Remote Sensing Model for Geological Hazard Monitoring Based on Bitemporal UAV Images
by Quanxi Zhou, Qianxiao Su, Xinran Wei, Wencan Mao, Yili Ren, Yunfei Chen, Jianzhong Bi, Mingjun Zhao and Manabu Tsukada
Remote Sens. 2026, 18(14), 2432; https://doi.org/10.3390/rs18142432 - 22 Jul 2026
Viewed by 241
Abstract
UAV-based remote sensing excels in rapid response, high timeliness, simple operation, and high degrees of automation, and has been widely applied for geological hazard monitoring. Deep learning methods based on unitemporal UAV images can only analyze the static appearance of a scene, while [...] Read more.
UAV-based remote sensing excels in rapid response, high timeliness, simple operation, and high degrees of automation, and has been widely applied for geological hazard monitoring. Deep learning methods based on unitemporal UAV images can only analyze the static appearance of a scene, while bitemporal change detection can capture the dynamic evolution of hazards; however, due to diverse geological landforms and topography, environmental noises such as vegetation cover, and dynamic weather conditions, change detection of geological hazards from UAV images based on traditional deep learning technology is not always effective. Therefore, there is an urgent need to utilize large vision-language models (LVLMs) to further improve the accuracy and robustness of the change detection model. Motivated by this, this paper proposes a novel remote sensing model for geological hazard monitoring, referred to as CLIFF (CLIP-BIT-EfficientNet), based on the multi-modal LVLM Contrastive Language–Image Pre-training (CLIP), the change detection network Bitemporal Image Transformer (BIT), and the classification network EfficientNet, along with corresponding datasets and model fine-tuning strategies. The proposed transfer fusion module bridges the CLIFF and BIT networks by aligning their feature distributions and dimensions, allowing the general knowledge of the LVLM and the task-specific knowledge of the learnable branch to reinforce each other. Furthermore, this integrated pipeline addresses the scarcity of labeled hazard data by allowing the BIT to train on larger public datasets, while fine-tuning EfficientNet on smaller hazard-classification datasets within the change area, making the approach more efficient and reliable than direct classification methods. Experimental results show that the proposed CLIFF algorithm outperforms state-of-the-art deep learning algorithms such as LightCDNet and ChangeFormer, with an IoU of 75.74% and an F1 score of 0.8689 for change detection. Meanwhile, CLIFF has an overall accuracy rate of 86.89% in identifying geological hazards along gas pipelines, such as crude oil spills, collapses, landslides, and floods, with per-class accuracies of 87.32% and 86.17% for crude oil spills and landslides, respectively. Full article
Show Figures

Figure 1

20 pages, 4195 KB  
Article
Motion-Aware Geometric Context Adaptation for Streaming 3D Reconstruction of Intelligent Rail Vehicles in Low-Parallax Scenes
by Peng Jiang, Fuyuan Wang, Zhiwei Chen and Wenbo Pan
Vehicles 2026, 8(7), 168; https://doi.org/10.3390/vehicles8070168 - 20 Jul 2026
Viewed by 173
Abstract
Recent context-aware streaming 3D reconstruction frameworks provide a promising solution for online vehicle perception by maintaining anchor references, local pose windows, and trajectory memory. However, directly applying such frameworks to intelligent rail vehicles remains challenging because rail transit scenes are dominated by long [...] Read more.
Recent context-aware streaming 3D reconstruction frameworks provide a promising solution for online vehicle perception by maintaining anchor references, local pose windows, and trajectory memory. However, directly applying such frameworks to intelligent rail vehicles remains challenging because rail transit scenes are dominated by long straight motion, low-parallax visual observations, repetitive trackside structures, weak textures, and illumination variations. These characteristics may cause redundant context accumulation, unstable frame registration, and gradual trajectory drift. To address this problem, this paper proposes a motion-aware geometric context adaptation method for streaming 3D reconstruction of intelligent rail vehicles in low-parallax scenes. Instead of requiring task-specific large-scale retraining, the proposed method adapts the inference-stage geometric context using scale-normalized visual motion cues, including scale-normalized translational displacement, turning tendency, and inter-frame viewpoint variation. A motion-aware keyframe selection strategy suppresses redundant low-parallax frames while preserving geometrically informative observations in curved or pose-changing segments. An adaptive local pose reference window further regulates recent visual context to improve frame registration consistency. Experiments on rail transit sequences and the Oxford Spires dataset show that the proposed method achieves lower trajectory error than LingBot-Map and VIPE, while reducing redundant keyframe storage and preserving the qualitative continuity of rail-related structures. The method provides a practical motion-aware streaming 3D perception solution for rail transit inspection and digital infrastructure management. Full article
Show Figures

Figure 1

19 pages, 763 KB  
Systematic Review
Generative AI in Architectural Conceptual Design: A Structured Scoping Review of LLMs, Text-to-Image Diffusion Models, Spatial Layout Generation, BIM-AI Coupling, and Human-AI Workflows
by Yingjie Wang, Tianyu Li, Yuexiao Zhao, Bo Zhang and Xiaoran Huang
Buildings 2026, 16(14), 2882; https://doi.org/10.3390/buildings16142882 - 20 Jul 2026
Viewed by 287
Abstract
This article reports a structured scoping-style review, rather than a meta-analysis, of generative artificial intelligence (GenAI) applications in architectural conceptual design. Searches were conducted in Web of Science, Scopus, Dimensions and CNKI for English- and Chinese-language records, supplemented by targeted forward and backward [...] Read more.
This article reports a structured scoping-style review, rather than a meta-analysis, of generative artificial intelligence (GenAI) applications in architectural conceptual design. Searches were conducted in Web of Science, Scopus, Dimensions and CNKI for English- and Chinese-language records, supplemented by targeted forward and backward citation chasing. The evidence base distinguishes coded studies from contextual references and comprises 58 coded studies and 15 contextual references published between 2018 and 2026. The literature was coded by technical type, design stage, input modality, output modality, evaluation strategy, editability, and stated limitation. The synthesis identifies four interrelated domains: LLM-based semantic and knowledge support, diffusion-based conceptual visualization, spatially conditioned layout and 3D scene generation, and BIM/parametric coupling within human-AI workflows. The review indicates that current research is strongest in atmospheric visualization and prompt-mediated exploration, while evidence for architectural validity, downstream editability, regulatory checking, and professional accountability remains limited. Four cross-cutting challenges—controllability, evaluability, translatability, and responsibility—are operationalized as review-derived evaluation dimensions. GenAI is therefore better understood as a representational and workflow technology for early-stage exploration than as an autonomous architectural designer. Full article
(This article belongs to the Special Issue New Trends in Digital Buildings)
Show Figures

Figure 1

30 pages, 1127 KB  
Review
An Operation-Centered Review of Deep-Learning Computer Vision for Dairy Cow Management
by Chi Zhang, Ziwen Yu and Jin Wang
Agriculture 2026, 16(14), 1544; https://doi.org/10.3390/agriculture16141544 - 19 Jul 2026
Viewed by 391
Abstract
Dairy cow management depends on repeated observations of behavior and physical condition to support health, welfare, and operational decisions, but these observations remain labor-intensive. Deep learning (DL)-based computer vision can automate parts of this work, although deployment requirements differ among management operations. This [...] Read more.
Dairy cow management depends on repeated observations of behavior and physical condition to support health, welfare, and operational decisions, but these observations remain labor-intensive. Deep learning (DL)-based computer vision can automate parts of this work, although deployment requirements differ among management operations. This operation-centered structured review retained 97 distinct publications. Five operations were compared: lameness detection, behavior recognition, individual identification and tracking, body condition score estimation, and body weight prediction. The comparison addresses management tasks, inputs, outputs, label requirements, validation designs, and deployment constraints. It identifies shared concerns involving reference-standard reliability, data-capture protocols, multi-animal scenes, identity continuity, cross-farm generalization, and data access. A qualitative three-part research agenda is proposed: candidate foundations in reference standards and identity continuity; transfer and data infrastructure; and conditional extensions involving multimodal and longitudinal evaluation. Full article
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 167
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

29 pages, 6783 KB  
Article
A 3D Point Cloud Gesture Estimation Method Based on EdgeConv Reconstruction of Joint Features
by Jiu Yong, Xiaomei Lei, Jianwu Dang and Zhenzhen Zhang
Sensors 2026, 26(14), 4472; https://doi.org/10.3390/s26144472 - 14 Jul 2026
Viewed by 263
Abstract
Gesture interaction allows users to interact with objects through natural hand movements without direct physical contact. However, in gesture interaction, the hand moves frequently with large variations in direction, and hand joints are prone to self-occlusion. The accuracy of existing 3D point cloud [...] Read more.
Gesture interaction allows users to interact with objects through natural hand movements without direct physical contact. However, in gesture interaction, the hand moves frequently with large variations in direction, and hand joints are prone to self-occlusion. The accuracy of existing 3D point cloud gesture estimation methods is insufficient to meet the requirements of natural gesture interaction. This article proposes a 3D point cloud gesture estimation method based on EdgeConv to reconstruct joint features. The method first converts depth maps into point cloud data to reduce the impact of viewpoint variations on depth map representation of the same gesture, thereby mitigating the effect on estimation accuracy. Then, the global features are reconstructed using EdgeConv in the initial joint estimation module so that the reconstructed features contain structural information between hand joints. Finally, local joint refinement is performed by constructing local features centered around the refined joint twice using EdgeConv. These features include reference information from within group points to group center points and structural information between joint points. Using EdgeConv to enrich hand joint feature information can effectively improve the performance of hand joint estimation. Comparative experimental analysis was conducted on the 3D point cloud datasets ICVL, NYU, and MSRA, and complex scene experimental analysis was conducted on the InterHand2.6M dataset. Furthermore, the development of virtual–real interaction applications was carried out. The experimental results show that the 3D point cloud gesture estimation method proposed in this paper has high accuracy, strong generalization and robustness, providing solid support for natural gesture interaction applications. Full article
Show Figures

Figure 1

29 pages, 41837 KB  
Article
Uncertainty-Guided Multi-Center Prototype Alignment for Cross-Domain Few-Shot Hyperspectral Image Classification
by Qinzheng Wang, Menglei Li, Shiping Du and Li Wang
Electronics 2026, 15(14), 3092; https://doi.org/10.3390/electronics15143092 - 14 Jul 2026
Viewed by 232
Abstract
Accurate hyperspectral image analysis plays a critical role in environmental monitoring, precision agriculture, and urban mapping, yet acquiring large-scale annotated datasets for newly emerging scenes and sensors remains challenging. Cross-domain few-shot hyperspectral image classification addresses this bottleneck by transferring knowledge from a labeled [...] Read more.
Accurate hyperspectral image analysis plays a critical role in environmental monitoring, precision agriculture, and urban mapping, yet acquiring large-scale annotated datasets for newly emerging scenes and sensors remains challenging. Cross-domain few-shot hyperspectral image classification addresses this bottleneck by transferring knowledge from a labeled source domain to a sparsely annotated target domain. Existing prototype-based approaches commonly use a single-mean prototype per class, which is often inadequate for target classes with spectral–spatial heterogeneity, intra-class dispersion, and unstable episode-level statistics. Consequently, the resulting alignment reference may fail to accurately characterize class structure, thereby exacerbating negative transfer under domain shift. To address this issue, we propose a plug-and-play prototype-stability module that combines Uncertainty-Guided Clustered Alignment (UGCA) with Center Regularization. UGCA identifies hard classes using a class-level dispersion proxy and dynamically constructs multi-center prototypes to better capture intra-class heterogeneity beyond the single-mean assumption. Meanwhile, Center Regularization adds a lightweight compactness constraint on query embeddings to reduce prototype drift under sparse supervision. Experiments on three cross-domain tasks demonstrate the effectiveness of the proposed method. Compared with the reproduced MLPA baseline over 10 independent runs, the proposed method improves the mean OA from 69.21 ± 2.71%, 81.90 ± 3.45%, and 76.92 ± 1.10% to 71.24 ± 3.39%, 83.73 ± 3.89%, and 78.24 ± 1.55% on the Indian Pines (IP), University of Pavia (UP), and Houston (HT) target domains, respectively. Furthermore, cross-framework insertion into Gia-CFSL further verifies that the module is host-agnostic across prototype-driven CD-FSL frameworks and improves prototype-based hyperspectral image analysis without changing the inference pipeline. These results indicate that improving prototype quality is a critical and complementary dimension for robust cross-domain few-shot hyperspectral classification under sparse supervision. Full article
(This article belongs to the Special Issue Wearable Technologies and Applications)
Show Figures

Figure 1

27 pages, 7124 KB  
Article
Executable Reference Trajectory Construction and Conflict-Aware Residual Reinforcement Learning for Urban Multi-UAV Navigation
by Xiangzhi Zhou, Siqin Li, Qianjin Xia and Shanmei Li
Aerospace 2026, 13(7), 636; https://doi.org/10.3390/aerospace13070636 - 13 Jul 2026
Viewed by 298
Abstract
Urban multi-UAV navigation in dense building environments requires not only collision-free geometric paths but also executable flight processes under motion constraints and inter-UAV safety requirements. A static path that is feasible in a geometric map may still fail during closed-loop execution because of [...] Read more.
Urban multi-UAV navigation in dense building environments requires not only collision-free geometric paths but also executable flight processes under motion constraints and inter-UAV safety requirements. A static path that is feasible in a geometric map may still fail during closed-loop execution because of velocity limits, acceleration constraints, local path-association errors, and coupled multi-UAV interactions. Meanwhile, end-to-end reinforcement learning often suffers from unstable training, weak geometric interpretability, poor early-stage safety, and high sample complexity. To address these issues, this paper proposes a hierarchical planning-and-learning framework that connects static reference path generation, executable reference tracking, successful demonstration distillation, and conflict-aware residual reinforcement learning. First, three-dimensional reference paths are generated offline in an OpenStreetMap-based urban scene represented by cuboid buildings. Second, a damped reference-tracking mechanism transforms these static paths into closed-loop executable reference processes through local path association, monotonic progress updating, path recapture, look-ahead guidance, and bounded action construction. Third, successful pure-reference executions are distilled for behavior-cloning initialization. Finally, a bounded residual TD3 module is introduced as a local conflict-correction mechanism around the verified executable reference baseline. Experiments in an urban scene containing 754 buildings show that simplified tracking strategies fail to execute the static paths reliably, whereas the proposed full-damped reference-tracking controller achieves a 91.67% all-success rate and eliminates building collision episodes in the tracking-ablation test. Speed-sensitivity experiments at 10, 15, and 20 m/s show the same 91.67% all-success rate, indicating that the conclusion is not dependent on a single speed setting. In constructed conflict-stress tests, the conflict-aware residual TD3 module increases the all-success rate from 33.33% to 80.09%, reduces inter-UAV collision episodes from 66.67% to 11.57%, and improves the hard-safety satisfaction rate from 33.33% to 87.04%. These results show that the main contribution of the proposed framework lies in converting static geometric paths into executable reference trajectories and further enabling bounded residual correction under inter-UAV conflict conditions. Full article
Show Figures

Figure 1

19 pages, 1827 KB  
Article
A Spectral-Linear Trajectory Representation for Guided Diffusion-Based Manipulator Motion Planning in Constrained Environments
by Zongjian Chen, Zichao Zou, Rongqian Yang and Shizhong Jiang
Machines 2026, 14(7), 782; https://doi.org/10.3390/machines14070782 - 13 Jul 2026
Viewed by 301
Abstract
Guided diffusion planners for robot manipulators often fail in bottleneck scenes. The reason is that local collision corrections must be coordinated through a limited latent parameterization. This study examines the issue at the level of trajectory representation. Spectral-linear representation is introduced as a [...] Read more.
Guided diffusion planners for robot manipulators often fail in bottleneck scenes. The reason is that local collision corrections must be coordinated through a limited latent parameterization. This study examines the issue at the level of trajectory representation. Spectral-linear representation is introduced as a compact joint-space representation that combines an endpoint-conditioned linear reference path with a small set of globally supported eigenmodes derived from a conditioned temporal kernel. Using a shared training corpus, conditioning scheme, guidance rule, and dense-evaluation protocol, spectral-linear representation is compared with five representative alternatives across six representation families and dimensionality sweeps from 14 to 448 dimensions. With 14 latent dimensions, spectral-linear representation attains the highest structured-scene success rates among the tested families, reaching 74.2% on narrow-passage scenes and 84.5% on rack-like scenes. Expert-only controls, per-family guidance tuning, and a Transformer-backbone rerun are consistent with the same broad ranking pattern. These results indicate that, under limited guidance budgets and within the tested protocol, the representation structure is more predictive of bottleneck success than latent width alone. Full article
(This article belongs to the Section Robotics, Mechatronics and Intelligent Machines)
Show Figures

Figure 1

25 pages, 1458 KB  
Article
A Digital Twin Framework for Multimodal Operator-Centered Human–Cobot Collaboration in Assembly Tasks
by David Alfaro-Viquez, Mauricio Zamora-Hernandez, Michael Fernandez-Vega, David Ortiz-Perez, Jose Garcia-Rodriguez and Jorge Azorin-Lopez
Machines 2026, 14(7), 780; https://doi.org/10.3390/machines14070780 - 12 Jul 2026
Viewed by 246
Abstract
Current digital twin frameworks focused on human–robot collaboration rarely take into account the sensory degradation of real industrial environments, nor do they integrate the operator as an active agent within the system. This research presents a multimodal digital twin framework for a dual-arm [...] Read more.
Current digital twin frameworks focused on human–robot collaboration rarely take into account the sensory degradation of real industrial environments, nor do they integrate the operator as an active agent within the system. This research presents a multimodal digital twin framework for a dual-arm collaborative robot at an assembly station; the system was developed using ROS 2 Jazzy and CoppeliaSim as the simulator. The architecture integrates three main components: the first is a perception layer that captures voice commands using Whisper ASR and the state of the workspace using a hybrid YOLO + ViT visual pipeline, both with per-channel metadata; the second consists of a Confidence-Weighted Late Fusion engine that dynamically adjusts the weight of each modality based on real-time signal quality, so that each fusion decision can be reconstructed from the signals that generated it; and the third component is a Reference Resolver that grounds linguistic intent within the visual context of the scene and in the fusion weights, using a local instance of Llama 3.1 8B that does not transmit audio, transcripts, or images outside the system. The framework was evaluated using 210 iterations distributed across seven degradation conditions of increasing severity, comparing adaptive fusion against a baseline of fixed weights (0.5/0.5). Under clean conditions and under visual degradation of any severity, both configurations achieved 100% accuracy. Under severe auditory degradation (SNR 0 dB), adaptive fusion activated the safety gate and refrained from executing most commands (13.3% accuracy), while the fixed-weight baseline executed more commands (60% accuracy) but made three incorrect object selections; under severe dual degradation, the pattern repeated (13.3% vs. 40%, with five incorrect selections in the baseline). The adaptive system made no grounding errors in the 210 executions, compared to eight in the baseline, substituting incorrect execution with conservative abstention when no modality provided a reliable signal. The implementation, featuring a versioned degradation protocol and a fixed seed, provides a reproducible benchmark for evaluating multimodal fusion strategies in human–cobot interaction. Full article
Show Figures

Figure 1

Back to TopTop