Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (830)

Search Parameters:
Keywords = dynamical image reconstruction

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 24261 KB  
Article
Lightweight 2.5D SLAM with Dynamic Map Refinement and Height-Aware Encoding for Resource-Constrained Indoor Robots
by Guitao Yu, Yuping Zhang, Zhiao Qi, Kui Yang, Yang He and Dongtai Liang
Sensors 2026, 26(15), 4765; https://doi.org/10.3390/s26154765 - 27 Jul 2026
Abstract
Indoor mobile robots equipped with low-cost and sparse sensors often suffer from limited vertical perception and dynamic residual artifacts in the final map. This paper presents a lightweight 2.5D simultaneous localization and mapping (SLAM) framework using a single-line laser distance sensor (LDS), time-of-flight [...] Read more.
Indoor mobile robots equipped with low-cost and sparse sensors often suffer from limited vertical perception and dynamic residual artifacts in the final map. This paper presents a lightweight 2.5D simultaneous localization and mapping (SLAM) framework using a single-line laser distance sensor (LDS), time-of-flight (ToF) sensing, wheel odometry, and an inertial measurement unit (IMU). In this work, 2.5D refers to a 2D grid map with discretized vertical occupancy bins for each grid cell, rather than a full continuous 3D reconstruction. The system integrates multi-sensor synchronization, motion correction, error-state Kalman filter (ESKF)-based state estimation, normal distributions transform (NDT) registration, and pose graph optimization to reconstruct a pose-consistent global map. Based on this map, an offline dynamic refinement module estimates temporal voxel support across keyframes, extracts low-support candidate regions, and applies geometric clustering and isolated-point filtering to suppress transient residual artifacts while preserving stable structures. A 24-bit RGB occupancy encoding is further proposed to store the discretized vertical occupancy state in a compact three-channel image format. The proposed framework emphasizes system-level deployment value by combining sparse multi-sensor mapping, conservative offline refinement, and compact height-aware map export on a low-cost indoor robot platform. Experiments on public datasets, embedded hardware, and self-collected indoor sequences evaluate odometry reference performance, resource usage, platform-specific 2.5D mapping, dynamic refinement, and height-aware encoding. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

30 pages, 12181 KB  
Article
Pore-Scale Characterization of Remaining Oil Evolution During Waterflooding in Offshore Sandstone Reservoirs Using Micro-CT and U-Net Semantic Segmentation
by Wensheng Zhou, Chen Liu, Deqiang Wang, Chi Zhang, Wenhui Gao, Yushan Ma, Yuxin Zhang and Yaopan Yu
Energies 2026, 19(15), 3518; https://doi.org/10.3390/en19153518 - 26 Jul 2026
Abstract
After offshore sandstone reservoirs enter the high-water-cut development stage, remaining oil gradually changes from a continuously movable state to a dispersed, isolated, and locally trapped state. Therefore, pore-scale dynamic characterization of remaining oil is important for clarifying its occurrence state and identifying potential [...] Read more.
After offshore sandstone reservoirs enter the high-water-cut development stage, remaining oil gradually changes from a continuously movable state to a dispersed, isolated, and locally trapped state. Therefore, pore-scale dynamic characterization of remaining oil is important for clarifying its occurrence state and identifying potential targets for further recovery. In this study, natural cores from a typical offshore continental sandstone reservoir were used to conduct micro-computed tomography (micro-CT) scanning experiments at different waterflooding stages. Fine identification of the rock skeleton, water phase, and remaining oil, as well as segmentation of remaining-oil units, was achieved using U-Net semantic segmentation and three-dimensional digital image reconstruction. Multiple parameters were extracted, including volume, surface area, equivalent diameter, major-axis length, intermediate-axis width, minor-axis thickness, shape factor, and Euler number. Combined with manually labeled results of typical remaining-oil unit, a decision-tree method was used to establish classification criteria for remaining-oil occurrence states. The results showed that, with increasing pore volumes injected (PV), remaining oil gradually evolved from an initially large-scale and continuously connected state to a small-scale dispersed trapping. The surface area and volume of remaining oil continuously decreased. Meanwhile, the shape factor decreased, and the Euler number increased, indicating that the continuous oil phase was continuously cut and stripped into remaining-oil units with lower topological connectivity during waterflooding. At the early stage of waterflooding, clustered flow dominated in cores with different permeabilities, accounting for 96–99% of the remaining-oil volume percentage. Under high pore volumes injected (PV), the remaining oil volume percentage decreased to 30–65%, whereas the remaining oil volume percentage of columnar flow and droplet flow increased to 16–37% and 5–13%, respectively. These results indicated that waterflooding promoted the transformation of remaining oil from large-scale connected occurrence to locally trapped states under pore–throat constraint. Further comparison of cores with similar permeability but different pore structures showed that differences in pore–throat geometry, local connectivity, and pore–throat matching still significantly affected the fragmentation degree and spatial trapping mode of remaining oil. This study established a pore-scale method for dynamic identification, quantitative characterization, and type classification of microscopic remaining oil during long-term waterflooding in offshore sandstone reservoirs. The results provide a pore-scale basis for fine remaining-oil potential tapping in the high-water-cut stage. Full article
(This article belongs to the Section H1: Petroleum Engineering)
Show Figures

Figure 1

34 pages, 13120 KB  
Article
Comparative Analysis of Strain-to-Velocity Conversion Methods for Active-Source DAS Data and Collocated Nodal Stations
by Prajwal Panthi and Brady R. Cox
Sensors 2026, 26(15), 4673; https://doi.org/10.3390/s26154673 - 23 Jul 2026
Viewed by 161
Abstract
Distributed Acoustic Sensing (DAS) provides dense spatial measurements of the dynamic strain along fiber optic cables, offering high-resolution wave sensing for ground motion monitoring and subsurface imaging applications. However, DAS records the axial strain or strain rate, whereas traditional seismic and engineering ground [...] Read more.
Distributed Acoustic Sensing (DAS) provides dense spatial measurements of the dynamic strain along fiber optic cables, offering high-resolution wave sensing for ground motion monitoring and subsurface imaging applications. However, DAS records the axial strain or strain rate, whereas traditional seismic and engineering ground motion equipment and derived metrics are based on particle displacement, velocity, or acceleration, necessitating reliable strain-to-velocity conversion methods. This study evaluates three widely used conversion approaches: the fk-rescaling, curvelet-based conversion, and slant-stack methods. These approaches are applied to a unique high-energy, near-field, active-source dataset collected at the Birds Landing Site in Sherman Island, California. The dataset includes wavefields generated by a large transmission tower collapse and sledgehammer impacts used for subsurface imaging. The wavefields were recorded simultaneously by a 1.4 km DAS array and 71 collocated nodal stations (NSs). Using 63 DAS–NS pairs, we quantify the strain-to-velocity conversion method performance using amplitude and phase transfer functions (TFs) between DAS-derived and NS particle velocity records, with the root-mean-square error (RMSE) evaluated across three frequency bands: 0.5–100 Hz, 1–10 Hz, and 10–100 Hz. The results show that fk-rescaling provides the most stable amplitude response across both source types, while both the fk-rescaling and slant-stack methods generally yield the best phase agreement. Curvelet-based conversion shows a greater variability and larger RMSE values. All methods yield a poorer amplitude reconstruction at higher frequencies, while the phase content is generally preserved more reliably than amplitudes. Differences between the tower collapse and sledgehammer sources demonstrate the influence of the source characteristics and spatial processing window length on the conversion performance. The findings provide practical guidance for selecting suitable strain-to-velocity conversion methods for active-source DAS applications, particularly where collocated reference sensors are unavailable. Full article
(This article belongs to the Special Issue Distributed Acoustic Sensing and Applications)
Show Figures

Figure 1

20 pages, 5197 KB  
Article
HDR Scene Reconstruction from Resolution-Mismatched Event Streams and Single-Exposure Images
by Zehao Chen, Binbin Zhou and Zengwei Zheng
Sensors 2026, 26(14), 4621; https://doi.org/10.3390/s26144621 - 21 Jul 2026
Viewed by 254
Abstract
High dynamic range (HDR) radiance-field reconstruction from single-exposure low dynamic range (LDR) images is limited by the information loss in saturated regions, while event cameras provide complementary measurements whose dynamic range far exceeds that of conventional sensors. In practical hybrid sensor systems, however, [...] Read more.
High dynamic range (HDR) radiance-field reconstruction from single-exposure low dynamic range (LDR) images is limited by the information loss in saturated regions, while event cameras provide complementary measurements whose dynamic range far exceeds that of conventional sensors. In practical hybrid sensor systems, however, the RGB camera usually has a higher spatial resolution than the event sensor, which makes existing event-aided HDR reconstruction methods difficult to apply directly. A straightforward solution is to super-resolve the event stream before reconstruction, but this 2D preprocessing introduces a global color cast and view-inconsistent high-frequency artifacts once the super-resolved events supervise a 3D radiance field. We propose a framework for HDR radiance-field reconstruction from resolution-mismatched event–image inputs. The framework incorporates a pretrained 2D event super-resolution prior and corrects its transfer to 3D reconstruction through a color correction module, which anchors the rendered radiance to the chrominance of the LDR images, and a dual-resolution event-stream constraint, which supervises the synthesized events at both the super-resolved and the native resolution. Experiments on the EvHDR-NeRF benchmark show that the proposed method achieves higher six-scene mean HDR fidelity than the strongest baseline while reducing the global color cast and alleviating multi-view artifacts. Full article
(This article belongs to the Special Issue Event-Based Vision and Multimodal Sensor Fusion)
Show Figures

Figure 1

26 pages, 10156 KB  
Article
Antecedent Topographic and Shoreline-Infrastructure Controls on Urban Beach Geomorphic Response to Lake Michigan Water-Level Rise
by Christopher R. Mattheus
Limnol. Rev. 2026, 26(3), 41; https://doi.org/10.3390/limnolrev26030041 - 20 Jul 2026
Viewed by 132
Abstract
This paper addresses the geomorphic response of an engineered Chicago beach to a >1.5 m rise in Lake Michigan’s base water level, from 2013 to 2020. Topographic monitoring data acquired since 2012 and subsurface geophysical data collected in 2022 are integrated to explore [...] Read more.
This paper addresses the geomorphic response of an engineered Chicago beach to a >1.5 m rise in Lake Michigan’s base water level, from 2013 to 2020. Topographic monitoring data acquired since 2012 and subsurface geophysical data collected in 2022 are integrated to explore the roles of lakefront infrastructure and beach topographic development on sedimentary dynamics, beach morphologic development, and stratigraphic architecture. While conceptual models of coastal geomorphology infer upward and landward beach-profile translation with lake-level rise, a high degree of along-shore variance occurs within pocket beaches. This stems from infrastructure-related modifications of storm hydrodynamics and scour patterns, close to shore, and backshore terrain physiography. The studied beach evolved contrary to how regional littoral drift patterns would have suggested, with erosion most severe along the embayment’s downdrift end, a product of infrastructure-induced scour and the reduced capacity for sediment retention through overwash accretion. Documented geomorphic patterns with lake-level rise are manifested in subsurface imaging data from the lake-level highstand, accordingly, providing a guide to more regional paleo-reconstruction. Great Lakes urban pocket beaches buffer shoreline infrastructure, offer recreational terrains, and support dune ecosystems of value to migrating shorebirds. Understanding their geomorphology benefits coastal managers looking to mitigate impacts of future climate and lake-level change. Full article
Show Figures

Figure 1

21 pages, 10432 KB  
Article
Projection-Free CLIP-Scale EEG Latents via a U-Net-Style Autoencoder
by Jeyoung Lee, Jaekwan Ahn, Jaeseung Sim and Hochul Kang
Sensors 2026, 26(14), 4583; https://doi.org/10.3390/s26144583 - 20 Jul 2026
Viewed by 240
Abstract
Electroencephalography is emerging as a promising conditioning modality for generative visual models. However, existing representation learning approaches often rely on high-capacity masked autoencoders and complex projection networks. When constrained to compact embedding dimensions to match vision-language models, these heavy transformer-based bottlenecks frequently suffer [...] Read more.
Electroencephalography is emerging as a promising conditioning modality for generative visual models. However, existing representation learning approaches often rely on high-capacity masked autoencoders and complex projection networks. When constrained to compact embedding dimensions to match vision-language models, these heavy transformer-based bottlenecks frequently suffer from representation collapse and lose critical signal dynamics. To address this, we propose a lightweight and projection-free autoencoder that directly outputs compact, Contrastive Language–Image Pre-training (CLIP)-scale latent vectors trained toward the CLIP embedding space. Our model adopts a U-Net-style architecture combining one-dimensional convolutional residual blocks for temporal dynamics and inter-channel attention modules for spatial dependencies, alongside skip connections to ensure stable reconstruction. Extensive experiments on visual perception datasets demonstrate that our approach successfully tracks complex signal amplitudes without collapsing. Under strict dimensional constraints, the proposed model achieves superior signal reconstruction fidelity across time and frequency domains using significantly fewer parameters than traditional masked autoencoder baselines. Furthermore, latent space visualizations and zero-shot retrieval tasks reveal that while the baseline collapses toward unstructured, near-chance representations, our architecture preserves emerging, partial semantic organization and retrieves several times above chance. This indicates that the proposed design preserves signal structure while exhibiting preliminary, above-chance semantic alignment, enabling integration into brain-driven generative pipelines. Full article
(This article belongs to the Special Issue Biosignal Sensing Analysis (EEG, EMG, ECG, PPG) (3rd Edition))
Show Figures

Figure 1

27 pages, 10988 KB  
Article
Text Image Super-Resolution via Fusion of OCR Priors and Cross-Scale Attention
by Xinyu Qiu, Jingchao Liu and Chen Fang
Symmetry 2026, 18(7), 1221; https://doi.org/10.3390/sym18071221 - 20 Jul 2026
Viewed by 262
Abstract
Text image super-resolution aims to improve the readability of low-quality text images while preserving character structures, stroke details, and semantic consistency. Compared with natural image super-resolution, this task is more sensitive to structural distortion because small changes in stroke topology may lead to [...] Read more.
Text image super-resolution aims to improve the readability of low-quality text images while preserving character structures, stroke details, and semantic consistency. Compared with natural image super-resolution, this task is more sensitive to structural distortion because small changes in stroke topology may lead to incorrect text recognition. To address this problem, this paper proposes an OCR prior-guided cross-scale framework for text image super-resolution. Specifically, character-level semantic priors extracted from a pretrained OCR model are introduced to provide structural guidance for degraded text reconstruction. A gated feature modulation mechanism is designed to adaptively regulate the contribution of OCR priors, reducing the influence of unreliable semantic predictions. A cross-scale dynamic attention module is also developed to aggregate multi-granularity visual features, enabling the model to jointly recover fine stroke boundaries and global character structures. In addition, a sequence-aware calibration module is introduced to improve structural consistency along the logical reading order of text. Experiments on mixed text image benchmarks and the TextZoom dataset show that the proposed method achieves competitive or better performance among the compared methods in terms of PSNR, SSIM, and recognition-oriented metrics. Additional ablation, OCR prior robustness, and computational complexity analyses further indicate that the proposed framework improves text readability while maintaining a reasonable accuracy–complexity trade-off. The results also suggest that OCR priors are useful for text image reconstruction, but should be used as soft constraints when external recognition predictions are uncertain. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 171
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

38 pages, 4675 KB  
Article
Enhancing PatchCore with Dynamic Scaling and Vision–Language Models for Explainable Industrial Defect Inspection
by Oğuz Ergin, Emre Güçlü, İlhan Aydın and Erhan Akın
Appl. Sci. 2026, 16(14), 7096; https://doi.org/10.3390/app16147096 - 15 Jul 2026
Viewed by 217
Abstract
Although unsupervised anomaly detection has shown promising performance in industrial visual inspection, many architectures still struggle with variable input resolutions, and anomaly scores are often difficult for end users to interpret. This study proposes a multi-stage hybrid workflow for pixel-level defect localization and [...] Read more.
Although unsupervised anomaly detection has shown promising performance in industrial visual inspection, many architectures still struggle with variable input resolutions, and anomaly scores are often difficult for end users to interpret. This study proposes a multi-stage hybrid workflow for pixel-level defect localization and structured reporting in bolt head images. The proposed SA-PatchCore framework customizes PatchCore by extracting multi-scale representations from a frozen deep feature extractor and supporting resolution-adaptive anomaly-map reconstruction through dynamic feature-map sizing. After anomaly detection, a Qwen3-VL-32B-based reporting module, adapted with GRPO, uses both the original image and the anomaly overlay as visual evidence. It generates structured JSON outputs containing defect presence, a 3 × 3 location label, and a concise textual description. On the industrial bolt dataset, SA-PatchCore achieved 98.69% pixel-level AUROC, 29.73% Pixel-AP, and 39.24% oracle Pixel-F1max. Compared with PatchCore, PaDiM, DRÆM, and CS-Flow, the method delivered strong results, especially in Pixel-AUROC. In the reporting stage, defect presence/absence accuracy improved from 76.63% to 96.41%, while defective-sample recall increased from 74.23% to 96.14% over the baseline Qwen3-VL-32B. Exact location match rose from 28.15% to 53.78%, and mean partial location score improved from 32.29% to 65.34%. Overall, the framework combines accurate anomaly localization with structured reporting, improving interpretability and usability. Full article
(This article belongs to the Topic Smart Production in Terms of Industry 4.0 and 5.0)
Show Figures

Figure 1

23 pages, 1305 KB  
Article
Semantic Communication for Intelligent Transmission and Recognition of High-Resolution Satellite Images in Satellite-to-Ground Systems
by Jiaxin Liu, Qiwang Chen and Yijun Chen
Entropy 2026, 28(7), 803; https://doi.org/10.3390/e28070803 - 14 Jul 2026
Viewed by 225
Abstract
Very-high-resolution (VHR) multispectral satellite imagery contains rich semantic information, yet its real-time transmission is constrained by limited satellite-to-ground bandwidth and dynamic channel impairments. Conventional communication schemes prioritize pixel-level reconstruction, resulting in large transmission overhead and poor robustness under unfavorable channel conditions. To address [...] Read more.
Very-high-resolution (VHR) multispectral satellite imagery contains rich semantic information, yet its real-time transmission is constrained by limited satellite-to-ground bandwidth and dynamic channel impairments. Conventional communication schemes prioritize pixel-level reconstruction, resulting in large transmission overhead and poor robustness under unfavorable channel conditions. To address these challenges, an end-to-end task-oriented semantic communication framework for remote sensing downstream recognition tasks, termed Semantic Transmission Architecture for Remote Sensing (STARS), is proposed. To improve transmission efficiency for very-high-resolution remote sensing images with highly redundant background regions, a Semantic Feature Reweighting Module (SFRM) is introduced to dynamically evaluate token-level semantic importance and adaptively allocate transmission resources to task-critical features. Furthermore, vector quantization and a practical digital transmission chain are jointly integrated to achieve efficient semantic compression, while dynamic channel variations are incorporated during training to improve robustness under fading channel conditions. Experimental results on the DOTA dataset demonstrate that STARS consistently outperforms conventional schemes and existing semantic baselines under Rician fading channels, validating the effectiveness of semantic-aware feature allocation for bandwidth-efficient VHR imagery transmission. Full article
Show Figures

Figure 1

15 pages, 8736 KB  
Article
Topographic–Climatic Interactions Drive Vegetation NPP Dynamics in the West Qinling Mountains (2003–2025)
by Ling Nan, Yongliu Li, Xiangshuai Zhang and Qiaorui Ba
Ecologies 2026, 7(3), 67; https://doi.org/10.3390/ecologies7030067 - 14 Jul 2026
Viewed by 269
Abstract
Mountain transition zones are highly sensitive to environmental change, yet the nonlinear coupling between topography and hydroclimate in controlling vegetation Net Primary Productivity (NPP) remains insufficiently constrained. Here, we reconstructed a 2003–2025 annual NPP time series for the West Qinling Mountains using a [...] Read more.
Mountain transition zones are highly sensitive to environmental change, yet the nonlinear coupling between topography and hydroclimate in controlling vegetation Net Primary Productivity (NPP) remains insufficiently constrained. Here, we reconstructed a 2003–2025 annual NPP time series for the West Qinling Mountains using a Carnegie–Ames–Stanford Approach (CASA)-based workflow that integrated Moderate Resolution Imaging Spectroradiometer (MODIS) vegetation products, fifth-generation European Centre for Medium-Range Weather Forecasts reanalysis for land (ERA5-Land) meteorological data, Shuttle Radar Topography Mission (SRTM) topography, and an Aridity Index (AI) dataset. Product-based validation against the annual MOD17A3HGF dataset indicated strong agreement, with a 23-year mean spatial Spearman correlation of 0.823 and a mean annual Pearson correlation of 0.773. The reconstructed dataset showed that 98.50% of the study area experienced increasing NPP, including 50.26% with significant increases, and the domain-wide mean Sen slope reached approximately 3.08 g C m−2 yr−1. Factor detection further showed that radiation (q = 0.253), elevation (q = 0.252), and temperature (q = 0.249) were the dominant single controls, whereas Aridity–Temperature (q = 0.367) and Elevation–Aridity (q = 0.367) represented the strongest interactions. The concentration of the strongest gains in gentle-slope and moderate-aridity settings suggests that vegetation recovery is maximized where topographic buffering and water-energy balance are jointly optimized. These results strengthen the interpretation of NPP dynamics in mountainous climate-transition environments and provide a basis for spatially targeted ecological restoration, regional carbon-budget assessment, and climate adaptation planning. Full article
Show Figures

Figure 1

25 pages, 1458 KB  
Article
A Digital Twin Framework for Multimodal Operator-Centered Human–Cobot Collaboration in Assembly Tasks
by David Alfaro-Viquez, Mauricio Zamora-Hernandez, Michael Fernandez-Vega, David Ortiz-Perez, Jose Garcia-Rodriguez and Jorge Azorin-Lopez
Machines 2026, 14(7), 780; https://doi.org/10.3390/machines14070780 - 12 Jul 2026
Viewed by 248
Abstract
Current digital twin frameworks focused on human–robot collaboration rarely take into account the sensory degradation of real industrial environments, nor do they integrate the operator as an active agent within the system. This research presents a multimodal digital twin framework for a dual-arm [...] Read more.
Current digital twin frameworks focused on human–robot collaboration rarely take into account the sensory degradation of real industrial environments, nor do they integrate the operator as an active agent within the system. This research presents a multimodal digital twin framework for a dual-arm collaborative robot at an assembly station; the system was developed using ROS 2 Jazzy and CoppeliaSim as the simulator. The architecture integrates three main components: the first is a perception layer that captures voice commands using Whisper ASR and the state of the workspace using a hybrid YOLO + ViT visual pipeline, both with per-channel metadata; the second consists of a Confidence-Weighted Late Fusion engine that dynamically adjusts the weight of each modality based on real-time signal quality, so that each fusion decision can be reconstructed from the signals that generated it; and the third component is a Reference Resolver that grounds linguistic intent within the visual context of the scene and in the fusion weights, using a local instance of Llama 3.1 8B that does not transmit audio, transcripts, or images outside the system. The framework was evaluated using 210 iterations distributed across seven degradation conditions of increasing severity, comparing adaptive fusion against a baseline of fixed weights (0.5/0.5). Under clean conditions and under visual degradation of any severity, both configurations achieved 100% accuracy. Under severe auditory degradation (SNR 0 dB), adaptive fusion activated the safety gate and refrained from executing most commands (13.3% accuracy), while the fixed-weight baseline executed more commands (60% accuracy) but made three incorrect object selections; under severe dual degradation, the pattern repeated (13.3% vs. 40%, with five incorrect selections in the baseline). The adaptive system made no grounding errors in the 210 executions, compared to eight in the baseline, substituting incorrect execution with conservative abstention when no modality provided a reliable signal. The implementation, featuring a versioned degradation protocol and a fixed seed, provides a reproducible benchmark for evaluating multimodal fusion strategies in human–cobot interaction. Full article
Show Figures

Figure 1

25 pages, 2814 KB  
Article
BRNet: A Dual-Backbone X-Ray Coronary Angiography Segmentation Network Based on Multi-Scale Fusion and Dynamic Detail Reconstruction
by Zhan Zhang, Hong Shao and Wencheng Cui
Appl. Sci. 2026, 16(14), 6960; https://doi.org/10.3390/app16146960 - 10 Jul 2026
Viewed by 320
Abstract
X-ray coronary angiography remains the gold standard for diagnosing coronary heart disease. However, accurate segmentation is challenged by the subtlety of fine vascular features, topological discontinuities, and blurred boundaries in these images. Existing methods often struggle to capture both the long-range global topology [...] Read more.
X-ray coronary angiography remains the gold standard for diagnosing coronary heart disease. However, accurate segmentation is challenged by the subtlety of fine vascular features, topological discontinuities, and blurred boundaries in these images. Existing methods often struggle to capture both the long-range global topology and the fine local details required for robust segmentation. To address these issues, we propose BRNet, a framework that integrates a dual-backbone collaborative mechanism with multi-scale fusion and dynamic detail reconstruction. Our approach first employs a Vascular Local Detail Attention Module that combines ResNet18’s local perception with BiFormer’s global modeling, using SimAM parameter-free attention to suppress background noise. We then design a Vascular Structure-Coherent Progressive Fusion Module, which uses a top-down pyramid semantic flow to ensure topological coherence across different vascular scales. Finally, a Vascular Enhancement Dynamic Upsampling Module replaces traditional interpolation with content-aware CARAFE operators to achieve high-resolution reconstruction of blurred edges. Experiments on the public ARCADE dataset and the constructed heterogeneous benchmark ZZ-CAHDS show that BRNet achieves superior performance, attaining IoU scores of 0.6305 and 0.6774, and clDice coefficients of 0.7655 and 0.8355, respectively, and achieving a good balance between segmentation accuracy and computational efficiency. These results highlight BRNet’s effectiveness on retrospective segmentation benchmarks, demonstrating its capability for computer-assisted coronary artery segmentation. Full article
(This article belongs to the Section Biomedical Engineering)
Show Figures

Figure 1

22 pages, 5755 KB  
Article
A Dynamic Displacement Measurement Method for Overhead Transmission Line Galloping Based on Deep Vision and Binocular Collaboration
by Jian Wang, Danyu Li, Bin Liu, Wenbo Gao and Xinyi Gong
Electronics 2026, 15(14), 3040; https://doi.org/10.3390/electronics15143040 - 10 Jul 2026
Viewed by 177
Abstract
Galloping of overhead transmission lines threatens grid safety and requires non-contact measurement methods that can quantify three-dimensional (3D) motion from field video. This paper proposes a deep-vision and binocular-collaboration framework for dynamic conductor displacement measurement. The framework combines three components that are matched [...] Read more.
Galloping of overhead transmission lines threatens grid safety and requires non-contact measurement methods that can quantify three-dimensional (3D) motion from field video. This paper proposes a deep-vision and binocular-collaboration framework for dynamic conductor displacement measurement. The framework combines three components that are matched to the physical structure of transmission lines: adaptive image enhancement using Retinex illumination decomposition and Wiener blind deconvolution; a structure-prior dual-branch extraction module that uses an improved YOLOv11 keypoint branch for spacer-equipped sections and an improved U-Net branch with Dynamic Snake Convolution (DSC) and Strip Pooling for bare conductors; and stereo reconstruction with Kalman-filter-based temporal association for continuous trajectory estimation. Compared with the original submission, the revised manuscript further clarifies the real-video data acquisition, annotation procedure, camera synchronization, calibration workflow, training/testing independence, and runtime measurement protocol. Additional validation on a public real power-line image dataset is also reported. The proposed method achieves a Z-axis Root Mean Square Error (RMSE) of 24.5 mm for spacer sections in the controlled binocular field test, a dominant-frequency relative error below 3.5%, and 32 FPS on edge hardware when preprocessing, visual extraction, stereo projection, and temporal filtering are included. On the supplementary public power-line dataset, the segmentation branch obtains a Dice coefficient of 0.9039 and an IoU of 0.8395. These results indicate that the proposed framework reduces the depth-scale limitation of monocular vision and provides a practical quantitative tool for field galloping monitoring. Full article
Show Figures

Figure 1

27 pages, 15267 KB  
Article
PanDiM: A Diffusion Mamba Network for High-Fidelity Pansharpening
by Haobo Xu, Yao Zhang, Lingfeng Lin, Jiajin Wu, Boxiang Xie, Wei Zhang, Honggang Li and Jing Qu
Remote Sens. 2026, 18(14), 2299; https://doi.org/10.3390/rs18142299 - 9 Jul 2026
Viewed by 277
Abstract
Pansharpening plays an important role in remote sensing image processing. Its purpose is to fuse a high-spatial-resolution panchromatic (PAN) image and a low-spatial-resolution multispectral (LRMS) image, thereby reconstructing a high-resolution multispectral (HRMS) image with both high spatial clarity and high spectral fidelity. In [...] Read more.
Pansharpening plays an important role in remote sensing image processing. Its purpose is to fuse a high-spatial-resolution panchromatic (PAN) image and a low-spatial-resolution multispectral (LRMS) image, thereby reconstructing a high-resolution multispectral (HRMS) image with both high spatial clarity and high spectral fidelity. In recent years, diffusion models have shown great potential in image generation. However, existing diffusion-based pansharpening methods usually adopt a fixed denoising strategy, making it difficult to adapt to the stage-wise changes in the denoising process and complex degradation distributions. Based on this, we propose PanDiM, an efficient generative framework for pansharpening. Specifically, we reformulate pansharpening as a high-frequency residual restoration process constrained by multimodal conditions. To improve the response accuracy of the model in complex regions, we design a Degradation-Posterior Guidance Module (DPGM), which extracts dual-scale physical detail priors from the PAN image, explicitly infers the degradation posterior, and converts it into dynamic control variables to adaptively regulate the state evolution of Mamba. In addition, we propose a time-aware mechanism, which allows temporal information to directly intervene in posterior estimation and state-space modeling, so as to accurately match the modeling requirements of different denoising stages. Considering the characteristics of residual reconstruction, we further propose a frequency-decoupled loss (FDL), which separates low- and high-frequency components in the frequency domain and applies targeted constraints. This significantly enhances the model’s ability to represent textures and achieves more robust spectral fidelity. Extensive experiments on three benchmark datasets, including WorldView-3, GaoFen-2, and QuickBird, show that PanDiM significantly outperforms existing mainstream methods in both reduced-resolution and full-resolution evaluations, providing a new solution for high-fidelity pansharpening in complex scenarios. Full article
Show Figures

Figure 1

Back to TopTop