Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (115)

Search Parameters:
Keywords = video–sensor fusion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 2034 KB  
Article
Camera–GPS Sensor Fusion for Kinematic Characterization, Microsimulation Validation, and Macroscopic Capacity Modeling of Traffic-Calming Corridors
by Deo Chimba, Wittness Mariki, Sunam Shrestha and Afia Yeboah
Sensors 2026, 26(17), 5340; https://doi.org/10.3390/s26175340 - 24 Aug 2026
Viewed by 310
Abstract
This study presents a sensor-fused field investigation and simulation-based analysis of four horizontal and vertical traffic-calming devices—two raised speed tables, a speed hump, and a raised crosswalk—installed along a 5250-ft two-lane residential collector in Nashville, TN, USA. A dual-sensor architecture combining a Miovision [...] Read more.
This study presents a sensor-fused field investigation and simulation-based analysis of four horizontal and vertical traffic-calming devices—two raised speed tables, a speed hump, and a raised crosswalk—installed along a 5250-ft two-lane residential collector in Nashville, TN, USA. A dual-sensor architecture combining a Miovision Scout video-based vehicle counter and WAAS/EGNOS-augmented GPS probe-vehicle logging (5 m 3-D RMS horizontal accuracy, 1 Hz sampling) was used to reconstruct 30 quality-controlled free-flow vehicle trajectories and 12-h per-lane volume counts. A spatial kinematic transform (a = v·dv/dx) was applied to extract device-specific approach-deceleration and post-device recovery-acceleration rates, and a three-parameter log-logistic cumulative-distribution function was fitted to the field-observed desired-speed percentiles (root-mean-square error below 0.043 for both speed-table devices). The camera- and GPS-derived observations were used to calibrate and statistically validate a PTV VISSIM microsimulation replica of the corridor, achieving a mean-speed calibration error of 0.71% or better at every device, a GEH statistic below 1.5 at all four analysis turning movements, and independent travel-time validation errors of 5.7–12.1%, within the accepted 15% threshold. The validated model was then used to reconstruct device- and spacing-specific May–Keller macroscopic speed–density–flow relationships, calibrated against simulated capacities of 650–775 vehicles per hour per lane at 350-, 700-, and 1050-ft device spacing. Results show capacity reductions of 20–33% relative to free-flow conditions and yield kinematically derived maximum recommended spacings of 265–630 ft to maintain crossing speeds at or below 15 mph, depending on device geometry. The findings demonstrate a reproducible, low-cost sensor-fusion workflow for quantifying the safety–capacity trade-off of traffic-calming corridors and for informing the design of sensor-in-the-loop adaptive-calming infrastructure. Full article
Show Figures

Figure 1

26 pages, 3704 KB  
Article
Privacy-Preserving Ambient Sensing for Activities of Daily Living: Multimodal Radar–Thermal Human Activity Recognition and Smart Plug Appliance Recognition
by Bilal Mohammed, Jordan J. Bird, Isibor Kennedy Ihianle, Martin Harris, Geoff Archenhold and Yangang Xing
Sensors 2026, 26(16), 5066; https://doi.org/10.3390/s26165066 - 10 Aug 2026
Viewed by 397
Abstract
Continuous monitoring of Activities of daily living (ADLs) requires sensing systems that are privacy-preserving, low-power, and robust to environmental variation. Ambient sensing technologies provide an alternative to RGB video and wearable devices, but individual sensing modalities exhibit characteristic limitations. Sparse mmWave radar provides [...] Read more.
Continuous monitoring of Activities of daily living (ADLs) requires sensing systems that are privacy-preserving, low-power, and robust to environmental variation. Ambient sensing technologies provide an alternative to RGB video and wearable devices, but individual sensing modalities exhibit characteristic limitations. Sparse mmWave radar provides strong motion sensitivity but limited posture detail, low-resolution thermal sensing preserves posture-related spatial information, and smart plug telemetry captures only appliance-mediated behavioural interaction. To address these limitations, this paper proposes a layered multimodal ambient-sensing framework comprising a sparse-track 24-GHz FMCW radar, a 32×24 low-resolution thermal sensor, and a Moko smart plug. It experimentally evaluates a radar–thermal HAR branch together with a separate smart plug appliance-recognition branch. The framework proposes three streams to enable continuous non-wearable monitoring while maintaining redundancy and reduced privacy exposure for intelligent-building and ambient assisted living environments. Radar and thermal streams are jointly evaluated on binary motion and four-class posture and activity recognition tasks collected across multiple environmental configurations using recording-grouped cross-validation, while the appliance stream is evaluated using per-plug telemetry from residential-grade appliances. The radar–thermal streams use a single-subject, fixed-placement dataset of binary-motion windows and four-class posture and motion windows collected across six furniture configurations. The separate intrusive load monitoring stream utilises smart plugs to classify appliances. Regarding binary motion recognition, radar (F1,Transformer=0.882±0.034) and thermal (F1,XGBoost=0.870±0.069) pipelines achieved similar macro F1 performance. On the four-class posture and activity recognition task, thermal features (F1,thermal=0.775±0.053) substantially outperformed radar (F1,radar=0.609±0.110). Weighted late fusion produced only modest descriptive gains. Separately, smart plug telemetry demonstrated strong appliance recognition performance using lightweight tree-based models suitable for constrained edge deployment. The results support a scoped redundancy argument. Sparse track-level radar carries gross motion, while low-resolution thermal sensing carries posture. The smart plug appliance monitoring extends the framework toward appliance-mediated instrumental activity of daily living (IADL) monitoring, with lightweight tree-based models achieving strong recognition performance under constrained edge deployment conditions. The findings support a layered multimodal sensing architecture for privacy-preserving ADL monitoring, where radar contributes motion-sensitive coverage, thermal sensing contributes posture-aware spatial context, and smart plug telemetry contributes appliance-level behavioural evidence within intelligent healthcare and ambient assisted living environments. Full article
(This article belongs to the Special Issue AI and Big Data for Smart Healthcare: Ensuring Privacy and Security)
Show Figures

Figure 1

24 pages, 5179 KB  
Article
Software-Only Registration and Cross-Spectral Classification of Unsynchronized RGB–LWIR Video: A Multisensor Benchmark for Conveyor-Based Waste Sorting
by Burak Akdemir and Seniha Esen Yuksel
Sensors 2026, 26(16), 5017; https://doi.org/10.3390/s26165017 - 7 Aug 2026
Viewed by 350
Abstract
Reliable multisensor perception is a key requirement for practical waste sorting, yet many low-cost sensor configurations cannot rely on hardware synchronization or carefully controlled acquisition. We present a pilot-scale multisensor waste-sorting testbed that combines an unsynchronized RGB camera with a long-wave infrared (LWIR) [...] Read more.
Reliable multisensor perception is a key requirement for practical waste sorting, yet many low-cost sensor configurations cannot rely on hardware synchronization or carefully controlled acquisition. We present a pilot-scale multisensor waste-sorting testbed that combines an unsynchronized RGB camera with a long-wave infrared (LWIR) camera for object classification on a continuously moving conveyor, and introduce ThermalRGBTrash, a new paired RGB–LWIR video dataset for this task. To enable fusion under asynchronous acquisition, we develop a fully software-based registration pipeline that combines SuperPoint–SuperGlue matching with an adaptive sliding-window strategy designed to recover from long-wave infrared sensor artifacts, including non-uniformity correction events. Across 281,439 matched frame pairs from 19 paired videos, the registration pipeline achieves a mean spatial alignment error of 2.27 pixels and matches 99.98% of attempted frame pairs. We then detect and segment objects with Mask R-CNN, track them across the conveyor, and classify each tracklet using frozen DINOv2 self-supervised Vision Transformer (ViT-L/14) features with a lightweight multilayer perceptron head. RGB and LWIR representations are combined through late fusion. On 550 tracklets under video-disjoint 10-fold cross-validation, the fused pipeline reaches a macro F1 score of 0.924, outperforming RGB alone (0.886) and LWIR alone (0.856). On a mixed-class test set of 351 tracklets reserved exclusively for final evaluation, fusion reaches a macro F1 score of 0.947. The fusion advantage persists across multiple backbone and pretraining choices, while ablation studies support the chosen temporal sampling and pooling design. These results show that accurate RGB–LWIR object classification is achievable without synchronization hardware, and establish ThermalRGBTrash as a benchmark for future work on practical multisensor perception in conveyor-based waste sorting. Full article
(This article belongs to the Special Issue Multisensor Image and Video Processing: Methods and Applications)
Show Figures

Figure 1

32 pages, 9798 KB  
Article
uVGS-2: The Micro Video Guidance Sensor: A 6-DoF Robust Pose Estimator for Autonomous Proximity Maneuvers in Drones, Spacecraft and Mobile Robot Navigation
by Hector Gutierrez, Jose Cornejo and Ivan Bertaska
Drones 2026, 10(7), 535; https://doi.org/10.3390/drones10070535 - 14 Jul 2026
Cited by 1 | Viewed by 654
Abstract
This paper presents the Micro Video Guidance Sensor Version 2 (uVGS-2), a ROS-based vision navigation framework for real-time six-degrees-of-freedom pose estimation in drones, spacecraft, and autonomous robotic platforms operating in GNSS-denied environments. The system evolves from the previous Smartphone Video Guidance Sensor (SVGS) [...] Read more.
This paper presents the Micro Video Guidance Sensor Version 2 (uVGS-2), a ROS-based vision navigation framework for real-time six-degrees-of-freedom pose estimation in drones, spacecraft, and autonomous robotic platforms operating in GNSS-denied environments. The system evolves from the previous Smartphone Video Guidance Sensor (SVGS) architecture through a modular C++ implementation, including advanced image preprocessing, deterministic blob sorting, and an optimized perspective-4-point solver using a Lie-algebra-based analytical Jacobian formulation. The proposed architecture achieves computationally efficient photogrammetric state estimation using onboard camera and processor resources, enabling deployment in resource-constrained systems. Experimental validation was conducted in NASA’s Astrobee free-flying robot, both at the International Space Station (ISS), for SVGS, and by ground testing through real-time sensor-fusion with Astrobee’s graph-based localizer (Astroloc), for uVGS-2. Results demonstrate robust centimeter-level accuracy in relative position and attitude estimation under illumination disturbances, partial occlusions, and intermittent loss of line-of-sight. The framework can be used in robotic platforms and autonomous UAV operations, including precision landing, formation flight, and cooperative navigation in environments where GNSS signals are unavailable or intermittent. Full article
(This article belongs to the Special Issue Autonomous Drone Navigation in GPS-Denied Environments)
Show Figures

Figure 1

15 pages, 6647 KB  
Article
Adaptive Multi-Temporal Fusion and Cross-Modal Adversarial Alignment for Robust Driver Fatigue Detection
by Yanqiao Feng, Yong Peng and Dennis Z. Yu
Sensors 2026, 26(13), 4298; https://doi.org/10.3390/s26134298 - 6 Jul 2026
Viewed by 606
Abstract
To address the critical challenges of multi-scale temporal dynamics and sensor-intrusiveness in driver fatigue detection, this paper proposes the Multi-Temporal Fusion Attention Network (MTFA-Net). The framework integrates two core innovations: a Multi-scale Temporal Adaptive Fusion (MTAF) module that dynamically weights short-, mid-, and [...] Read more.
To address the critical challenges of multi-scale temporal dynamics and sensor-intrusiveness in driver fatigue detection, this paper proposes the Multi-Temporal Fusion Attention Network (MTFA-Net). The framework integrates two core innovations: a Multi-scale Temporal Adaptive Fusion (MTAF) module that dynamically weights short-, mid-, and long-term behavioral features via a scene-aware modulator, and a Physiological–Behavioral Cross-modal Adversarial Alignment (PBCAA) network that implicitly infers latent physiological states (e.g., HRV) from facial videos using adversarial learning and mutual information maximization. Experimental results on RLDD and NTHU-DDD datasets demonstrate that MTFA-Net achieves state-of-the-art accuracy (92.8%) while maintaining high interpretability and real-time efficiency, providing a robust, non-intrusive solution for intelligent cockpit safety. Full article
Show Figures

Figure 1

47 pages, 7116 KB  
Review
Vision-Based Displacement Measurement for Structural Health Monitoring: A Metrology-Oriented Review of Uncertainty Quantification
by Arman Neyestani, Francesco Picariello, Ioan Tudosa, Michela Monaco, Luca De Vito and Mauro D’Arco
Buildings 2026, 16(13), 2659; https://doi.org/10.3390/buildings16132659 - 4 Jul 2026
Viewed by 757
Abstract
This paper presents a metrology-oriented review of vision-based displacement and deformation measurement for civil structural health monitoring (SHM), with an emphasis on field robustness and uncertainty quantification (UQ). The review focuses on image- and video-based methods that convert visual information into quantitative physical [...] Read more.
This paper presents a metrology-oriented review of vision-based displacement and deformation measurement for civil structural health monitoring (SHM), with an emphasis on field robustness and uncertainty quantification (UQ). The review focuses on image- and video-based methods that convert visual information into quantitative physical measurements, such as displacement, strain, or derived dynamic indicators. The literature is organized according to the main stages of the measurement chain: image formation, image-plane motion estimation, and geometric conversion to metric motion. Within this framework, measurement pipelines are interpreted through three levels of geometric mapping, namely, a scalar scale-factor model, a planar homography-based model, and a full Jacobian-based model. The review synthesizes major method families, including marker-based and markerless tracking, feature-based tracking, optical flow, digital image correlation (DIC), phase-based motion magnification, edge-based estimators, fixed- and moving-camera configurations, UAV-based acquisition with ego-motion compensation, hybrid vision–sensor fusion, and deep-learning-enhanced pipelines. A structured taxonomy of uncertainty sources is then presented along the processing chain, covering camera geometry and calibration, imaging noise and blur, quantization, timing and synchronization, environmental disturbances, optical turbulence and heat haze, platform motion, algorithmic failure modes, and reference-sensor uncertainty. The paper also compares UQ practices, including GUM-aligned analytical propagation, Monte Carlo methods, DIC-specific error budgets, bootstrap and resampling strategies, and probabilistic deep learning. The main contribution of this review is to connect computer-vision-based displacement pipelines with metrological requirements by explicitly linking measurement models, uncertainty sources, UQ methods, and field-validation evidence within a unified framework. A practical uncertainty-budget template is compiled to support traceable reporting across different pipelines and deployment scenarios. The paper concludes with prioritized research gaps and future directions, including standardized benchmarks and datasets, traceable UQ for moving-camera systems, multi-sensor fusion with end-to-end uncertainty propagation, long-term drift characterization, optical-turbulence and adverse-weather modeling, validated subpixel limits at extreme range, probabilistic deep learning–metrology integration, and standardized reporting practices. Full article
(This article belongs to the Special Issue Smart Structures and IoT-Based Health Monitoring for Buildings)
Show Figures

Figure 1

16 pages, 2305 KB  
Article
Continuous Full-Domain Highway Trajectory Tracking Based on Improved Deep-SORT and Inverse Covariance Intersection
by Zheye Tian, Changhuizi Duan, Shijie Gao, Jianling Gu and Nengchao Lyu
Sensors 2026, 26(13), 4251; https://doi.org/10.3390/s26134251 - 4 Jul 2026
Viewed by 299
Abstract
Continuous full-domain vehicle trajectories are essential for smart highway monitoring, but single-sensor roadside perception is limited by physical coverage, occlusion, and environmental sensitivity. To address continuous trajectory tracking across multiple roadside-sensing domains, this study proposes a real-time, full-domain highway trajectory tracking framework based [...] Read more.
Continuous full-domain vehicle trajectories are essential for smart highway monitoring, but single-sensor roadside perception is limited by physical coverage, occlusion, and environmental sensitivity. To address continuous trajectory tracking across multiple roadside-sensing domains, this study proposes a real-time, full-domain highway trajectory tracking framework based on radar–camera fusion, improved Deep-SORT, and inverse covariance intersection. At the local perception level, a two-stage object-level and decision-level fusion model is constructed, and Deep-SORT is improved using a CIoU matching strategy and an occluded target tracking controller to enhance local multi-object tracking continuity. At the cross-domain association level, a geometry-motion consistency stepwise calibration method is developed to unify adjacent sensing domains, and a CATS-ICI trajectory stitching strategy is introduced to improve trajectory association and state smoothness during sensor handover. The proposed framework was validated on a real highway test section with roadside radar, video, and drone-based ground-truth trajectories. Experimental results show that the full local method achieves an EMOTA of 92.35%, and the reconstructed full-domain trajectories achieve a successful trajectory matching rate of 98.4% under the 452 vehicles/10 min test condition. Additional ablation experiments further verify the contributions of radar–camera fusion, CIoU, OTTC, GMCSC, CATS, and ICI. These results demonstrate that the proposed framework can provide continuous and reliable full-domain vehicle trajectories for real-world highway monitoring. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

34 pages, 86423 KB  
Article
FS-YOLOv3: A Reliability-Driven, Temporally Consistent, and Scene-Adaptive Dual-Source Forest Smoke Detector
by Yalei Jia, Fansen Meng, Xufeng Yang, Jisong Zang, Renjie Song and Jianhui Meng
Electronics 2026, 15(13), 2886; https://doi.org/10.3390/electronics15132886 - 1 Jul 2026
Viewed by 372
Abstract
Early smoke detection for forest fire prevention requires accurate and temporally stable decisions under dynamic clutter, tiny long-range targets, atmospheric degradation, and partial sensor unreliability. This paper presents FS-YOLOv3, a reliability-driven RGB–thermal smoke detector that extends a reproduced FS-YOLO baseline with two new [...] Read more.
Early smoke detection for forest fire prevention requires accurate and temporally stable decisions under dynamic clutter, tiny long-range targets, atmospheric degradation, and partial sensor unreliability. This paper presents FS-YOLOv3, a reliability-driven RGB–thermal smoke detector that extends a reproduced FS-YOLO baseline with two new modules: Cross-Temporal Consistency Alignment (CTCA) and Scene-Adaptive Expert Routing Fusion (SAERF). CTCA performs local short-horizon feature alignment and is evaluated with additional offset-field diagnostics to test whether the learned offsets correlate more strongly with annotated smoke expansion than with non-smoke motion. SAERF routes fused features to compact experts according to illumination, haze, texture ambiguity, and thermal reliability, with descriptor ablations and collinearity diagnostics used to examine routing stability. On the proposed clip-level RGB–thermal benchmark, FS-YOLOv3 improves over the reproduced FS-YOLO baseline from 93.7% to 96.3% mAP@0.5 and from 89.5% to 94.8% temporal alarm consistency (TAC), with 165 model FPS on Jetson AGX Orin under the default one-frame-look-ahead buffered inference setting. Comparisons with lightweight YOLO detectors, RGB-only and infrared-only controls, simple fusion strategies, and stronger temporal baselines provide deployment context, while the main technical evidence is the controlled gain obtained by enabling CTCA and SAERF on the same baseline architecture. To support reproducibility, the paper specifies the baseline interface, sensor and annotation protocol, sequence-disjoint split policy, temporal metrics, threshold sensitivity, causal CTCA behavior, SAERF descriptor analysis, and model-side versus end-to-end latency boundaries. The reproducibility package is organized to provide code, configuration files, split identifiers, evaluation scripts, diagnostic-statistic scripts, and illustrative sample annotations; redistribution of the full curated benchmark is handled through institutional data-review approval or controlled access when direct video release is restricted. Full article
Show Figures

Figure 1

34 pages, 5532 KB  
Article
Attention-Based Multimodal Framework for Athlete-Performance Analysis and Rehabilitation Monitoring Using Vision and Wearable Sensors
by Mohammed Alonazi, Iqra Aijaz Abro, Maha Abdelhaq, Raed Alsaqour, Ahmad Jalal and Hui Liu
Bioengineering 2026, 13(7), 718; https://doi.org/10.3390/bioengineering13070718 - 23 Jun 2026
Viewed by 606
Abstract
Background: Advances in monitoring systems featuring wearable sensors, computer vision, and artificial intelligence (AI) have been increasingly used in sports science and rehabilitation practices as a means of movement pattern analysis, injury prevention, and training optimization. These technologies are becoming essential components of [...] Read more.
Background: Advances in monitoring systems featuring wearable sensors, computer vision, and artificial intelligence (AI) have been increasingly used in sports science and rehabilitation practices as a means of movement pattern analysis, injury prevention, and training optimization. These technologies are becoming essential components of athlete-performance analysis and rehabilitation-monitoring systems designed to support biomechanical assessment, athlete development, and movement-quality evaluation. Athlete-performance analysis and rehabilitation monitoring increasingly rely on intelligent multimodal sensing systems capable of continuously evaluating movement quality, biomechanical patterns, training execution, and recovery progress. Human activity recognition (HAR) serves as a key enabling technology for these applications by providing automated assessment of human movement using wearable and vision-based sensing modalities. Therefore, the purpose of this study was to develop and evaluate an attention-based multimodal framework that integrates wearable inertial sensing and RGB video analysis for robust athlete-performance assessment and rehabilitation monitoring through accurate recognition of human movement patterns. Methods: Athlete-performance analysis and rehabilitation monitoring combining inertial sensor data and RGB-based visual information was introduced. Inertial signals were segmented with adaptive windowing, whereas silhouette refinement was performed to analyze motion structures from visual inputs in support of athlete-performance analysis and rehabilitation monitoring. Temporal, spatial, and motion features such as trajectory, orientation, and skeleton-based space-time representations were calculated from multimodal inputs. The proposed framework was designed to capture complex movement dynamics associated with rehabilitation exercises and sports-related motion patterns across heterogeneous sensing environments. Extracted features were then combined and optimized with a multimodal feature fusion approach, while the Ranger optimization algorithm was utilized during the process. An attention-based deep learning classifier was implemented to classify movement activities. Results: The results showed that the proposed framework reached accuracy scores of 88.40% and 87.96% on the VIDIMU dataset and the UTD-MHAD dataset respectively. Recognition performance across both inertial and vision-based modalities provided greater robustness than single-modality solutions. The integration of wearable sensing and computer vision modalities further improved the ability of the framework to analyze complex movement behaviors under varying execution conditions and environmental variations. Conclusion: The proposed multimodal framework provides a foundation for intelligent athlete-performance and rehabilitation-monitoring systems by integrating wearable sensing, computer vision, and attention-based artificial intelligence for robust movement analysis. The findings highlight its potential to support biomechanical assessment, movement-quality evaluation, training-performance monitoring, rehabilitation tracking, and injury-risk management in modern sports and healthcare environments. Full article
Show Figures

Figure 1

31 pages, 30018 KB  
Article
Sensors-Driven Multimodal Deepfake Detection: A Cross-Attention Fusion Approach with Adaptive Modality Gating
by Syeda Sitara Waseem, Noman Shabbir, Syed Rizwan Hassan and KangYoon Lee
Sensors 2026, 26(12), 3695; https://doi.org/10.3390/s26123695 - 10 Jun 2026
Cited by 1 | Viewed by 694
Abstract
Deepfakes threaten sensor-based authentication systems, including biometric sensors, surveillance cameras, and IoT edge devices. Unimodal detectors remain vulnerable to modality-specific attacks. We propose a multimodal deepfake detection framework optimized for resource-constrained edge devices, featuring a novel cross-modal attention fusion mechanism with adaptive gating. [...] Read more.
Deepfakes threaten sensor-based authentication systems, including biometric sensors, surveillance cameras, and IoT edge devices. Unimodal detectors remain vulnerable to modality-specific attacks. We propose a multimodal deepfake detection framework optimized for resource-constrained edge devices, featuring a novel cross-modal attention fusion mechanism with adaptive gating. The architecture combines enhanced Res2Net for audio, temporal 3D CNN with SE attention for video, and bidirectional cross-modal attention with quality-based gates. On our benchmark (5472 audio + 1842 video samples), the fusion model achieves 96.7% accuracy, 96.6% F1-score, 0.988 AUC-ROC, and 3.3% EER. Adversarial testing shows 92.3% accuracy under the Fast Gradient Sign Method (FGSM) attack. The model has a 30.3 MB footprint and runs at 20 FPS on edge hardware. Modality contribution analysis reveals adaptive weighting (72% audio for TTS forgery, 78% video for lip-synced attacks). Cross-dataset evaluation on FakeAVCeleb achieves 92.3% overall accuracy, confirming generalization. Full article
Show Figures

Figure 1

25 pages, 14805 KB  
Article
Hybrid IoT-VIoT System for Real-Time Water-Level Monitoring Using Computer Vision
by Aigul Tungatarova, Gaukhar Borankulova, Aslanbek Murzakhmetov, Bakhyt Yeraliyeva, Saltanat Dulatbayeva, Samat Bekbolatov and Balzhan Turarova
Computers 2026, 15(6), 373; https://doi.org/10.3390/computers15060373 - 7 Jun 2026
Cited by 1 | Viewed by 693
Abstract
Efficient water resource management is critically important for arid regions such as southern Kazakhstan. This paper presents a hybrid Internet of Things (IoT) and Vision-based Internet of Things (VIoT) architecture for real-time monitoring of water levels in irrigation channels. The proposed system integrates [...] Read more.
Efficient water resource management is critically important for arid regions such as southern Kazakhstan. This paper presents a hybrid Internet of Things (IoT) and Vision-based Internet of Things (VIoT) architecture for real-time monitoring of water levels in irrigation channels. The proposed system integrates an ultrasonic water-level sensor, an IP camera with edge-based computer vision processing on a Raspberry Pi, wireless communication, an autonomous solar power supply, and discharge estimation using Manning’s equation. The VIoT subsystem applies image processing techniques, including gauge calibration, Canny edge detection, and pixel-to-metric conversion, to automatically estimate water level from captured video frames. Water-level measurements obtained from IoT sensors and video-based analysis are combined through synchronised data fusion to improve monitoring accuracy and reliability. The hybrid approach leverages the complementary strengths of IoT and VIoT by combining continuous quantitative sensing with visual verification capabilities. Field experiments conducted on the Merke River in the Zhambyl region of Kazakhstan over a 14-day observation period demonstrated stable real-time operation with RMSE = 0.311 cm, MAE = 0.279 cm, and Pearson r = 0.99 between the ultrasonic sensor and the vision-based estimates. Sensitivity analysis indicated that water level is the most influential parameter in Manning-based discharge estimation, confirming the importance of accurate level detection. The proposed system improves reliability by cross-checking independent data sources, making it applicable to monitoring water levels in agricultural regions. Full article
Show Figures

Figure 1

49 pages, 2508 KB  
Review
Sensing the Action: Rethinking Sensor Modalities and Multi-Modal Fusion in Vision–Language–Action Models for Robotic Manipulation
by Byoung Chul Ko
Sensors 2026, 26(11), 3541; https://doi.org/10.3390/s26113541 - 3 Jun 2026
Viewed by 1914
Abstract
Recent Vision–Language–Action (VLA) models have rapidly emerged as general-purpose robotic policies that integrate language understanding, visual perception, and robot control. However, prior studies and surveys have primarily emphasized backbone architectures, action decoders, training recipes, and benchmark performance, whereas relatively limited systematic attention has [...] Read more.
Recent Vision–Language–Action (VLA) models have rapidly emerged as general-purpose robotic policies that integrate language understanding, visual perception, and robot control. However, prior studies and surveys have primarily emphasized backbone architectures, action decoders, training recipes, and benchmark performance, whereas relatively limited systematic attention has been given to sensor modality selection, heterogeneous signal alignment and fusion, and their connection to action generation, all of which are critical to the performance and safety of real-world robotic manipulation. This survey addresses this gap by reinterpreting VLA within the framework of a sensor–fusion–action pipeline. This study first presents a systematic taxonomy of major sensor modalities, including RGB, depth, tactile sensing, force/torque, proprioception and inertial measurement unit, multi-spectral/thermal, and event-based vision, and compares them in terms of the physical information they provide, their characteristic failure modes, and their deployment constraints. This survey further reviews teleoperation-, human video-, and simulation-based data collection pipelines, together with representative dataset configurations, and analyzes the multi-modal design space from a sensor-centric perspective, including early and late fusion, cross-attention, token-level fusion, adapters, mixture of experts, and multi-rate action representations. In addition, this study identifies a strong bias in existing benchmarks toward RGB-centric inputs and single success-rate metrics and emphasizes the need for a multidimensional evaluation framework incorporating robustness, worst-case performance, safety, latency, and efficiency. By shifting the focus away from a model-centric narrative and explicitly accounting for real-world sensor complexity, this survey seeks to establish a sensor-centered foundation for the next generation of Physical AI. Full article
(This article belongs to the Special Issue Feature Review Papers in Sensors and Robotics)
Show Figures

Figure 1

23 pages, 22564 KB  
Article
A Multi-Module Fusion Framework for Restoring Human and Machine Vision Quality in Compressed Video
by Keren He, Kun Xiang, Yufei Gao, Yang Yu and Jinjia Zhou
Sensors 2026, 26(11), 3494; https://doi.org/10.3390/s26113494 - 1 Jun 2026
Viewed by 476
Abstract
With the increasing demand for video processing in both human perception and machine vision applications, enhancing heavily compressed video has become a critical problem in practical multimedia systems. In many real-world scenarios, video data acquired by image sensors are often compressed for efficient [...] Read more.
With the increasing demand for video processing in both human perception and machine vision applications, enhancing heavily compressed video has become a critical problem in practical multimedia systems. In many real-world scenarios, video data acquired by image sensors are often compressed for efficient transmission and storage, which introduces compression artifacts and degrades both visual quality and downstream task performance. This issue is especially significant in sensor-based systems such as surveillance cameras and mobile imaging devices. To address these challenges, we propose a novel joint human–machine video enhancement framework for compressed video enhancement that jointly targets human perceptual quality and machine vision performance. The framework integrates four complementary components: a Spatio-Temporal Fusion Module that leverages inter-frame correlations, a High-Frequency Semantic Fusion module for recovering structurally important details relevant to machine tasks, a Texture-Guided Model that enhances low-level visual features, and a Refined Attention Residual Quality Enhancement Module that adaptively emphasizes salient regions. By progressively combining these modules, the framework effectively restores compressed content while preserving task-relevant semantics. The experimental results demonstrate that our method consistently outperforms existing approaches, achieving higher PSNR and SSIM as well as improved object detection and video object segmentation performance. These results highlight the framework’s practical applicability for compressed video enhancement in sensor-based systems, including intelligent surveillance and autonomous imaging platforms. Full article
(This article belongs to the Special Issue Advances in Learning-Based Sensing-Driven Multimedia Processing)
Show Figures

Figure 1

33 pages, 5543 KB  
Article
The New Frontier of Quality Evaluation for Visual Sensors: A Survey of Large Multimodal Model-Based Methods
by Qihang Ge, Xiongkuo Min, Sijing Wu, Yunhao Li and Guangtao Zhai
Sensors 2026, 26(8), 2530; https://doi.org/10.3390/s26082530 - 20 Apr 2026
Viewed by 1279
Abstract
Visual quality assessment is entering a new frontier as media evolve from static images to temporally dynamic videos and 3D content. These visual signals are typically captured by sensing devices such as cameras and depth sensors, whose acquisition characteristics significantly influence perceptual quality. [...] Read more.
Visual quality assessment is entering a new frontier as media evolve from static images to temporally dynamic videos and 3D content. These visual signals are typically captured by sensing devices such as cameras and depth sensors, whose acquisition characteristics significantly influence perceptual quality. Traditional quality models, including distortion-centric and regression-based approaches, perform well on conventional degradations but struggle to evaluate higher-level attributes such as semantic plausibility and structural coherence in modern AI-generated and multimodal scenarios. The emergence of large multimodal models (LMMs), including vision–language models (VLMs) and multimodal large language models (MLLMs), reshapes the evaluation paradigm by enabling semantic grounding, instruction-driven assessment, and explainable reasoning. This survey presents a unified perspective on visual quality assessment for sensor-captured visual data across image, video, and 3D modalities. We review conventional deep learning approaches and recent LMM-based methods, highlighting how multimodal fusion and language-conditioned reasoning transform quality assessment from scalar prediction to perceptual intelligence. Finally, we discuss key challenges and future opportunities for building efficient, robust, and sensor-aware visual quality assessment systems. Full article
(This article belongs to the Special Issue Perspectives in Intelligent Sensors and Sensing Systems)
Show Figures

Figure 1

52 pages, 18820 KB  
Article
Multimodal Industrial Scene Characterisation for Pouring Process Monitoring Using a Mixture of Experts
by Javier Nieves, Javier Selva, Guillermo Elejoste-Rementeria, Jorge Angulo-Pines, Jon Leiñena, Xuban Barberena and Fátima A. Saiz
Appl. Sci. 2026, 16(7), 3430; https://doi.org/10.3390/app16073430 - 1 Apr 2026
Viewed by 758
Abstract
Industrial pouring processes operate under highly dynamic conditions where small deviations can lead to defects, scrap, and production losses. Although modern foundries are equipped with multiple sensors and visual inspection systems, most monitoring approaches remain fragmented, unimodal, and difficult to interpret. Furthermore, annotated [...] Read more.
Industrial pouring processes operate under highly dynamic conditions where small deviations can lead to defects, scrap, and production losses. Although modern foundries are equipped with multiple sensors and visual inspection systems, most monitoring approaches remain fragmented, unimodal, and difficult to interpret. Furthermore, annotated anomalous samples in industrial settings are scarce, hindering the development of traditional methods. As a result, many critical pouring anomalies are detected too late or lack sufficient contextual information for effective decision making. In this work, we propose a multimodal framework for industrial scene characterisation that combines visual information and process signals through an explainable Mixture-of-Experts (MoE)-style expert-fusion strategy. First, we deploy an ensemble of specialised modules that collaborate to identify regions of interest, assess pouring quality, and contextualise events within the production process, thereby generating an interpretable description of pouring events. Second, we introduce a novel anomaly detection method for multimodal video data, combining a self-supervised transformer with an outlier-aware clustering algorithm. Our approach effectively identifies rare anomalies without requiring extensive manual labelling. The resulting information is structured into a digital twin-ready representation, supporting synchronisation between the physical system and its virtual counterpart. This solution provides a scalable, deployable pathway to transform heterogeneous industrial data into actionable knowledge, supporting advanced monitoring, anomaly detection, and quality control in real foundry environments. Full article
Show Figures

Figure 1

Back to TopTop