Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (47)

Search Parameters:
Keywords = event–frame fusion

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
28 pages, 8603 KB  
Article
Event-Guided Image Reconstruction for Nighttime Dynamic Scenes
by Qingjiao Meng, Ji Li and Yan Jin
J. Imaging 2026, 12(9), 399; https://doi.org/10.3390/jimaging12090399 (registering DOI) - 23 Aug 2026
Abstract
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, [...] Read more.
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, we propose an event-guided image reconstruction method for nighttime dynamic visual perception. The method constructs a multi-channel event voxel representation by jointly encoding event count, event intensity, timestamp distribution, and blurred-frame intensity priors. A parameter-efficient local–global reconstruction network is then designed to restore fine-grained textures and model holistic structures. In addition, edge-alignment and blur-alignment constraints are introduced to improve geometric consistency and imaging plausibility. Experiments on the HQF and REDS datasets show that the proposed method outperforms existing methods in the MSE, PSNR, and SSIM. Compared with DeblurSR, it reduces the MSE by 14.81% on HQF and 10.00% on REDS, while improving the PSNR by 1.603 dB and 1.053 dB, respectively. Qualitative results further show sharper edges, lower structural errors, and better edge consistency. Low illumination, dynamic blur, rapid brightness variation, and event noise are also common degradation factors in nighttime UAV imaging, making the investigated problem technically relevant to that setting. However, because neither REDS nor HQF was acquired during an actual UAV flight, the reported results establish benchmark-level reconstruction performance rather than UAV-specific operational effectiveness. Full article
39 pages, 11582 KB  
Article
A Dual-Camera Edge Sensing Framework with Zone-Aware Multi-Object Tracking for Sensorless Smart Vending Cabinets
by Abror Shavkatovich Buriboev, Farkhat Rajabov, Shavkat Buriboev, Rustem Allanyazov, Giyosjon Sharipov, Abbos Abduvaytov, Aziza Akhmedova, Ruzimboy Sobirov, Su-Mi Shin, Cheolwon Lee and Heung Seok Jeon
Sensors 2026, 26(16), 5213; https://doi.org/10.3390/s26165213 - 17 Aug 2026
Viewed by 338
Abstract
Top-loading smart vending cabinets require precise transaction-level product detection under strict hardware and deployment constraints. In this paper, “sensorless” refers specifically to the absence of auxiliary product-level sensing hardware, such as RFID tags, weight sensors, shelf load cells, or product-slot instrumentation; the system [...] Read more.
Top-loading smart vending cabinets require precise transaction-level product detection under strict hardware and deployment constraints. In this paper, “sensorless” refers specifically to the absence of auxiliary product-level sensing hardware, such as RFID tags, weight sensors, shelf load cells, or product-slot instrumentation; the system still uses two camera sensors. Conventional snapshot-difference methods compare only a small number of frames at the beginning and end of a transaction and therefore cannot explicitly represent intermediate product motion, such as pickup, return, inspection, occlusion, and shelf resettling. This paper proposes ZAB-Fusion, a dual-camera edge sensing framework with zone-aware multi-object tracking for sensorless smart vending cabinets. The framework combines a YOLO11-seg and RT-DETR detection ensemble with ByteTrack temporal association, projects product tracks into a three-zone vertical cabinet model, interprets compressed zone sequences using a finite-state event classifier, and integrates camera-specific event streams through an evidence-gated cross-camera fusion rule. The proposed method was evaluated on 220 in-service vending transactions containing 227 ground-truth TAKEN events and 87 RETURNED events across seven product classes. Compared with the snapshot-difference baseline, ZAB-Fusion improved recall from 0.665 to 0.925 and F1-score from 0.780 to 0.944, while maintaining a high precision of 0.963. At the transaction level, exact receipt accuracy increased from 0.645 to 0.900. Runtime analysis on an Intel N100 CPU-only edge device showed an average processing latency of 562 ms per transaction under the selected-frame inference protocol. The zone classification, finite-state event interpretation, and evidence-gated cross-camera fusion stages required only 3 ms in total. These results demonstrate that explicit motion semantics and auditable cross-camera evidence gating can improve sensorless retail transaction level recognition in sensorless smart vending cabinets without adding auxiliary product-level sensing hardware. Full article
Show Figures

Figure 1

57 pages, 39305 KB  
Review
Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision
by Paul K. J. Park, Junseok Kim and Juhyun Ko
Sensors 2026, 26(16), 5127; https://doi.org/10.3390/s26165127 - 13 Aug 2026
Viewed by 403
Abstract
Frame-based RGB image sensors and event-based vision sensors provide complementary sensing capabilities for human-perceptual imaging and machine vision. RGB image sensors capture dense spatial, color, and texture information that is essential for human-viewable imaging, semantic recognition, and conventional image signal processing pipelines. In [...] Read more.
Frame-based RGB image sensors and event-based vision sensors provide complementary sensing capabilities for human-perceptual imaging and machine vision. RGB image sensors capture dense spatial, color, and texture information that is essential for human-viewable imaging, semantic recognition, and conventional image signal processing pipelines. In contrast, dynamic vision sensors (DVSs) and event vision sensors (EVSs) asynchronously detect local brightness changes and provide sparse temporal information with low latency, high temporal resolution, and reduced redundant data output. Because neither modality alone satisfies all requirements of emerging vision systems, hybrid event–frame sensing has become an important direction for compact, low-latency, and energy-efficient sensing. This review presents a sensor-oriented taxonomy of hybrid event–frame sensing architectures and systems, including dual-camera event–frame systems, optically aligned event–frame systems, pixel-level shared hybrid image sensors, stacked CIS–DVS hybrid image sensors, homogeneous-pixel sensing systems, and event-only reconstruction systems. We analyze key sensor specifications, including latency, spatial resolution, color fidelity, power consumption, and form factor, and discuss how these specifications guide sensor configuration and design. The review identifies stacked CIS–DVS sensors as one of the most balanced and competitive architectures because they can support compact integration, synchronized event–frame sensing, and on-chip processing. However, important challenges remain, including color fidelity, demosaicing, event-pixel ratio optimization, calibration, benchmarking, and edge-AI deployment. Finally, we emphasize that future hybrid event–frame sensing systems should be developed through sensor–algorithm–ISP–AI co-design. This review provides practical guidelines for developing next-generation hybrid event–frame sensing systems for both human-perceptual imaging and machine vision. Full article
(This article belongs to the Special Issue Computer Vision-Based Human Activity Recognition)
Show Figures

Figure 1

24 pages, 5179 KB  
Article
Software-Only Registration and Cross-Spectral Classification of Unsynchronized RGB–LWIR Video: A Multisensor Benchmark for Conveyor-Based Waste Sorting
by Burak Akdemir and Seniha Esen Yuksel
Sensors 2026, 26(16), 5017; https://doi.org/10.3390/s26165017 - 7 Aug 2026
Viewed by 292
Abstract
Reliable multisensor perception is a key requirement for practical waste sorting, yet many low-cost sensor configurations cannot rely on hardware synchronization or carefully controlled acquisition. We present a pilot-scale multisensor waste-sorting testbed that combines an unsynchronized RGB camera with a long-wave infrared (LWIR) [...] Read more.
Reliable multisensor perception is a key requirement for practical waste sorting, yet many low-cost sensor configurations cannot rely on hardware synchronization or carefully controlled acquisition. We present a pilot-scale multisensor waste-sorting testbed that combines an unsynchronized RGB camera with a long-wave infrared (LWIR) camera for object classification on a continuously moving conveyor, and introduce ThermalRGBTrash, a new paired RGB–LWIR video dataset for this task. To enable fusion under asynchronous acquisition, we develop a fully software-based registration pipeline that combines SuperPoint–SuperGlue matching with an adaptive sliding-window strategy designed to recover from long-wave infrared sensor artifacts, including non-uniformity correction events. Across 281,439 matched frame pairs from 19 paired videos, the registration pipeline achieves a mean spatial alignment error of 2.27 pixels and matches 99.98% of attempted frame pairs. We then detect and segment objects with Mask R-CNN, track them across the conveyor, and classify each tracklet using frozen DINOv2 self-supervised Vision Transformer (ViT-L/14) features with a lightweight multilayer perceptron head. RGB and LWIR representations are combined through late fusion. On 550 tracklets under video-disjoint 10-fold cross-validation, the fused pipeline reaches a macro F1 score of 0.924, outperforming RGB alone (0.886) and LWIR alone (0.856). On a mixed-class test set of 351 tracklets reserved exclusively for final evaluation, fusion reaches a macro F1 score of 0.947. The fusion advantage persists across multiple backbone and pretraining choices, while ablation studies support the chosen temporal sampling and pooling design. These results show that accurate RGB–LWIR object classification is achievable without synchronization hardware, and establish ThermalRGBTrash as a benchmark for future work on practical multisensor perception in conveyor-based waste sorting. Full article
(This article belongs to the Special Issue Multisensor Image and Video Processing: Methods and Applications)
Show Figures

Figure 1

21 pages, 2146 KB  
Article
A Multi-Stage Dual Encoder–Decoder Network Based on Event Image Cross-Modal Fusion for Image Deblurring
by Yan Liu, Yanfei Jia, Sheng Qiang, Yongpei Lin and Liquan Zhao
Sensors 2026, 26(15), 4762; https://doi.org/10.3390/s26154762 - 27 Jul 2026
Viewed by 355
Abstract
Most existing event-driven image deblurring methods ignore inherent differences between the two modalities and lack explicit alignment strategies, leading to cross-modal mismatches and degraded feature reconstruction. To address this issue, a multi-stage dual encoder–decoder image deblurring method based on event image cross-modal fusion [...] Read more.
Most existing event-driven image deblurring methods ignore inherent differences between the two modalities and lack explicit alignment strategies, leading to cross-modal mismatches and degraded feature reconstruction. To address this issue, a multi-stage dual encoder–decoder image deblurring method based on event image cross-modal fusion is proposed. The proposed network consists of an encoder and a decoder. The encoder employs dilated convolutional residual modules for feature extraction. It also integrates a cross-modal feature fusion module and a local scoring mechanism. These components combine event features with frame image features while suppressing noise. The decoder reconstructs image features via two directional decoding sub-networks. It also incorporates a feedback attention module. This module selects informative features along the feedback path. As a result, the image reconstruction quality is enhanced. In addition to the standard loss, mean absolute error, structural similarity, and frequency reconstruction losses are used to optimize deblurring performance. Extensive experiments are conducted on the GoPro, REBlur, and RwEvent datasets. For PSNR, our method exceeds REFID by 0.23 dB, 0.20 dB, and 0.49 dB on the three datasets. For SSIM, our model achieves gains of 0.002, 0.002, and 0.017 against REFID. In terms of computational cost and inference speed, our network adds only 3.4 M parameters and 105.2 GFLOPs, with an FPS reduction of only 2.12. This trivial efficiency loss delivers significant improvements in both pixel and structural restoration performance. Ablation experiments verify the independent positive contribution of each designed module. Both qualitative visual comparisons and quantitative metrics demonstrate that the proposed network has stronger deblurring and generalization capabilities. Full article
(This article belongs to the Special Issue AI-Based Sensing and Imaging Applications)
Show Figures

Figure 1

27 pages, 11969 KB  
Article
ULSTM: Multi-Scale and Full-Level Temporal Consistency for Traffic Anomaly Detection
by Borja Pérez, Mario Resino, Jaime Godoy, Abdulla Al-Kaff and Fernando García
Smart Cities 2026, 9(7), 120; https://doi.org/10.3390/smartcities9070120 - 22 Jul 2026
Viewed by 431
Abstract
Urban traffic anomaly detection is essential for intelligent transportation systems, particularly in smart city environments where fast identification of abnormal events can improve road safety and traffic management. This work proposes a novel ULSTM-driven architecture that explicitly models temporal dependencies across consecutive traffic [...] Read more.
Urban traffic anomaly detection is essential for intelligent transportation systems, particularly in smart city environments where fast identification of abnormal events can improve road safety and traffic management. This work proposes a novel ULSTM-driven architecture that explicitly models temporal dependencies across consecutive traffic frames to achieve more stable and temporally coherent reconstructions. The proposed framework leverages sequential spatio-temporal representations to improve the distinction between normal traffic patterns and anomalous events. To further enhance reliability, we introduce a Hybrid Weighted Fusion strategy that synergistically combines structural, perceptual and pixel-wise metrics. The framework’s parameters are optimized using a Discrete Dirichlet Sampling approach, achieving a peak F1 Score of 70.28%. Evaluations were conducted on a manually curated traffic anomaly dataset with frame-level annotations. Experimental results demonstrate that the ULSTM framework significantly outperforms frame-independent generative models by suppressing high-frequency reconstruction noise, providing a robust solution for real-world smart city deployments. While highly effective in complex scenarios, the proposed framework is strictly applicable to highly dynamic traffic environments with active motion, as static background ensembles can degrade performance. Full article
(This article belongs to the Section Smart Urban Mobility, Transport, and Logistics)
Show Figures

Figure 1

53 pages, 2103 KB  
Article
Sequence-Anchored Shared Tumor-Specific Epitopes for Pre-Manufactured HLA-Matched mRNA Cancer Vaccine Libraries: A Pan-Cancer Framework
by Sarfaraz K. Niazi
Biomolecules 2026, 16(7), 1015; https://doi.org/10.3390/biom16071015 - 11 Jul 2026
Viewed by 626
Abstract
A single vaccine cannot prevent or treat all cancers; however, recurrent tumor-specific epitopes may facilitate the development of pre-manufactured, HLA-matched mRNA vaccines tailored for specific molecular subgroups. We define the shared tumor-specific epitope as a recurring peptide derived from a viral oncoprotein, a [...] Read more.
A single vaccine cannot prevent or treat all cancers; however, recurrent tumor-specific epitopes may facilitate the development of pre-manufactured, HLA-matched mRNA vaccines tailored for specific molecular subgroups. We define the shared tumor-specific epitope as a recurring peptide derived from a viral oncoprotein, a driver mutation, a frameshift, an altered protein C-terminus, or a fusion junction, and we employ a rigorous cancer-cell-only criterion: a target must be recurrent within a defined subgroup, absent from essential normal tissues at the peptide–HLA level, naturally presented on tumor cells, and sufficiently clonal to minimize immune escape. Under this criterion, we present fifteen sequence-anchored reference designs alongside one conceptual placeholder across thirteen candidates divided into four superclasses: viral oncoproteins (such as HPV16/18 E6 and E7 as attenuated antigenic reference designs; Merkel cell polyomavirus serving as a design-specific placeholder), recurrent driver neoepitopes (including KRAS G12/G13, IDH1 R132H, and H3 K27M), hematologic neoantigens (such as NPM1 Type A C-terminus; and a single CALR exon 9 construct encoding the shared novel C-terminus of types 1 and 2 mutations), and fusion junctions (notably EWS-FLI1 and BCR-ABL). Each open reading frame is anchored to a canonical accession with its documented event; representative ORFs are provided as reference designs, with the intended residue-level verification records. These sequence designs are intended as reference constructs and are not suitable as clinical-grade or manufacturing-ready products; they require independent residue-level validation and comprehensive safety assessments prior to laboratory or clinical application. The historical record of non-personalized vaccination—including HPV and hepatitis B prophylaxis, intravesical BCG, and unsuccessful tumor-associated antigen trials—frames both the potential and limitations of such approaches. The practical product is not a universal vaccine but rather a governed library aligned with specific genotype, viral etiology, HLA context, and clinical setting. Currently, none of these designs have established proof-of-benefit-tier evidence. Full article
Show Figures

Figure 1

21 pages, 2178 KB  
Article
Event-Driven Highlight Generation in Football Broadcasts: A Context-Aware Temporal Framework for Six-Class Action Spotting
by Khalil M. Abdelnaby
Appl. Sci. 2026, 16(14), 6838; https://doi.org/10.3390/app16146838 - 8 Jul 2026
Viewed by 400
Abstract
This paper presents an end-to-end system for the automatic summarization of football matches through temporally accurate event spotting across six action categories: goal, card, substitution, kick-off, direct free-kick, and corner. These six classes were selected from the SoccerNet-v2 annotation taxonomy as a non-standard [...] Read more.
This paper presents an end-to-end system for the automatic summarization of football matches through temporally accurate event spotting across six action categories: goal, card, substitution, kick-off, direct free-kick, and corner. These six classes were selected from the SoccerNet-v2 annotation taxonomy as a non-standard evaluation subset: the three standard SoccerNet-v2 classes (goal, card, substitution) are retained and extended with three additional context-dependent set-piece classes (kick-off, direct free-kick, and corner) that represent the primary unresolved challenge in temporal action spotting. We developed and empirically evaluated a framework based on a Temporal Convolutional Network (TCN) with capsule-based relational encoding applied to class-specific Time-Shift Encoded (TSE) labels using a composite segmentation-detection loss. Evaluated on the SoccerNet-v2 benchmark (300/100/100 train/validation/test splits) using the official PCA-compressed ResNet-152 features sampled at 2 FPS, the implemented TCN-Capsule baseline achieved a test-set mean Average Precision (mAP) of 58.16% at a 10 s tolerance threshold (δ = 20 frames), with a validation mAP of 58.55%, indicating strong generalization with only a 0.39 pp performance gap. The six per-class AP values underlying this result are Goal: 75.6%, Corner: 78.0%, Substitution: 56.9%, Card: 53.9%, Direct Free-kick: 52.8%, and Kick-off: 34.3%; the headline mAP is computed by the official SoccerNet area-under-precision-recall-curve script and is not a simple arithmetic average of these six values. In addition, five complementary evaluation metrics are introduced: Average Temporal Error (ATE), Temporal Recall@K (TR@K), Latency-to-Event (L2E), Temporal Causality Consistency (TCC), and Annotation Noise Sensitivity. The analysis reveals that visually salient events such as goals and corners are detected with high accuracy (AP > 75%, ATE < 5 s, AUC > 0.95 at frame-level), whereas context-dependent events, including kick-offs and direct free-kicks, remain challenging (AP < 55%, ATE > 8 s, AUC = 0.68–0.72). As future work, two architectural extensions are formally specified but not empirically evaluated: a Temporal Segment Network with Non-Local Blocks and a hybrid Transformer-CNN framework with optical flow fusion. Furthermore, an OpenCV-based highlight generation prototype is implemented; formal quality evaluation is planned as future work. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

33 pages, 11337 KB  
Article
Video-Based Detection of Dairy Cow Hoof-Slipping Behaviour Using Improved DeepLabCut and NeuFlow v2
by Yue Nian, Kaixuan Zhao, Jiangtao Ji, Yinan Chen and Ruihong Zhang
Animals 2026, 16(13), 2103; https://doi.org/10.3390/ani16132103 - 7 Jul 2026
Viewed by 410
Abstract
Hoof slipping in dairy cows is a subtle, transient hoof motion event distinct from lameness or falling, with short duration, limited displacement, and close resemblance to normal gait, making automated detection particularly challenging; relevant methods remain scarce. This study proposes a cascaded detection [...] Read more.
Hoof slipping in dairy cows is a subtle, transient hoof motion event distinct from lameness or falling, with short duration, limited displacement, and close resemblance to normal gait, making automated detection particularly challenging; relevant methods remain scarce. This study proposes a cascaded detection framework based on improved DeepLabCut and NeuFlow v2 for automated hoof-slipping detection and distance estimation in Holstein dairy cows. The four-stage framework covers hoof key point localization, pixel-level optical flow fusion, motion parameter curve feature extraction, and Random Forest classification. The framework was developed on Dataset 1, which contained 115 single-cow side-view videos. Of these, 31 contained slipping events and 84 were normal walking. It was further assessed on a smaller second-farm dataset of 17 single-cow videos (Dataset 2). ResNet-50 with a Coordinate Attention mechanism was adopted as the backbone, reducing mean four-hoof localization RMSE to 2.80 pixels across five independent training runs, showing a 15.2% improvement over the baseline, and outperforming YOLOv8s-Pose. NeuFlow v2 was applied to extract the localized optical flow from hoof regions, yielding velocity and directional curves from which slipping features were derived. The Random Forest classifier achieved an accuracy of 98.9%, precision of 93.3%, recall of 90.3%, F1 score of 91.8%, and AUC of 0.995, outperforming MViT, SlowFast, and STME. The slipping distance estimation RMSE was 1.22 pixels. With the localisation model retrained on new farm frames, the method reached comparable performance on the second farm, suggesting preliminary cross-farm generalisability that warrants larger-scale validation. The proposed framework provides a non-invasive basis for early hoof-health monitoring and welfare-oriented farm management. Full article
(This article belongs to the Section Cattle)
Show Figures

Figure 1

23 pages, 5420 KB  
Article
Real-Time Detection of Rare Traffic Situations Using RGB-LiDAR Fusion and a Rule-Based Safety Agent in CARLA
by Matúš Čávojský, Matúš Dopiriak, Eugen Šlapak, Arisha Al Faruque, Tomáš Doboš and Gabriel Bugár
Appl. Sci. 2026, 16(13), 6722; https://doi.org/10.3390/app16136722 - 5 Jul 2026
Viewed by 453
Abstract
Rare and safety-critical traffic situations remain challenging for autonomous driving (AD) because they are underrepresented in common training data and may include objects outside standard detector classes. This paper presents a real-time RGB-LiDAR fusion framework for detecting and reacting to rare traffic situations [...] Read more.
Rare and safety-critical traffic situations remain challenging for autonomous driving (AD) because they are underrepresented in common training data and may include objects outside standard detector classes. This paper presents a real-time RGB-LiDAR fusion framework for detecting and reacting to rare traffic situations in CARLA (Car Learning to Act), a reproducible simulator for AD research. The approach combines YOLOv8n-based RGB perception, bird’s-eye-view (BEV) LiDAR clustering, decision-level fusion, an interpretable rule-based safety agent with hysteresis, Time-to-Collision (TTC)-aware escalation, and an automatic emergency braking (AEB) override above the CARLA autopilot. Fused observations are classified as semantic–geometric detections, semantic-only detections, or geometric-only obstacle candidates, where unmatched LiDAR clusters are treated conservatively as candidate-level physical evidence rather than confirmed rare objects. The framework was evaluated on three CARLA maps and 3CSim-inspired corner-case scenarios comprising 19,253 frames, with additional weather/lighting stress tests and a public nuScenes mini cross-platform check. On a manually annotated subset of 4800 CARLA frames, corresponding to approximately 24.9% of the recorded CARLA log, the full framework achieved 96.2% precision, 97.3% recall, and a 96.7% F1-score for safety-relevant threat detection. The control experiments show that the fusion-based safety agent reduced unnecessary braking to 1.7% compared with 8.6% for the LiDAR-only baseline and achieved event-level success on the annotated critical intervals. The proposed CPU-only implementation maintained real-time performance, with an average processing time of 34.7ms. Full article
Show Figures

Figure 1

21 pages, 72670 KB  
Article
Dense Optical Flow Retrieval of Wildfire Smoke Plume Motion from Spaceborne and Airborne Imagery
by Igor Yanovsky, Nicholas LaHaye, Olga V. Kalashnikova, Derek J. Posselt and William C. Porter
Remote Sens. 2026, 18(12), 1868; https://doi.org/10.3390/rs18121868 - 6 Jun 2026
Viewed by 622
Abstract
This paper evaluates a dense, total-variation-based optical flow method for retrieving wildfire smoke plume motion vectors from geostationary, deep-space, and airborne remote sensing imagery. Using multiple major fire events, we assess the robustness of the approach across a range of spatial resolutions and [...] Read more.
This paper evaluates a dense, total-variation-based optical flow method for retrieving wildfire smoke plume motion vectors from geostationary, deep-space, and airborne remote sensing imagery. Using multiple major fire events, we assess the robustness of the approach across a range of spatial resolutions and time intervals. The test cases include Geostationary Operational Environmental Satellite (GOES) observations of the 2025 Los Angeles Fires and the 2024 Park Fire, imagery from NASA’s Enhanced MODIS Airborne Simulator (eMAS) for the 2019 Sheridan and Williams Flats Fires, and a complementary Park Fire image pair from the Earth Polychromatic Imaging Camera (EPIC) aboard the Deep Space Climate Observatory (DSCOVR). Optical flow is computed directly on radiance fields, and smoke plumes are isolated using smoke masks derived from the Segmentation, Instance Tracking, and data Fusion Using multi-SEnsor imagery (SIT-FUSE) framework where available. Performance is evaluated by comparing the root mean square error (RMSE) between original image pairs and between the first image and the second image after warping with the retrieved motion field. RMSE is computed both globally and over smoke-only regions. Across GOES and eMAS cases, optical flow systematically reduces RMSE, often by more than a factor of two within smoke regions, indicating substantially improved frame-to-frame alignment of plume structures after motion correction. The DSCOVR/EPIC case, despite its coarser spatial resolution and longer temporal separation, also shows a marked reduction in global RMSE, demonstrating that the method remains informative under a broader range of observational conditions. For a selected subset of 10 consecutive GOES Park Fire pairs, we additionally compare the retrieved smoke motion vectors with collocated winds from the High-Resolution Rapid Refresh (HRRR) model and find the closest agreement in a broad lower-tropospheric layer centered near 875 hPa. These results show that dense optical flow can capture fine-scale plume evolution in high-temporal-resolution datasets while also providing useful motion estimates in coarser, global-view imagery. RMSE reduction is interpreted here as evidence of improved motion-compensated alignment, while the HRRR comparison provides initial physical context rather than independent validation. The resulting smoke motion vector fields provide a foundation for future comparison with model winds and for applications in plume analysis, fire hazard monitoring, and air quality studies. Full article
Show Figures

Figure 1

37 pages, 6289 KB  
Article
An Indoor Occupancy Detection Method and Application by Fusing Field-of-View Information and Events with a Single Camera
by Pengchen Chen, Chuang Wang and Jingjing An
Buildings 2026, 16(11), 2133; https://doi.org/10.3390/buildings16112133 - 26 May 2026
Viewed by 450
Abstract
Accurate and stable indoor occupancy information is essential for occupant-based intelligent ventilation control. Under a single-camera setting, existing indoor occupancy detection methods commonly suffer from missed detections caused by occlusion and blind zones, false detections caused by people outside the room, and cumulative [...] Read more.
Accurate and stable indoor occupancy information is essential for occupant-based intelligent ventilation control. Under a single-camera setting, existing indoor occupancy detection methods commonly suffer from missed detections caused by occlusion and blind zones, false detections caused by people outside the room, and cumulative entry–exit errors that are difficult to correct. These problems lead to false fluctuations in detected occupancy, affect control performance, and may further reduce indoor comfort or cause unnecessary energy use. To address the practical situation in which indoor spaces are commonly equipped with a single security camera, this study proposes an indoor occupancy detection method by fusing field-of-view information and entry–exit events with a single camera. The study covers method development, multi-scenario validation, parameter analysis, and a ventilation control application. The proposed method uses YOLOv8x and DeepSORT as front-end models and performs post-processing on their outputs to extract field-of-view occupancy information, entry–exit events, and blind-zone events. An occupancy confirmation and correction module is then constructed. The blind-zone event mechanism reduces the influence of missed entry–exit events and camera blind zones on occupancy judgment. The correction module integrates frame-by-frame ID counts, historical outputs, and multiple event signals to verify and suppress false occupancy changes caused by false detections, missed detections, and blind zones, thereby producing more stable indoor occupancy results. Experimental results show that the proposed method outperforms the baseline methods based on front-end object detection and tracking in terms of score, RMSE, and F1 score in three typical scenarios: an office, a home, and a classroom. In the office scenario, the proposed method achieved a score of 99.36%, an RMSE of 0.081, and an F1 score of 0.781. The detection stability was also improved in the home and classroom scenarios. In the high-density and strongly occluded classroom scenario, the absolute detection performance of the fusion-based detection method was limited by the front-end models, indicating that the method still has certain applicability boundaries in complex high-density scenes. Parameter sensitivity analysis shows that key parameters, including the entry–exit area depth, confidence threshold, and time threshold, affect the detection results of the fusion-based detection method. Under the test conditions of this study, the method performs well when the entry–exit area depth is approximately 1.5d, the YOLOv8x confidence threshold is 40%, and the time threshold is 5 × FPS. These results can provide a reference for initial parameter setting and on-site calibration in similar scenarios. Using the office scenario as a case study, the method was further applied to occupant-based ventilation control. The average CO2 concentration during occupied periods under the proposed method was 622.43 ppm, which was closest to the result under ground-truth occupancy control, with a deviation of only 0.9 ppm. This indicates that the method can help improve indoor air quality. Compared with conventional schedule-based control, occupant-based ventilation control driven by the proposed fusion method reduced cumulative fan energy consumption by approximately 65.2%, showing good energy-saving potential at the ventilation-control level. In summary, the proposed method can effectively improve the accuracy and stability of indoor occupancy detection under a single-camera setting and provide more reliable input for occupant-based ventilation control. The framework is modular, and the front-end object detection and tracking models can be replaced according to actual deployment needs. However, the validation in this study is still mainly based on scenarios where existing security cameras can cover the main activity areas and all entry–exit passages. The applicability of the method under more complex camera arrangements, lighting variations, and automatic region configuration requires further investigation. Full article
(This article belongs to the Section Building Energy, Physics, Environment, and Systems)
Show Figures

Figure 1

27 pages, 4914 KB  
Article
A Viewpoint on Event-Driven Perception and Digital Twin Integration for Autonomous Mining Robotics
by Vasiliki Balaska and Antonios Gasteratos
Electronics 2026, 15(10), 1993; https://doi.org/10.3390/electronics15101993 - 8 May 2026
Viewed by 590
Abstract
Robotic systems are increasingly being deployed in mining operations to support tasks such as inspection, navigation, environmental monitoring, and safety supervision. However, mining environments present significant challenges for robotic perception due to dynamic terrain conditions, poor illumination, airborne dust, and frequent disturbances caused [...] Read more.
Robotic systems are increasingly being deployed in mining operations to support tasks such as inspection, navigation, environmental monitoring, and safety supervision. However, mining environments present significant challenges for robotic perception due to dynamic terrain conditions, poor illumination, airborne dust, and frequent disturbances caused by excavation and heavy machinery. Conventional frame-based vision systems often struggle under these conditions due to motion blur, latency, and limited dynamic range. This study proposes a system-level conceptual framework for integrating event-based sensing into robotic mining systems in order to support perception in highly dynamic and safety-critical environments, with the aim of improving responsiveness and robustness under such conditions. Event-based cameras, inspired by biological vision, asynchronously detect brightness changes at the pixel level and provide microsecond temporal resolution with high dynamic range and low latency. The proposed framework combines event cameras with complementary sensing modalities including LiDAR, inertial measurement units, and RGB cameras to form a multi-sensor perception architecture. The framework is structured into multiple functional layers encompassing environmental sensing, event-driven perception, sensor fusion and AI processing, digital twin integration, and autonomous decision-making. Potential application scenarios including robotic tunnel inspection, autonomous navigation of mining robots, hazard detection, multi-agent cooperation in mining sites, and real-time digital twin updating are also discussed. The proposed framework provides a unified system-level reference architecture intended to guide future implementation and validation. Full article
Show Figures

Figure 1

16 pages, 3818 KB  
Article
Independent Motion Segmentation Based on Pure Event Data
by Wenjun Yin, Dongdong Teng and Lilin Liu
Sensors 2026, 26(9), 2620; https://doi.org/10.3390/s26092620 - 23 Apr 2026
Viewed by 880
Abstract
Event cameras are bio-inspired vision sensors offering low latency, low power consumption, and high dynamic range, capturing motion with microsecond-level precision via a per-event triggering mechanism. Despite these advantages, the inherent sparsity and lack of color in event data hinder direct analysis, necessitating [...] Read more.
Event cameras are bio-inspired vision sensors offering low latency, low power consumption, and high dynamic range, capturing motion with microsecond-level precision via a per-event triggering mechanism. Despite these advantages, the inherent sparsity and lack of color in event data hinder direct analysis, necessitating advanced deep learning approaches. To achieve low-latency and high-precision motion segmentation for indoor robotic applications, this paper introduces a dual-branch decoupled CNN framework. Specifically, Principal Component Analysis (PCA) is utilized to project 3D event point clouds into 2D motion trend maps, capturing local motion priors while suppressing ambiguity in structured environments. Concurrently, an Event Leaky Integration (ELI) model, inspired by biological membrane potentials, is designed to enhance the structural representation of sparse events. Within this framework, separate branches respectively perform motion validation and shape extraction and are fused via a Spatial Gated Fusion (SGF) module to suppress static background interference. It is demonstrated experimentally that with an input window of only 10 ms, the proposed method achieves a 77% average mIoU across five indoor test scenarios from the EV-IMO dataset with an inference latency of 10 ms per frame. Compared to state-of-the-art methods like MSRNN and GCN, which required 30–300 ms event slices, our framework achieves a favorable trade-off between computational efficiency and segmentation accuracy, maintaining competitive performance under ultra-short time windows for indoor event-based motion processing. Full article
(This article belongs to the Special Issue Event-Based Vision Technology: From Imaging to Perception and Control)
Show Figures

Figure 1

38 pages, 3132 KB  
Article
Lightweight Semantic-Aware Route Planning on Edge Hardware for Indoor Mobile Robots: Monocular Camera–2D LiDAR Fusion with Penalty-Weighted Nav2 Route Server Replanning
by Bogdan Felician Abaza, Andrei-Alexandru Staicu and Cristian Vasile Doicin
Sensors 2026, 26(7), 2232; https://doi.org/10.3390/s26072232 - 4 Apr 2026
Viewed by 2308
Abstract
The paper introduces a computationally efficient semantic-aware route planning framework for indoor mobile robots, designed for real-time execution on resource-constrained edge hardware (Raspberry Pi 5, CPU-only). The proposed architecture fuses monocular object detection with 2D LiDAR-based range estimation and integrates the resulting semantic [...] Read more.
The paper introduces a computationally efficient semantic-aware route planning framework for indoor mobile robots, designed for real-time execution on resource-constrained edge hardware (Raspberry Pi 5, CPU-only). The proposed architecture fuses monocular object detection with 2D LiDAR-based range estimation and integrates the resulting semantic annotations into the Nav2 Route Server for penalty-weighted route selection. Object localization in the map frame is achieved through the Angular Sector Fusion (ASF) pipeline, a deterministic geometric method requiring no parameter tuning. The ASF projects YOLO bounding boxes onto LiDAR angular sectors and estimates the object range using a 25th-percentile distance statistic, providing robustness to sparse returns and partial occlusions. All intrinsic and extrinsic sensor parameters are resolved at runtime via ROS 2 topic introspection and the URDF transform tree, enabling platform-agnostic deployment. Detected entities are classified according to mobility semantics (dynamic, static, and minor) and persistently encoded in a GeoJSON-based semantic map, with these annotations subsequently propagated to navigation graph edges as additive penalties and velocity constraints. Route computation is performed by the Nav2 Route Server through the minimization of a composite cost functional combining geometric path length with semantic penalties. A reactive replanning module monitors semantic cost updates during execution and triggers route invalidation and re-computation when threshold violations occur. Experimental evaluation over 115 navigation segments (legs) on three heterogeneous robotic platforms (two single-board RPi5 configurations and one dual-board setup with inference offloading) yielded an overall success rate of 97% (baseline: 100%, adaptive: 94%), with 42 replanning events observed in 57% of adaptive trials. Navigation time distributions exhibited statistically significant departures from normality (Shapiro–Wilk, p < 0.005). While central tendency differences between the baseline and adaptive modes were not significant (Mann–Whitney U, p = 0.157), the adaptive planner reduced temporal variance substantially (σ = 11.0 s vs. 31.1 s; Levene’s test W = 3.14, p = 0.082), primarily by mitigating AMCL recovery-induced outliers. On-device YOLO26n inference, executed via the NCNN backend, achieved 5.5 ± 0.7 FPS (167 ± 21 ms latency), and distributed inference reduced the average system CPU load from 85% to 48%. The study further reports deployment-level observations relevant to the Nav2 ecosystem, including GeoJSON metadata persistence constraints, graph discontinuity (“path-gap”) artifacts, and practical Route Server configuration patterns for semantic cost integration. Full article
(This article belongs to the Special Issue Advances in Sensing, Control and Path Planning for Robotic Systems)
Show Figures

Figure 1

Back to TopTop