Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (904)

Search Parameters:
Keywords = camera sensor network

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 107365 KB  
Article
M3-RGB: An Imaging Sensor System Using Multicore, Multimode Optical Fiber and Neural Networks
by Seigo Ito, Isamu Takai, Akari Kawasaki, Tadashi Ichikawa, Shin Motooka and Minoru Tanaka
Sensors 2026, 26(17), 5582; https://doi.org/10.3390/s26175582 - 2 Sep 2026
Viewed by 262
Abstract
Conventional image acquisition requires an electrically powered image sensor to be placed directly behind the camera lens, constraining camera placement. To overcome this issue, we introduce M3-RGB as an incoherent-light fiber imaging system in which a multicore, multimode optical fiber passively relays lens [...] Read more.
Conventional image acquisition requires an electrically powered image sensor to be placed directly behind the camera lens, constraining camera placement. To overcome this issue, we introduce M3-RGB as an incoherent-light fiber imaging system in which a multicore, multimode optical fiber passively relays lens images to a remotely located image sensor. Unlike conventional approaches, M3-RGB is designed to operate directly on incoherent light and requires no electrical power or active components at the sensing interface. Because propagation through the fiber yields spatially scrambled patterns, a neural network is used to reconstruct the original scene by exploiting the spatial locality preserved by the multicore structure. In a controlled optical bench setup, where a liquid crystal display monitor displays road-scene images, we construct a paired dataset of scrambled and ground-truth images and quantitatively evaluate reconstruction performance across different fiber core counts, fiber lengths, and calibration settings, utilizing the peak signal-to-noise ratio and structural similarity index measure as performance metrics. By decoupling imaging electronics from the sensing point, this passive remote image relay approach may expand sensor placement options for potential applications such as all-around perception for mobile robots and autonomous vehicles, surveillance, and inspection in confined spaces. Evaluations in real outdoor environments constitute future work. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

22 pages, 17875 KB  
Article
Sparse Image Registration-Based Marker-Free Hand Acupoint Localization Using TCM Template Queries
by Shujian Zhang, Shuyue Zhang, Chi Zhang and Jianqing Peng
Sensors 2026, 26(17), 5552; https://doi.org/10.3390/s26175552 - 1 Sep 2026
Viewed by 231
Abstract
Deep learning for sensor-based medical imaging provides a non-contact and data-driven route for anatomical surface analysis and personalized traditional Chinese medicine (TCM) applications. However, accurate hand-acupoint localization from camera-acquired hand images remains challenging because expert-annotated acupoint datasets are limited and inter-subject anatomical variations [...] Read more.
Deep learning for sensor-based medical imaging provides a non-contact and data-driven route for anatomical surface analysis and personalized traditional Chinese medicine (TCM) applications. However, accurate hand-acupoint localization from camera-acquired hand images remains challenging because expert-annotated acupoint datasets are limited and inter-subject anatomical variations are significant. This paper proposes a marker-free hand-acupoint localization method based on deep feature correspondence learning and sparse image registration. The task is formulated as template-to-target correspondence estimation, in which expert-annotated acupoints in a TCM template image are used as query points and mapped to a target hand image acquired by an optical imaging sensor. A Transformer-based architecture is employed to correlate multi-scale image features, and an uncertainty-aware matching formulation is used to estimate both acupoint positions and unreliable matches. Unlike conventional keypoint detection networks, the proposed method exploits TCM template priors and reduces the dependence on dense target-image acupoint annotations. The constructed dataset contains 1400 images from 378 participants. Participant-level partitioning was performed before image-pair generation, yielding a held-out test set of 38 participants (140 images). Palm and dorsal-hand views were evaluated separately against keypoint-detection baselines and COTR. On this participant-independent internal test set, the proposed method achieved AAPE values of 18.42 pixels for palm images and 12.60 pixels for dorsal-hand images, while reducing inference time from 11,000 ms for COTR to 600 ms. These results indicate the feasibility of template-guided image-based hand-acupoint localization with reference to expert annotations. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

32 pages, 74908 KB  
Article
CIAFNet: An RGB-D Cross-Modal Interaction and Adaptive Fusion Network for Camellia oleifera Fruit Detection
by Yan Chen, Chengxin Yang, Chao Yuan, Yiming Lu, Dandan Fu, Yinghui Fang, Shuman Liu and Hui Ai
Agriculture 2026, 16(17), 1884; https://doi.org/10.3390/agriculture16171884 - 30 Aug 2026
Viewed by 262
Abstract
To better address the accuracy bottleneck of RGB-only Camellia oleifera C.Abel fruit detection in complex orchard environments, this paper proposes CIAFNet—a dual-stream RGB-D fusion detection network—and evaluates its potential as a visual front end for relative 3D localization using sensor-measured depth and camera [...] Read more.
To better address the accuracy bottleneck of RGB-only Camellia oleifera C.Abel fruit detection in complex orchard environments, this paper proposes CIAFNet—a dual-stream RGB-D fusion detection network—and evaluates its potential as a visual front end for relative 3D localization using sensor-measured depth and camera back projection under controlled conditions. With RGB images and depth maps as parallel dual-branch inputs, the network integrates the C3k2_PartialNetBlock for efficient intra-modal feature extraction with reduced computational redundancy, devises the cross-modal interaction and difference-aware adaptive fusion (CIDAF) module for adaptive cross-modal feature fusion, and adopts an SC-EUCB-augmented BiFPN in the neck to optimize multiscale feature aggregation and detail restoration during upsampling. Pseudo-depth maps generated from natural orchard RGB images via Depth Anything V2 were paired with RGB counterparts to build an RGB–pseudo-depth dataset. Synchronized RGB-D data collected by an Intel RealSense D435i under controlled conditions were used to quantify pseudo-to-sensor depth discrepancies and evaluate input adaptability. On the natural orchard test set, CIAFNet achieved 93.33% mAP@0.5 with only 10.49 GFLOPs and 3.80 M parameters. Second-stage fine-tuning improved CIAFNet’s adaptation to D435i-measured depth under controlled conditions. In the subsequent relative displacement consistency experiment, the mean absolute consistency errors along the X, Y, and Z axes were 3.10, 3.15, and 3.27 mm, respectively, and the mean 3D Euclidean consistency error was 5.60 mm. These results demonstrate that CIAFNet improves Camellia oleifera fruit detection using natural orchard RGB–pseudo-depth data and has potential as a visual front end for relative 3D localization under controlled conditions. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

35 pages, 10906 KB  
Article
An AR 3D Tracking and Registration Method That Integrates Optical Flow Tracking and Mean Shift
by Jiu Yong, Xiaomei Lei and Jianwu Dang
Sensors 2026, 26(17), 5509; https://doi.org/10.3390/s26175509 - 30 Aug 2026
Viewed by 308
Abstract
Augmented reality (AR) enhances the real world scene by overlaying virtual information onto it. Vision-based 3D tracking and registration is the key technology for ensuring the fusion of virtual and real content in monocular AR systems. Existing mainstream visual tracking and registration methods [...] Read more.
Augmented reality (AR) enhances the real world scene by overlaying virtual information onto it. Vision-based 3D tracking and registration is the key technology for ensuring the fusion of virtual and real content in monocular AR systems. Existing mainstream visual tracking and registration methods are susceptible to illumination variations, motion blur, target occlusion, and dynamic background interference in complex scenarios. They also suffer from low computational efficiency, cumulative pose errors, and insufficient stability, making them difficult to deploy on low power edge devices such as embedded systems and mobile terminals. To address these issues, this paper proposes a lightweight monocular AR 3D tracking and registration method that integrates ORB-FREAK features, mismatching outlier filtering, background weighted mean shift, and template-based relocalization. The method does not rely on depth sensors or neural network inference, enabling efficient and accurate lightweight pose estimation. Specifically, we first combine the ORB (Oriented FAST and Rotated BRIEF) descriptor with the FREAK (Fast Retina Keypoint) algorithm for feature detection and initial matching. Hamming distance is used for coarse filtering of mismatched point pairs, and an ascending sort combined with an iterative sequential sampling strategy is applied to solve the optimal homography matrix, significantly improving the accuracy and efficiency of matrix estimation. Then, distance constraints among feature points are imposed on the target registration region to optimize the selection, and camera pose is computed based on the matching between 2D feature points and their corresponding 3D spatial coordinates, eliminating the error accumulation problem of conventional algorithms. Real-time feature matching is further used to correct the optical flow tracking sequence and camera pose, ensuring the continuity of the AR tracking process. Finally, a background weighted mean shift algorithm is introduced to narrow the feature detection range and suppress background interference, complemented by a template-matching relocalization module and a dynamic model update strategy, which effectively enhance the robustness of continuous tracking and registration under complex conditions. Experimental results demonstrate that, in extreme scenarios such as low light conditions, high speed motion, and occlusion, the proposed method achieves AR 3D tracking and registration success rates of 86.7%, 82.3%, and 78.5%, respectively. It exhibits superior performance in pose estimation accuracy and anti-interference capability in complex environments, with significantly reduced computational overhead. Moreover, it can achieve robust and continuous AR 3D tracking and registration on low power edge devices, effectively adapting to demanding AR application scenarios and providing reliable technical support for lightweight AR applications. Full article
(This article belongs to the Topic Extended Reality: Models and Applications)
Show Figures

Figure 1

22 pages, 9335 KB  
Article
A Single-Pass Approach That Mines Unstructured Robotic Trajectories for Calibrating Surveillance Cameras
by Wayne Lam, Yingqi Liu, Chong Di, Hao Tang, Jie Gong, Fred Roberts and Zhigang Zhu
Sensors 2026, 26(17), 5473; https://doi.org/10.3390/s26175473 - 29 Aug 2026
Viewed by 193
Abstract
The ubiquity of surveillance networks in public infrastructure presents a significant, yet underutilized, opportunity to assist vulnerable populations, particularly individuals who are blind or have low vision (BLV). However, transforming these passive video feeds into active guidance systems requires accurate camera calibration, a [...] Read more.
The ubiquity of surveillance networks in public infrastructure presents a significant, yet underutilized, opportunity to assist vulnerable populations, particularly individuals who are blind or have low vision (BLV). However, transforming these passive video feeds into active guidance systems requires accurate camera calibration, a process that is traditionally labor-intensive and unscalable in large facilities. This paper introduces a novel, automated framework that leverages a mobile quadruped robot (Boston Dynamics Spot) as a dynamic calibration agent. We propose a single-pass approach with two planar calibration algorithms, which mines unstructured robotic trajectories to construct distinct geometric features on the ground plane. By synthesizing “virtual rectangles” from the robot’s odometry, our method first recovers camera focal length through vanishing point estimation by constructing virtual rectangles, and then solves for extrinsic 6-DoF pose using two algorithms: our proposed Virtual Rectangle (ViR) algorithm and a standard planar Perspective-n-Point (PnP) algorithm. Experimental validation using real-world data demonstrates the system’s robustness to sensor noise, maintaining focal length errors around 5%, rotational errors around 5° and relative translation errors of approximately 10%, while synthetic simulations indicate focal length errors generally under 10%, rotational errors under 1.5°, and translation errors also around 10% despite heavy image and object disturbances. This work eliminates the need for manual calibration targets for dynamic camera calibration, effectively converting static security infrastructure into a metric sensing network capable of supporting high-fidelity social robotics applications. Full article
Show Figures

Figure 1

28 pages, 8603 KB  
Article
Event-Guided Image Reconstruction for Nighttime Dynamic Scenes
by Qingjiao Meng, Ji Li and Yan Jin
J. Imaging 2026, 12(9), 399; https://doi.org/10.3390/jimaging12090399 - 23 Aug 2026
Viewed by 195
Abstract
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, [...] Read more.
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, we propose an event-guided image reconstruction method for nighttime dynamic visual perception. The method constructs a multi-channel event voxel representation by jointly encoding event count, event intensity, timestamp distribution, and blurred-frame intensity priors. A parameter-efficient local–global reconstruction network is then designed to restore fine-grained textures and model holistic structures. In addition, edge-alignment and blur-alignment constraints are introduced to improve geometric consistency and imaging plausibility. Experiments on the HQF and REDS datasets show that the proposed method outperforms existing methods in the MSE, PSNR, and SSIM. Compared with DeblurSR, it reduces the MSE by 14.81% on HQF and 10.00% on REDS, while improving the PSNR by 1.603 dB and 1.053 dB, respectively. Qualitative results further show sharper edges, lower structural errors, and better edge consistency. Low illumination, dynamic blur, rapid brightness variation, and event noise are also common degradation factors in nighttime UAV imaging, making the investigated problem technically relevant to that setting. However, because neither REDS nor HQF was acquired during an actual UAV flight, the reported results establish benchmark-level reconstruction performance rather than UAV-specific operational effectiveness. Full article
Show Figures

Figure 1

36 pages, 28430 KB  
Article
Robot-Centric Elevation Map Completion with Sensor Geometry-Aware Augmentation and Uncertainty Estimation
by Jozef Goga, Michal Kovac, Martin Dekan, Jarmila Pavlovicova and Frantisek Duchon
Appl. Sci. 2026, 16(16), 8262; https://doi.org/10.3390/app16168262 - 19 Aug 2026
Viewed by 276
Abstract
Robot-centric elevation maps built from onboard sensing are always incomplete: occlusions, a limited field of view, and range limits leave large unobserved regions that traversability analysis and motion planning must still reason about. We present a supervised framework that completes these maps and [...] Read more.
Robot-centric elevation maps built from onboard sensing are always incomplete: occlusions, a limited field of view, and range limits leave large unobserved regions that traversability analysis and motion planning must still reason about. We present a supervised framework that completes these maps and reports a per-cell uncertainty. Its core is a ray-cone augmentation that removes angular sectors anchored at the sensor origin during training; unlike the random masks of image inpainting, these sectors match the coverage gaps of real deployments, such as camera failures or reduced camera configurations. Partial maps generated from four depth cameras along legged-robot trajectories in the TartanGround dataset are paired with dense ground truth, yielding 32,329 samples across five outdoor environments. An encoder–decoder network is trained with a masked β-NLL loss and evaluated with a five-fold leave-one-environment-out protocol. The augmentation lowers the completion error on missing sensor sectors by 8.3 to 9.7%, depending on the sector width, at no measurable cost on uncorrupted partial inputs. The completed maps reach a hole root-mean-square error of 2.86 m, a 45% improvement over the strongest classical interpolation baseline. Full article
(This article belongs to the Special Issue Application of Computer Science in Mobile Robots, 3rd Edition)
Show Figures

Figure 1

28 pages, 102280 KB  
Article
PRMEFNet: A Real-Time Unsupervised Multi-Exposure Fusion Network Driven by Prior Knowledge
by Junwei Qi, Hangdong Wang, Xu Xiao and Jingpeng Gao
J. Imaging 2026, 12(8), 389; https://doi.org/10.3390/jimaging12080389 - 19 Aug 2026
Viewed by 167
Abstract
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely [...] Read more.
Due to the limited dynamic range of imaging sensors, most cameras can only capture low-dynamic-range (LDR) images. Multi-exposure fusion (MEF) is an effective technique for generating high-dynamic-range (HDR) images. However, to simultaneously preserve texture details and global exposure, most existing methods primarily rely on more complex models to improve performance, resulting in higher computational costs and longer processing times. To address this issue, we propose a real-time unsupervised MEF network driven by prior knowledge. To this end, a hierarchical feature extraction module is designed that utilizes filtering operations to decompose the source images into base layers and detail layers. Features are extracted from each layer separately to reduce the difficulty of extracting effective features. Then, the receptive field of feature maps is expanded by dilated convolutions, and a window-based self-attention mechanism is applied to perform context modeling, achieving effective contextual modeling with low computational cost. Subsequently, texture features and global features are extracted separately to enable the model to maintain both local texture clarity and global smoothness. In addition, a one-dimensional lookup table is utilized to accelerate the inference process. Comprehensive experiments are conducted to verify the effectiveness of the proposed method. The subjective evaluation results demonstrate that the fused images exhibit superior visual quality, while objective experiments further quantify its superior performance, demonstrating that the proposed method effectively reduces computation time. Full article
(This article belongs to the Topic Computational Imaging)
Show Figures

Figure 1

18 pages, 7408 KB  
Article
Effectiveness of Spectral Analysis for Evaluating Internal Quality of Korla Fragrant Pears Under Different Detection Distances
by Yifei Li, Xueting Ma, Jianping Bao, Yuesen Tong, Lei Kang, Huaiyu Liu, Zhe Han, Jun Guo, Xuhang Liu and Kaijie Qi
Horticulturae 2026, 12(8), 1026; https://doi.org/10.3390/horticulturae12081026 - 17 Aug 2026
Viewed by 391
Abstract
This study investigated how detection distance affects spectral models for soluble solids content (SSC) and firmness evaluation in Korla fragrant pears and provides a reference for calibrating standardized indoor non-destructive detection equipment. Two hundred visually intact fruit samples at the early-ripening stage were [...] Read more.
This study investigated how detection distance affects spectral models for soluble solids content (SSC) and firmness evaluation in Korla fragrant pears and provides a reference for calibrating standardized indoor non-destructive detection equipment. Two hundred visually intact fruit samples at the early-ripening stage were collected from the Korla production area in Xinjiang. An FS-640 multispectral camera system equipped with a VS-SWR fixed-focus industrial lens (16 mm focal length, F1.8 maximum aperture, 1/2-inch sensor format) was used to acquire fruit reflectance spectra at seven vertical lens-to-fruit-surface distances of 90, 100, 110, 120, 130, 140, and 150 cm. A 625-pixel region of interest (ROI) was selected using ENVI at an undamaged equatorial or near-equatorial position of each fruit, and the regional mean spectrum was used as the spectral feature of one fruit sample. The sample-set partitioning based on joint X–Y distances (SPXY) algorithm was used to divide the calibration and prediction sets at a 3:1 ratio after outlier removal via a residual-threshold method. Four preprocessing methods, namely LOESS smoothing, standardization, vector normalization, and Savitzky–Golay (SG) smoothing, were compared. Competitive adaptive reweighted sampling (CARS) was performed with 50 Monte-Carlo sampling runs, a maximum of 30 principal components, and 10-fold cross-validation, yielding 99 characteristic wavelengths. Partial least squares regression (PLSR), support vector regression (SVR), random forest (RF), and artificial neural network (ANN) models were then established using identical input variables and sample partitions. Model performance was evaluated using the coefficient of determination for calibration (Rc2), coefficient of determination for prediction (RP2), root-mean-square error of calibration (RMSEC), root-mean-square error of prediction (RMSEP), relative prediction deviation (RPD), and ratio of performance to interquartile distance (RPIQ). Under the static laboratory acquisition conditions in this work, the SSC model achieved the best prediction performance at 110 cm with SG smoothing (RP2) = 0.8949, RPD = 3.0633, RPIQ = 5.8661), whereas the firmness model obtained optimal prediction performance at 140 cm with standardization (RP2) = 0.7460, RPD = 1.9425, RPIQ = 3.2867). Changes in detection distance altered illumination uniformity, effective reflected signal, photon-scattering paths, and background-noise proportion. These effects may partially explain why the chemical-absorption-dominated SSC index and the tissue-scattering-dominated firmness index responded differently to detection distance. The results provide a reference for setting spectral detection parameters for Korla fragrant pears; however, samples were obtained from only a single producing region, harvest season, and maturity stage, and no independent external validation dataset was used. Therefore, the generalization ability of the developed models needs to be further verified using cross-season and cross-orchard sample sets. Full article
Show Figures

Figure 1

27 pages, 25544 KB  
Article
AOPQ-Net Acoustic–Optical Proposal Query Network for Underwater Multimodal Object Detection
by Yanze Lu, Zhengyan Zhang, Shuoshuo Ding, Haochen Hu, Chih-Yung Wen and Tiedong Zhang
Remote Sens. 2026, 18(16), 2703; https://doi.org/10.3390/rs18162703 - 11 Aug 2026
Viewed by 464
Abstract
Optical cameras and imaging sonars are widely used sensors in autonomous underwater vehicles. However, their different imaging mechanisms introduce substantial cross-modal discrepancies in the acquired data. In addition, underwater optical images are often degraded by low illumination, scattering, and turbidity, whereas sonar images [...] Read more.
Optical cameras and imaging sonars are widely used sensors in autonomous underwater vehicles. However, their different imaging mechanisms introduce substantial cross-modal discrepancies in the acquired data. In addition, underwater optical images are often degraded by low illumination, scattering, and turbidity, whereas sonar images commonly suffer from speckle noise and low spatial resolution. As a result, object detection based on a single optical or acoustic modality is often insufficient in challenging underwater environments. To address this problem, this paper proposes an acoustic–optical fusion network for underwater object detection, termed an Acoustic–Optical Proposal Query Network (AOPQ-Net). First, a Sonar Position Encoding (SPE) module is designed to explicitly encode the geometric priors in sonar images. Second, a Bi-directional Discrepancy-aware Spatial Alignment (BDSA) module is introduced to alleviate spatial misalignment between the two modalities at the feature level. Third, a Proposal Query Transformer (PQT) module performs target-oriented cross-modal interaction at the proposal level. Furthermore, this study constructs a dedicated dataset for underwater acoustic–optical fusion object detection, named Haiqin Underwater Fusion (HUF), and conducts systematic experiments on this dataset. The experimental results show that AOPQ-Net outperforms single-modality baselines and representative multimodal fusion methods in both optical and acoustic image spaces, which demonstrate the effectiveness of the proposed method. Full article
Show Figures

Figure 1

24 pages, 3613 KB  
Article
RG-PSR: Reliability-Guided Poisson Surface Reconstruction for Degraded 3D-Imaging Point Clouds
by Na Liu, Fan Zhang, Jiawei Wang, Dan Zhang, Jinliang Wu and Xiaohui Li
J. Imaging 2026, 12(8), 369; https://doi.org/10.3390/jimaging12080369 - 10 Aug 2026
Viewed by 302
Abstract
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit [...] Read more.
Three-dimensional (3D) imaging systems, including depth cameras, LiDAR sensors, and multi-view scanning pipelines, often produce point clouds with noisy normals, outliers, sparse sampling, and non-uniform density, which can degrade downstream mesh reconstruction. Poisson surface reconstruction is lightweight and training-free, but its global implicit formulation is sensitive to unreliably oriented samples and fixed density-trimming thresholds. This paper presents RG-PSR, a reliability-guided enhancement framework for Poisson-family surface reconstruction from degraded 3D-imaging point clouds. RG-PSR estimates a deterministic per-point reliability score from local density regularity, spacing variation, and normal consistency, and propagates this score through conservative point filtering, reliability-guided normal refinement, adaptive density-reliability trimming, and structure-aware postprocessing. The main pipeline requires no manual labels, neural network training, or ground-truth meshes at inference time. Experiments on three groups of object meshes under five deterministic degradation types show that RG-PSR improves Poisson-family reconstruction under degraded inputs. Compared with fixed density-trimmed Poisson reconstruction, RG-PSR reduces the overall Chamfer-L1 from 0.0218 to 0.0172, improves F0.01 from 0.6618 to 0.6836, and reduces Artifact0.02 from 0.3090 to 0.2632. In the broader classical comparison, local triangulation methods achieve stronger point-wise accuracy, while RG-PSR yields the fewest connected components and the highest largest-component ratio. These results position RG-PSR as a practical reliability layer for coherent Poisson-family reconstruction rather than a universal replacement for all surface-reconstruction methods. Full article
(This article belongs to the Special Issue Advances in 3D Point Cloud Processing)
Show Figures

Figure 1

29 pages, 597 KB  
Article
Quantized vs. Full-Precision YOLO Models on Edge Devices: A Performance Benchmark for Real-Time License Plate Detection in Smart Parking Systems
by Ervin Burkus, Bence Lestyán, Lehel Dénes-Fazakas and György Eigner
Sensors 2026, 26(16), 5034; https://doi.org/10.3390/s26165034 - 8 Aug 2026
Viewed by 353
Abstract
The deployment of deep learning-based vision systems on edge devices introduces a complex trade-off between computational efficiency and detection accuracy. In this work, we investigate this trade-off in the context of a multi-stage Automatic License Plate Recognition (ALPR) pipeline, evaluated in two heterogeneous [...] Read more.
The deployment of deep learning-based vision systems on edge devices introduces a complex trade-off between computational efficiency and detection accuracy. In this work, we investigate this trade-off in the context of a multi-stage Automatic License Plate Recognition (ALPR) pipeline, evaluated in two heterogeneous edge execution environments: a general-purpose Raspberry Pi 5 single-board computer and the ARTPEC-8 system-on-chip integrated into an Axis smart camera, where neural network inference is accelerated by the on-chip DLPU. All experiments were performed using pre-recorded images loaded from the file system; neither the Axis camera sensor nor a live video stream was used. This study evaluates the impact of model architecture, numerical precision, and input resolution on both inference latency and detection performance. YOLOv5- and YOLOv8-based models were analyzed under multiple quantization schemes (FP32, FP16, dynamic, and INT8), while a cross-platform benchmark was conducted to assess the benefits and limitations of hardware acceleration. The results show that dedicated accelerators provide significant latency reduction at higher resolutions; however, this advantage is accompanied by reduced flexibility and increased sensitivity to quantization effects. In contrast, CPU-based execution enables the use of more recent and quantization-robust model architectures, which can partially compensate for the lack of hardware acceleration when combined with resolution scaling. Furthermore, the analysis hig ights the importance of hybrid-resolution processing in multi-stage pipelines, where different stages can operate at different input resolutions to balance accuracy and performance. The findings demonstrate that optimal system design requires a joint consideration of hardware characteristics, model architecture, and quantization strategy, rather than relying on a single optimization dimension. The presented results provide practical insights for the design of efficient and robust edge-based ALPR systems, with direct implications for real-world industrial deployments. Full article
Show Figures

Graphical abstract

27 pages, 26649 KB  
Article
Evaluating Deep Learning Local Features for RGB-Thermal Image Matching and 3D InfraRed Thermography
by Luca Morelli, Neil Sutherland, Francesco Ioli, Alfonso Vitti, Stuart Marsh, Jon Mills, Paul Bryan and Fabio Remondino
Geomatics 2026, 6(4), 83; https://doi.org/10.3390/geomatics6040083 - 29 Jul 2026
Viewed by 492
Abstract
InfraRed Thermography (IRT), a non-invasive, non-contact, and non-destructive testing (NDT) technique, has become an established tool in the assessment of a building’s behavior and energy performance. However, the inherent low spatial resolution of thermal infrared (TIR) cameras has led recent work to fuse [...] Read more.
InfraRed Thermography (IRT), a non-invasive, non-contact, and non-destructive testing (NDT) technique, has become an established tool in the assessment of a building’s behavior and energy performance. However, the inherent low spatial resolution of thermal infrared (TIR) cameras has led recent work to fuse thermographic and geometric data to generate accurate 3D representations of buildings encapsulating temperature information. Whilst existing data fusion methods have relied on sensors in fixed relative orientation (RO), the co-registration of independent TIR and RGB blocks using ground control points (GCPs), or the reprojection of TIR images onto additional geometric or parametric models, approaches that directly match multi-modal images remain limited. In principle, if multi-modal tie points were available, it would be possible to directly align the RGB block with the TIR block; however, such matching is extremely challenging due to the substantial differences in radiometric properties. The main contribution of this paper is to demonstrate the applicability of off-the-shelf deep learning-based image matching algorithms, originally trained on mono-modal datasets, to multi-modal matching tasks for InfraRed Thermography 3D-Data Fusion (IRT-3DDF). We conduct a comparative evaluation of the principal algorithms developed in recent years, with particular emphasis on 3D accuracy and computational efficiency, under the hypothesis that, owing to the inherently local nature of the problem they address, these algorithms can generalize from a mono-modal training domain to a multi-modal application domain. The results are benchmarked against existing hand-crafted open-source multi-modal reference methods. Importantly, the proposed method is fully-automatic, obviating the need for sensor pre-calibration, manual co-registration, or associated positioning information. Results demonstrate that DL-based image matching, using pre-trained neural networks outside of their expected training domain, provides a viable approach for IRT-3DDF capable of co-registering blocks of multi-modal images across varying scales, settings, sensors, and subjects. Our results indicate accuracy in 3D is up to seven times better than multi-modal hand-crafted algorithms, while hand-crafted mono-modal methods fail to co-register images in their entirety. Full article
Show Figures

Figure 1

32 pages, 19861 KB  
Article
A Geographic Consistency-Constrained Cross-Modal Super-Resolution Matching Method for UAV Geo-Localization
by Jindi Wang, Haigang Sui, Chang Liu, Zhina Song and Lieyun Hu
Remote Sens. 2026, 18(15), 2475; https://doi.org/10.3390/rs18152475 - 28 Jul 2026
Viewed by 456
Abstract
Visual geo-localization is a predominant approach for unmanned aerial vehicles (UAVs) operating in Global Navigation Satellite System (GNSS)-denied environments, typically achieved by matching UAV-captured visible optical images with satellite base maps. However, under low-light conditions, visible cameras struggle to capture distinct features. While [...] Read more.
Visual geo-localization is a predominant approach for unmanned aerial vehicles (UAVs) operating in Global Navigation Satellite System (GNSS)-denied environments, typically achieved by matching UAV-captured visible optical images with satellite base maps. However, under low-light conditions, visible cameras struggle to capture distinct features. While infrared sensors can capture clear features in such scenarios, the significant modality gap between thermal infrared images and optical satellite base maps makes accurate matching highly challenging. In this paper, we propose a novel cross-modal super-resolution matching and geo-localization method constrained by geographic consistency. First, a geographic consistency normalization module is introduced to narrow the modality gap between satellite optical images and thermal infrared images, thereby enhancing cross-modal matchability. Subsequently, a thermal infrared super-resolution enhancement module is employed to improve the spatial resolution and detail representation of the images, effectively increasing feature discriminability in low-texture regions. Finally, an end-to-end dense matching module is utilized to strengthen the stability of cross-modal correspondence estimation, ultimately improving geo-localization accuracy in low-light environments. Extensive experiments conducted on both a self-constructed network dataset and a real-world flight dataset demonstrate that the proposed method outperforms current competitive approaches. The proposed framework is not a simple combination of existing enhancement and matching modules, but a task-driven design that jointly addresses cross-modal discrepancy, low-resolution thermal imagery, and robust correspondence estimation. Experiments on self-constructed and public datasets demonstrate its robustness and superiority, achieving average geo-localization errors of 1.31 m and 8.04 m, respectively. Full article
Show Figures

Figure 1

14 pages, 13570 KB  
Article
A Portable Solar-Powered Edge-AI System for Livestock Monitoring in Off-Grid Mountain Pastures: System Design and Field Validation
by Tomo Popović, Dejan Drajić, Janko Kaljević, Ivan Jovović and Dejan Babić
Appl. Sci. 2026, 16(14), 7257; https://doi.org/10.3390/app16147257 - 20 Jul 2026
Viewed by 942
Abstract
Highland pastures in Montenegro, known as katuns, are seasonal settlements without grid power or network coverage and which are located where conventional monitoring is unfeasible. This study presents a solar-powered, off-grid system for livestock and environmental monitoring. It integrates, into a single portable [...] Read more.
Highland pastures in Montenegro, known as katuns, are seasonal settlements without grid power or network coverage and which are located where conventional monitoring is unfeasible. This study presents a solar-powered, off-grid system for livestock and environmental monitoring. It integrates, into a single portable unit, a solar power station, an edge-AI computer, a camera, environmental sensors, a LoRaWAN gateway, and a cellular router for backhaul. All parts are pre-wired in a modular enclosure, deployable by one operator in under 30 min. Data are fed to the agroNET farm-management platform and a purpose-built mobile web application; livestock detection runs on-device using a model from our earlier work. The system was evaluated at three sites, including a highland katun near Žabljak (~1450 m), under a two-phase energy-measurement protocol. During field logging it drew ~75 W on average against ~125 W solar input—a measured surplus that is used to recharge the battery—with a daily monitoring load of ~1560 Wh. The four-panel array’s nameplate potential in summer is an estimated ~3700 Wh/day, indicating substantial headroom relative to the measured load. At 80% depth of discharge the battery gives ~20 h autonomy, and the detection pipeline ran continuously, processing ~10,000 frames at under 3 s latency. The results demonstrate the feasibility of off-grid precision livestock farming, reaching TRL 6. Full article
(This article belongs to the Special Issue Automation and Smart Technologies in Agriculture)
Show Figures

Figure 1

Back to TopTop