Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (14)

Search Parameters:
Keywords = pseudo-depth net

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
32 pages, 74908 KB  
Article
CIAFNet: An RGB-D Cross-Modal Interaction and Adaptive Fusion Network for Camellia oleifera Fruit Detection
by Yan Chen, Chengxin Yang, Chao Yuan, Yiming Lu, Dandan Fu, Yinghui Fang, Shuman Liu and Hui Ai
Agriculture 2026, 16(17), 1884; https://doi.org/10.3390/agriculture16171884 - 30 Aug 2026
Viewed by 294
Abstract
To better address the accuracy bottleneck of RGB-only Camellia oleifera C.Abel fruit detection in complex orchard environments, this paper proposes CIAFNet—a dual-stream RGB-D fusion detection network—and evaluates its potential as a visual front end for relative 3D localization using sensor-measured depth and camera [...] Read more.
To better address the accuracy bottleneck of RGB-only Camellia oleifera C.Abel fruit detection in complex orchard environments, this paper proposes CIAFNet—a dual-stream RGB-D fusion detection network—and evaluates its potential as a visual front end for relative 3D localization using sensor-measured depth and camera back projection under controlled conditions. With RGB images and depth maps as parallel dual-branch inputs, the network integrates the C3k2_PartialNetBlock for efficient intra-modal feature extraction with reduced computational redundancy, devises the cross-modal interaction and difference-aware adaptive fusion (CIDAF) module for adaptive cross-modal feature fusion, and adopts an SC-EUCB-augmented BiFPN in the neck to optimize multiscale feature aggregation and detail restoration during upsampling. Pseudo-depth maps generated from natural orchard RGB images via Depth Anything V2 were paired with RGB counterparts to build an RGB–pseudo-depth dataset. Synchronized RGB-D data collected by an Intel RealSense D435i under controlled conditions were used to quantify pseudo-to-sensor depth discrepancies and evaluate input adaptability. On the natural orchard test set, CIAFNet achieved 93.33% mAP@0.5 with only 10.49 GFLOPs and 3.80 M parameters. Second-stage fine-tuning improved CIAFNet’s adaptation to D435i-measured depth under controlled conditions. In the subsequent relative displacement consistency experiment, the mean absolute consistency errors along the X, Y, and Z axes were 3.10, 3.15, and 3.27 mm, respectively, and the mean 3D Euclidean consistency error was 5.60 mm. These results demonstrate that CIAFNet improves Camellia oleifera fruit detection using natural orchard RGB–pseudo-depth data and has potential as a visual front end for relative 3D localization under controlled conditions. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

30 pages, 21057 KB  
Article
CIDR-MobileNet: A Monocular Pseudo-Depth and Cross-Modal Feature Fusion Approach for Chili Pepper Above-Ground Biomass Estimation
by Yi Wang, Jingtao Deng, Lin Yang, Shangjing Ruan, Weijie Wang, Wenwu Hu and Ping Jiang
Agriculture 2026, 16(13), 1457; https://doi.org/10.3390/agriculture16131457 - 2 Jul 2026
Viewed by 426
Abstract
Accurate real-time estimation of above-ground biomass is critical for intelligent chilli pepper harvesting. This study proposes CIDR-MobileNet, a lightweight end-to-end framework that addresses the limitations of destructive sampling, reliance on additional depth sensors, and weak regression robustness in existing methods. Pseudo-depth maps are [...] Read more.
Accurate real-time estimation of above-ground biomass is critical for intelligent chilli pepper harvesting. This study proposes CIDR-MobileNet, a lightweight end-to-end framework that addresses the limitations of destructive sampling, reliance on additional depth sensors, and weak regression robustness in existing methods. Pseudo-depth maps are generated from single-view RGB images using Depth Anything V2, providing low-cost structural information without requiring extra hardware. A cross-modal feature interaction module adaptively fuses RGB texture with pseudo-depth geometry, while a multi-branch distribution regression head models AGB prediction as a probabilistic task to improve robustness against occlusion and noise. A ranking loss is also introduced to preserve the relative order of predictions. Validated on 275 in-field chilli pepper samples via ten-fold cross-validation, the model achieves an R2 of 0.972, MAE of 174.56 g, RMSE of 230.74 g, and MAPE of 9.56%, with only 3.28 M parameters. Comparative experiments demonstrate that CIDR-MobileNet outperforms mainstream lightweight networks while maintaining high inference efficiency (10.56 ms CPU latency). The method strikes a favourable balance between prediction accuracy, hardware cost, and real-time performance, offering a practical solution for non-destructive biomass monitoring in precision agriculture. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

18 pages, 1420 KB  
Article
SAS-Net: An Agitated Behavior Early Warning Model for Community-Dwelling Dementia Patients Based on Symmetric Autoencoders and Spatio-Temporal Network
by Jing Xu, Bin Li, Ping Feng, Yonghan Zhang and Shengchun Yang
Symmetry 2026, 18(7), 1102; https://doi.org/10.3390/sym18071102 - 29 Jun 2026
Viewed by 356
Abstract
Using home sensors to provide agitation warnings for community-dwelling dementia patients without ongoing clinical supervision can enable their carers to intervene early during the agitation latent period, thereby reducing unnecessary hospital admissions and harmful events. Most existing studies are based on behavioral, sleep, [...] Read more.
Using home sensors to provide agitation warnings for community-dwelling dementia patients without ongoing clinical supervision can enable their carers to intervene early during the agitation latent period, thereby reducing unnecessary hospital admissions and harmful events. Most existing studies are based on behavioral, sleep, and physiological data from the previous 24 h to predict patients’ agitation events, which fail to fully capture patients’ recent behavioral details. In this study, monitoring data from the previous 8 × 24 h for dementia patients are used to achieve in-depth mining of patients’ living habits. Furthermore, to address the common clinical problem of extreme imbalance between agitation and normal samples (agitation samples often are extremely scarce), we designed a two-stage agitation early warning model based on symmetric autoencoders and a spatio-temporal network, dubbed SAS-Net. In the first stage, we randomly sample 80% of normal samples and employ multiple symmetric autoencoders to perform feature transformation and pseudo-label construction, and then pre-train the spatio-temporal learning network composed of convolutional networks and a gated Multilayer Perceptron. It aims to learn the patient data’s intrinsic structure, underlying patterns, and effective feature extraction methods. In the second stage, we freeze the parameters of the spatio-temporal learning network and use a balanced dataset consisting of the remaining 20% of normal samples and all agitation samples to reconstruct and fine-tune the top fully connected classifier to improve the recognition performance of agitation samples. The two-stage strategy resolves the problem of ineffective training faced by deep learning models on imbalanced datasets. The experimental results demonstrate the effectiveness of the proposed SAS-Net for agitated behavior early warning. Full article
Show Figures

Figure 1

29 pages, 8082 KB  
Article
CMYD-SurfaceNet: Scale-Aware Cascaded Multimodal MRI Segmentation via Representation-Level Structural Decoupling and Boundary-Constrained Learning
by Chaymae El Mechal, Mostefa Mesbah, Loubna Mazgouti, Fatima Zahra Ammor and Najiba El Amrani El Idrissi
Digital 2026, 6(2), 49; https://doi.org/10.3390/digital6020049 - 16 Jun 2026
Viewed by 568
Abstract
Reliable delineation of brain tumor boundaries in multimodal magnetic resonance imaging (MRI) remains challenging despite substantial advances in deep learning–based segmentation. Although modern encoder–decoder architectures achieve strong volumetric overlap, precise geometric alignment of tumor contours remains inconsistent, particularly for small lesions and heterogeneous [...] Read more.
Reliable delineation of brain tumor boundaries in multimodal magnetic resonance imaging (MRI) remains challenging despite substantial advances in deep learning–based segmentation. Although modern encoder–decoder architectures achieve strong volumetric overlap, precise geometric alignment of tumor contours remains inconsistent, particularly for small lesions and heterogeneous clinical cases. In neuro-oncology, even minor boundary deviations may influence surgical planning, radiotherapy targeting, and longitudinal treatment assessment. These limitations suggest that segmentation performance is not determined solely by network depth or loss design, but also by how multimodal information is structured prior to learning. We introduce CMYD-SurfaceNet, a scale-aware cascaded framework that restructures multimodal MRI inputs at the representation level to enhance boundary-sensitive segmentation. Rather than treating modalities as independently concatenated channels, selected sequences are first organized into a task-guided pseudo-RGB projection. This intermediate representation is subsequently transformed into the CMYK color space to disentangle shared luminance structure from modality-specific contrast dominance. To further encode geometric priors, a gradient-derived boundary density channel is incorporated to explicitly emphasize spatial discontinuities corresponding to tumor margins. The resulting CMYD representation is integrated within a two-stage nnU-Net cascade, where global tumor localization is followed by high-resolution region-of-interest refinement with auxiliary contour supervision. This scale-aware design improves sensitivity to small tumor components while stabilizing contour delineation. Extensive evaluation on the BraTS benchmark demonstrates consistent improvements in boundary-sensitive metrics. Compared with baseline nnU-Net, the proposed framework reduces HD95 from 3.6 mm to 2.4 mm and increases Surface Dice at 1 mm tolerance from 0.82 to 0.89, while maintaining competitive Dice performance. These findings suggest that representation-level structural decoupling, when combined with scale-aware refinement, may provide clinically relevant boundary-aware multimodal MRI segmentation support without increasing architectural complexity. Full article
Show Figures

Figure 1

22 pages, 3661 KB  
Article
Industrial Weld Defect Detection Based on Monocular Depth Estimation and Dual-Attention Point Cloud Network
by Nannan Zhao and Shijie Chen
Sensors 2026, 26(11), 3321; https://doi.org/10.3390/s26113321 - 23 May 2026
Viewed by 646
Abstract
In industrial quality control, the precise identification of severe structural weld defects is paramount. Traditional 2D image-based detection methods are susceptible to illumination and texture interference, while high-precision 3D laser scanning solutions are costly and impractical for large-scale deployment. To achieve reliable geometric [...] Read more.
In industrial quality control, the precise identification of severe structural weld defects is paramount. Traditional 2D image-based detection methods are susceptible to illumination and texture interference, while high-precision 3D laser scanning solutions are costly and impractical for large-scale deployment. To achieve reliable geometric defect detection at low cost, this paper proposes a detection framework based on monocular depth estimation and a dual-attention point cloud network. First, YOLOv8 is employed for rapid region of interest extraction, and an advanced monocular depth estimation model generates 3D pseudo-point clouds containing geometric information. Secondly, addressing the challenge of distinct spatial orientation features in missed weld defects that are prone to confusion, this paper introduces a dual-attention-enhanced point cloud classification network named DA-PointNet++. This model embeds dual-attention modules within the PointNet++ backbone network, enhancing key feature representation in both the channel and spatial dimensions. Experimental results demonstrate that this approach achieves an accuracy of 93.67% and a recall rate of 90.51% in a unified binary classification task for general weld defect detection, effectively identifying both normal welds and complex missed weld defects. Compared to PointConv, Dynamic Graph Convolutional Neural Network (DGCNN), and mainstream Point Cloud Transformer, this method significantly reduces false negative rates while maintaining low computational costs, offering a cost-effective solution for industrial automation. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

14 pages, 3758 KB  
Article
1D U-Net Enhanced QEPAS Sensor for Trace Water Vapor Detection
by Huiming Xiao, Jiahui Wu, Haoyang Lin, Lihao Wang, Jianfeng He, Leqing Lin, Ruobin Zhuang, Guantian Hong, Jiabao Xie, Jianhui Yu, Wenguo Zhu, Yongchun Zhong, Zhigang Song and Huadan Zheng
Optics 2026, 7(1), 15; https://doi.org/10.3390/opt7010015 - 9 Feb 2026
Cited by 2 | Viewed by 1115
Abstract
We report a deep learning-assisted quartz-enhanced photoacoustic spectroscopy (QEPAS) sensor for trace water vapor detection in air. A 1392 nm butterfly-packaged DFB laser is wavelength-modulated at f0/2, and the QEPAS signal is retrieved by second-harmonic (2f) lock-in demodulation using [...] Read more.
We report a deep learning-assisted quartz-enhanced photoacoustic spectroscopy (QEPAS) sensor for trace water vapor detection in air. A 1392 nm butterfly-packaged DFB laser is wavelength-modulated at f0/2, and the QEPAS signal is retrieved by second-harmonic (2f) lock-in demodulation using a commercial quartz tuning fork gas cell. After optimizing the modulation depth to 400 mV, a 1D U-Net denoising network trained with pseudo-clean supervision is applied to the measured 2f traces, yielding an SNR improvement of 2.05× (3.11 dB). Allan deviation analysis indicates a minimum detection limit (MDL) of ~2.21 ppm at an optimum averaging time of ~619 s, corresponding to an ~2.1× improvement compared with the raw output. These results demonstrate that neural-network-based post-processing can improve QEPAS water vapor sensing performance without modifying the optical hardware. Full article
(This article belongs to the Section Laser Sciences and Technology)
Show Figures

Figure 1

16 pages, 1443 KB  
Article
DCRDF-Net: A Dual-Channel Reverse-Distillation Fusion Network for 3D Industrial Anomaly Detection
by Chunshui Wang, Jianbo Chen and Heng Zhang
Sensors 2026, 26(2), 412; https://doi.org/10.3390/s26020412 - 8 Jan 2026
Viewed by 987
Abstract
Industrial surface defect detection is essential for ensuring product quality, but real-world production lines often provide only a limited number of defective samples, making supervised training difficult. Multimodal anomaly detection with aligned RGB and depth data is a promising solution, yet existing fusion [...] Read more.
Industrial surface defect detection is essential for ensuring product quality, but real-world production lines often provide only a limited number of defective samples, making supervised training difficult. Multimodal anomaly detection with aligned RGB and depth data is a promising solution, yet existing fusion schemes tend to overlook modality-specific characteristics and cross-modal inconsistencies, so that defects visible in only one modality may be suppressed or diluted. In this work, we propose DCRDF-Net, a dual-channel reverse-distillation fusion network for unsupervised RGB–depth industrial anomaly detection. The framework learns modality-specific normal manifolds from nominal RGB and depth data and detects defects as deviations from these learned manifolds. It consists of three collaborative components: a Perlin-guided pseudo-anomaly generator that injects appearance–geometry-consistent perturbations into both modalities to enrich training signals; a dual-channel reverse-distillation architecture with guided feature refinement that denoises teacher features and constrains RGB and depth students towards clean, defect-free representations; and a cross-modal squeeze–excitation gated fusion module that adaptively combines RGB and depth anomaly evidence based on their reliability and agreement.Extensive experiments on the MVTec 3D-AD dataset show that DCRDF-Net achieves 97.1% image-level I-AUROC and 98.8% pixel-level PRO, surpassing current state-of-the-art multimodal methods on this benchmark. Full article
(This article belongs to the Section Sensor Networks)
Show Figures

Figure 1

18 pages, 10762 KB  
Article
NRAP-RCNN: A Pseudo Point Cloud 3D Object Detection Method Based on Noise-Reduction Sparse Convolution and Attention Mechanism
by Ziyue Zhou, Yongqing Jia, Tao Zhu and Yaping Wan
Information 2025, 16(3), 176; https://doi.org/10.3390/info16030176 - 26 Feb 2025
Cited by 2 | Viewed by 2208
Abstract
In recent years, pseudo point clouds generated from depth completion of RGB images and LiDAR data have provided a robust foundation for multimodal 3D object detection. However, the generation process often introduces noise, reducing data quality and detection accuracy. Moreover, existing methods fail [...] Read more.
In recent years, pseudo point clouds generated from depth completion of RGB images and LiDAR data have provided a robust foundation for multimodal 3D object detection. However, the generation process often introduces noise, reducing data quality and detection accuracy. Moreover, existing methods fail to effectively capture channel correlations and global contextual information during the 2D feature extraction stage after the 3D backbone network, limiting detection performance. To address these challenges, this paper proposes NRAP-RCNN, a pseudo point cloud-based 3D object detection method with two key innovations: (1) A noise-reduction sparse convolution network (NRConvNet), comprising NRConv (noise-resistant submanifold sparse convolution), SRB (sparse convolution residual block), and MHSA (multi-head self-attention). NRConv suppresses pseudo point cloud noise by jointly encoding 2D and 3D features, SRB enhances feature extraction depth and robustness, and MHSA optimizes global feature representation. (2) An attention fusion module (ECA_GCA) is introduced to enhance the feature representation of the 2D backbone network by combining channel and global contextual information. The experimental results demonstrate that NRAP-RCNN achieves 88.4% car AP (R40) on the KITTI validation set and 85.1% on the test set, significantly outperforming advanced 3D detection methods, showcasing its effectiveness in improving detection performance. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Graphical abstract

22 pages, 2395 KB  
Article
Semi-Supervised Burn Depth Segmentation Network with Contrast Learning and Uncertainty Correction
by Dongxue Zhang and Jingmeng Xie
Sensors 2025, 25(4), 1059; https://doi.org/10.3390/s25041059 - 10 Feb 2025
Cited by 4 | Viewed by 1741
Abstract
Burn injuries are a common traumatic condition, and the early diagnosis of burn depth is crucial for reducing treatment costs and improving survival rates. In recent years, image-based deep learning techniques have been utilized to realize the automation and standardization of burn depth [...] Read more.
Burn injuries are a common traumatic condition, and the early diagnosis of burn depth is crucial for reducing treatment costs and improving survival rates. In recent years, image-based deep learning techniques have been utilized to realize the automation and standardization of burn depth segmentation. However, the scarcity and difficulty in labeling burn data limit the performance of traditional deep learning-based segmentation methods. Mainstream semi-supervised methods face challenges in burn depth segmentation due to single-level perturbations, lack of explicit edge modeling, and ineffective handling of inaccurate predictions in unlabeled data. To address these issues, we propose SBCU-Net, a semi-supervised burn depth segmentation network with contrastive learning and uncertainty correction. Building on the LTB-Net from our previous work, SBCU-Net introduces two additional decoder branches to enhance the consistency between the probability map and soft pseudo-labels under multi-level perturbations. To improve segmentation in complex regions like burn edges, contrastive learning refines the outputs of the three-branch decoder, enabling more discriminative feature representation learning. In addition, an uncertainty correction mechanism weights the consistency loss based on prediction uncertainty, reducing the impact of inaccurate pseudo-labels. Extensive experiments on burn datasets demonstrate that SBCU-Net effectively leverages unlabeled data and achieves superior performance compared to state-of-the-art semi-supervised methods. Full article
(This article belongs to the Section Biomedical Sensors)
Show Figures

Figure 1

19 pages, 5171 KB  
Article
Quantification of Nearshore Sandbar Seasonal Evolution Based on Drone Pseudo-Bathymetry Time-Lapse Data
by Evangelos Alevizos
Remote Sens. 2024, 16(23), 4551; https://doi.org/10.3390/rs16234551 - 4 Dec 2024
Cited by 5 | Viewed by 3579
Abstract
Nearshore sandbars are dynamic features that characterize shallow morphobathymetry and vary over a wide range of geometries and temporal lifespans. Nearshore sandbars influence beach geometry by altering the energy of incoming waves; thus, monitoring the evolution of sandbars is a fundamental approach in [...] Read more.
Nearshore sandbars are dynamic features that characterize shallow morphobathymetry and vary over a wide range of geometries and temporal lifespans. Nearshore sandbars influence beach geometry by altering the energy of incoming waves; thus, monitoring the evolution of sandbars is a fundamental approach in effective coastal planning. Due to several natural and technical limitations related to shallow seafloor mapping, there is a significant gap in the availability of high-resolution, shallow bathymetric data for monitoring the dynamic behaviour of nearshore sandbars effectively. This study introduces a novel image-processing technique that produces time series of pseudo-bathymetric data by utilizing multi-temporal (monthly) drone imagery, and it provides an assessment of local morphodynamics at a sandy beach in the southeast Mediterranean. The technique is called standardized-ratio bathymetric index (SRBI), and it transforms natural-colour drone imagery to pseudo-bathymetric data by applying an empirical formula used for satellite-derived bathymetry. This technique correlates well with laser altimetry depth measurements; however, it does not require in situ depth data for implementation. The resulting pseudo-bathymetric data allows for extracting cross-shore profiles and delineating the sandbar crest with 4 m horizontal accuracy. Stacking of temporal profiles allowed for the quantification of the sandbar’s crest and trough changes at different alongshore sections. The main findings suggest that the nearshore crescentic sandbar at Episkopi Beach (north Crete) shows strong seasonality regarding net offshore migration that is promoted by enhanced wave action during winter months. In addition, the crescentic sandbar is susceptible to morphology arrestment during prolonged weeks of low wave action. The average migration rate during winter is 10 m.month−1, with some sections exhibiting a maximum of 60 m.month−1. This study (a) offers a novel remote-sensing approach, suitable for nearshore seafloor monitoring with low computational complexity, (b) reveals sandbar geometry and temporal change in superior detail compared to other observational methods, and (c) advances knowledge about nearshore sandbar monitoring in the Mediterranean region. Full article
(This article belongs to the Section Remote Sensing in Geology, Geomorphology and Hydrology)
Show Figures

Graphical abstract

22 pages, 5609 KB  
Article
Road Anomaly Detection with Unknown Scenes Using DifferNet-Based Automatic Labeling Segmentation
by Phuc Thanh-Thien Nguyen, Toan-Khoa Nguyen, Dai-Dong Nguyen, Shun-Feng Su and Chung-Hsien Kuo
Inventions 2024, 9(4), 69; https://doi.org/10.3390/inventions9040069 - 28 Jun 2024
Cited by 4 | Viewed by 5019
Abstract
Obstacle avoidance is essential for the effective operation of autonomous mobile robots, enabling them to detect and navigate around obstacles in their environment. While deep learning provides significant benefits for autonomous navigation, it typically requires large, accurately labeled datasets, making the data’s preparation [...] Read more.
Obstacle avoidance is essential for the effective operation of autonomous mobile robots, enabling them to detect and navigate around obstacles in their environment. While deep learning provides significant benefits for autonomous navigation, it typically requires large, accurately labeled datasets, making the data’s preparation and processing time-consuming and labor-intensive. To address this challenge, this study introduces a transfer learning (TL)-based automatic labeling segmentation (ALS) framework. This framework utilizes a pretrained attention-based network, DifferNet, to efficiently perform semantic segmentation tasks on new, unlabeled datasets. DifferNet leverages prior knowledge from the Cityscapes dataset to identify high-entropy areas as road obstacles by analyzing differences between the input and resynthesized images. The resulting road anomaly map was refined using depth information to produce a robust drivable area and map of road anomalies. Several off-the-shelf RGB-D semantic segmentation neural networks were trained using pseudo-labels generated by the ALS framework, with validation conducted on the GMRPD dataset. Experimental results demonstrated that the proposed ALS framework achieved mean precision, mean recall, and mean intersection over union (IoU) rates of 80.31%, 84.42%, and 71.99%, respectively. The ALS framework, through the use of transfer learning and the DifferNet network, offers an efficient solution for semantic segmentation of new, unlabeled datasets, underscoring its potential for improving obstacle avoidance in autonomous mobile robots. Full article
Show Figures

Graphical abstract

19 pages, 14351 KB  
Article
A Deep Joint Network for Monocular Depth Estimation Based on Pseudo-Depth Supervision
by Jiahai Tan, Ming Gao, Tao Duan and Xiaomei Gao
Mathematics 2023, 11(22), 4645; https://doi.org/10.3390/math11224645 - 14 Nov 2023
Cited by 1 | Viewed by 2706
Abstract
Depth estimation from a single image is a significant task. Although deep learning methods hold great promise in this area, they still face a number of challenges, including the limited modeling of nonlocal dependencies, lack of effective loss function joint optimization models, and [...] Read more.
Depth estimation from a single image is a significant task. Although deep learning methods hold great promise in this area, they still face a number of challenges, including the limited modeling of nonlocal dependencies, lack of effective loss function joint optimization models, and difficulty in accurately estimating object edges. In order to further increase the network’s prediction accuracy, a new structure and training method are proposed for single-image depth estimation in this research. A pseudo-depth network is first deployed for generating a single-image depth prior, and by constructing connecting paths between multi-scale local features using the proposed up-mapping and jumping modules, the network can integrate representations and recover fine details. A deep network is also designed to capture and convey global context by utilizing the Transformer Conv module and Unet Depth net to extract and refine global features. The two networks jointly provide meaningful coarse and fine features to predict high-quality depth images from single RGB images. In addition, multiple joint losses are utilized to enhance the training model. A series of experiments are carried out to confirm and demonstrate the efficacy of our method. The proposed method exceeds the advanced method DPT by 10% and 3.3% in terms of root mean square error (RMSE(log)) and 1.7% and 1.6% in terms of squared relative difference (SRD), respectively, according to experimental results on the NYU Depth V2 and KITTI depth estimation benchmarks. Full article
(This article belongs to the Special Issue Advances in Computer Vision and Machine Learning)
Show Figures

Figure 1

19 pages, 33702 KB  
Article
Detection of Fittings Based on the Dynamic Graph CNN and U-Net Embedded with Bi-Level Routing Attention
by Zhihui Xie, Min Fu and Xuefeng Liu
Electronics 2023, 12(22), 4611; https://doi.org/10.3390/electronics12224611 - 11 Nov 2023
Cited by 3 | Viewed by 3081
Abstract
Accurate detection of power fittings is crucial for identifying defects or faults in these components, which is essential for assessing the safety and stability of the power system. However, the accuracy of fittings detection is affected by a complex background, small target sizes, [...] Read more.
Accurate detection of power fittings is crucial for identifying defects or faults in these components, which is essential for assessing the safety and stability of the power system. However, the accuracy of fittings detection is affected by a complex background, small target sizes, and overlapping fittings in the images. To address these challenges, a fittings detection method based on the dynamic graph convolutional neural network (DGCNN) and U-shaped network (U-Net) is proposed, which combines three-dimensional detection with two-dimensional object detection. Firstly, the bi-level routing attention mechanism is incorporated into the lightweight U-Net network to enhance feature extraction for detecting the fittings boundary. Secondly, pseudo-point cloud data are synthesized by transforming the depth map generated by the Lite-Mono algorithm and its corresponding RGB fittings image. The DGCNN algorithm is then employed to extract obscured fittings features, contributing to the final refinement of the results. This process helps alleviate the issue of occlusions among targets and further enhances the precision of fittings detection. Finally, the proposed method is evaluated using a custom dataset of fittings, and comparative studies are conducted. The experimental results illustrate the promising potential of the proposed approach in enhancing features and extracting information from fittings images. Full article
(This article belongs to the Special Issue Advances in Computer Vision and Deep Learning and Its Applications)
Show Figures

Figure 1

13 pages, 1759 KB  
Article
Efficient Stereo Depth Estimation for Pseudo-LiDAR: A Self-Supervised Approach Based on Multi-Input ResNet Encoder
by Sabir Hossain and Xianke Lin
Sensors 2023, 23(3), 1650; https://doi.org/10.3390/s23031650 - 2 Feb 2023
Cited by 9 | Viewed by 5421
Abstract
Perception and localization are essential for autonomous delivery vehicles, mostly estimated from 3D LiDAR sensors due to their precise distance measurement capability. This paper presents a strategy to obtain a real-time pseudo point cloud from image sensors (cameras) instead of laser-based sensors (LiDARs). [...] Read more.
Perception and localization are essential for autonomous delivery vehicles, mostly estimated from 3D LiDAR sensors due to their precise distance measurement capability. This paper presents a strategy to obtain a real-time pseudo point cloud from image sensors (cameras) instead of laser-based sensors (LiDARs). Previous studies (such as PSMNet-based point cloud generation) built the algorithm based on accuracy but failed to operate in real time as LiDAR. We propose an approach to use different depth estimators to obtain pseudo point clouds similar to LiDAR to achieve better performance. Moreover, the depth estimator has used stereo imagery data to achieve more accurate depth estimation as well as point cloud results. Our approach to generating depth maps outperforms other existing approaches on KITTI depth prediction while yielding point clouds significantly faster than other approaches as well. Additionally, the proposed approach is evaluated on the KITTI stereo benchmark, where it shows effectiveness in runtime. Full article
Show Figures

Figure 1

Back to TopTop