Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (61)

Search Parameters:
Keywords = automotive image dataset

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
38 pages, 33093 KB  
Article
Fast-DG2GAN: A Computationally Efficient DG2GAN Variant for Industrial Injection Molding Splay Defect Generation
by Timothy Reinhart, Seshasai Srinivasan and Zhen Gao
Machines 2026, 14(8), 888; https://doi.org/10.3390/machines14080888 - 4 Aug 2026
Viewed by 240
Abstract
Synthetic data generation is a potential solution for addressing limited data in manufacturing defect detection. In injection molding of automotive tubes, surface defects such as splay present as white or silver streaks in the tube’s texture, that are difficult to capture in sufficient [...] Read more.
Synthetic data generation is a potential solution for addressing limited data in manufacturing defect detection. In injection molding of automotive tubes, surface defects such as splay present as white or silver streaks in the tube’s texture, that are difficult to capture in sufficient quantity for training robust object detection models. This study proposes Fast-DG2GAN: a DG2GAN based, computationally optimized, generative style defect generator for manufacturing defect images. Fast-DG2GAN achieves a 56% reduction in training time relative to the original DG2GAN, completing training in 334.9 min compared to 768.0 min, while maintaining comparable image quality metrics with a best FID score of 132.83 and IS 1.45 ± 0.08. Key contributions to this DG2GAN variant include depth-wise separable convolutions, reduced residual blocks, and automatic mixed precision (AMP) training, to improve training time. Training stabilization techniques include perceptual loss, feature matching, and exponential moving average of weights (EMA). Legacy Generative Adversarial Network (GAN) architectures are benchmarked for feasibility and include WGAN, DCGAN, and FastGAN, for fine-grained defect image generation essential to downstream object detection. Metrics, such as Inception Score (IS) and Fréchet Inception Distance (FID) are used for quantitative performance evaluation. This study highlights the potential of GAN-generated datasets to augment real-world training for defect detection models. Full article
Show Figures

Figure 1

17 pages, 15332 KB  
Article
Research on Recognition Method of Carbon Deposit Degree Based on Improved EfficientNet-B0
by Yongfeng Yue, Ping Chen, Simeng Ma and Youxing Chen
Electronics 2026, 15(15), 3333; https://doi.org/10.3390/electronics15153333 - 28 Jul 2026
Viewed by 240
Abstract
Accurate classification of carbon deposit severity for automotive engines is critical for routine vehicle inspection and exhaust emission pollution control. Although deep learning-based methods have made certain progress in carbon deposit visual detection, they face prominent drawbacks in real industrial deployment because endoscopic [...] Read more.
Accurate classification of carbon deposit severity for automotive engines is critical for routine vehicle inspection and exhaust emission pollution control. Although deep learning-based methods have made certain progress in carbon deposit visual detection, they face prominent drawbacks in real industrial deployment because endoscopic carbon deposit images carry subtle inter-class fine-grained textures and lack sufficient labeled training samples, which severely impair conventional models’ classification performance and prevent them from satisfying the accuracy requirements of practical inspection tasks. This paper proposes an improved EfficientNet-B0 framework for fine-grained identification of cylinder wall carbon deposit severity. By introducing the DiffuseMix data augmentation algorithm, label confusion and overfitting are avoided. At the same time, a dual-path weighted feature fusion structure (DPWFFS) is designed to achieve bidirectional complementary of shallow and deep features, enhancing the representation ability of local discriminative features of carbon deposits. Quantitative experiments conducted on our self-built cylinder wall carbon deposit dataset demonstrate that the proposed improved model achieves a classification accuracy of 90.33%, which substantially outperforms the original EfficientNet-B0 baseline. The presented method provides an effective technical reference for the intelligent automatic identification of engine carbon deposit severity. Full article
Show Figures

Figure 1

29 pages, 1048 KB  
Article
Composite Gramian Angular Field for Time-Series Classification
by Pero Bogunović, Saša Mladenović and Andrina Granić
Information 2026, 17(7), 640; https://doi.org/10.3390/info17070640 - 30 Jun 2026
Viewed by 454
Abstract
Gramian Angular Field (GAF) encodings transform time series into two-dimensional images suitable for convolutional neural network (CNN) classification. Existing applications typically use either the Gramian Angular Summation Field (GASF) or the Gramian Angular Difference Field (GADF) independently, although these two encodings capture complementary [...] Read more.
Gramian Angular Field (GAF) encodings transform time series into two-dimensional images suitable for convolutional neural network (CNN) classification. Existing applications typically use either the Gramian Angular Summation Field (GASF) or the Gramian Angular Difference Field (GADF) independently, although these two encodings capture complementary pairwise angular relationships. This paper proposes the Composite Gramian Angular Field (CGAF), a single-image time-series representation obtained by a weighted algebraic combination of the summation and difference GAF components. The weights are optimised using coarse grid search followed by Gaussian-process Bayesian refinement, with all candidate evaluation restricted to training-only inner validation partitions. The selected weights are frozen before held-out test evaluation. CGAF produces a single encoded output image (approximately 0.08 MB, compared with approximately 0.16 MB for retaining separate GASF and GADF images) and encodes at 5.9±0.3 ms per sample. We evaluate CGAF in three domain-specific settings—EEG cognitive engagement, PTB-DB heartbeat classification, and FordA automotive fault detection—and on a selected subset of 20 datasets from the UCR Time Series Classification Archive. The method is compared with GASF, GADF, recurrence plots, spectrogram-based encodings, and non-image time-series baselines including SVM, ResNet-1D, InceptionTime, and ROCKET. On the evaluated datasets, CGAF consistently improves over the individual GASF and GADF encodings. It achieves macro-F1 =0.867±0.027 on the EEG pilot study, heartbeat-segment-level macro-F1 =0.941±0.018 on PTB-DB, and test accuracy =91.2% on FordA. Because patient identifiers are unavailable for PTB-DB, that result does not establish patient-level generalisation. On the selected UCR subset, CGAF outperforms both GASF and GADF on all 20 datasets. It achieves the best overall accuracy among all evaluated methods on 14 of 20 datasets, whereas ROCKET achieves the best overall accuracy on the remaining six datasets. The results suggest that algebraic integration of summation-based and difference-based angular dependencies can improve image-based time-series classification without modifying the CNN backbone or adding gradient-trained parameters. The EEG results should be interpreted as pilot evidence, whereas broader generalisation requires evaluation on the full UCR/UEA archive, additional biomedical cohorts, and further backbone architectures. Full article
(This article belongs to the Special Issue Signal Processing and Machine Learning, 2nd Edition)
Show Figures

Figure 1

18 pages, 15288 KB  
Article
HUD-DPCNet: A Joint Learning Framework for Distortion Pre-Correction in AR-HUD Systems
by Ying Huang, Huaixin Chen and Zhixi Wang
Appl. Sci. 2026, 16(13), 6361; https://doi.org/10.3390/app16136361 - 25 Jun 2026
Viewed by 383
Abstract
As a next-generation automotive display technology, Augmented Reality Head-Up Display (AR-HUD) has demonstrated immense potential in reshaping driving safety and enhancing the human–computer interaction experience. To address the challenges of barrel distortion and perspective distortion inherent in HUD systems, we propose a joint-learning-based [...] Read more.
As a next-generation automotive display technology, Augmented Reality Head-Up Display (AR-HUD) has demonstrated immense potential in reshaping driving safety and enhancing the human–computer interaction experience. To address the challenges of barrel distortion and perspective distortion inherent in HUD systems, we propose a joint-learning-based dual-path pre-correction method. This approach employs a shared encoder to extract image features, which are then decoupled into two parallel branches: a classification branch and a distortion flow prediction branch. Building upon this architecture, a model-fitting method is introduced to estimate the distortion model parameters in the parameter space using the predicted distortion types and flows, thereby reconstructing a refined distortion flow. Finally, image rectification is achieved through a resampling method. On the ARHDD dataset, the proposed method achieves a PSNR of 24.617 dB (barrel) and 25.062 dB (perspective), an SSIM of 0.845 and 0.873, and an NRMSE of 0.163 and 0.157, respectively. On the Places 365 dataset, it achieves a PSNR of 23.914 dB (barrel) and 21.870 dB (perspective), an SSIM of 0.812 and 0.748, and an NRMSE of 0.174 and 0.211, respectively. Both quantitative and qualitative comparative experiments against other state-of-the-art methods demonstrate that the proposed approach achieves superior correction performance for both types of distortion. Finally, the simulation verification of the HUD system proved that this correction method demonstrated excellent potential, but further verification is still needed in a real or semi-real environment. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

24 pages, 1020 KB  
Article
Research on the Diagnosis of Abnormal Sound Defects in Automobile Engines Based on Fusion of Multi-Modal Images and Audio
by Yi Xu, Wenbo Chen and Xuedong Jing
Electronics 2026, 15(7), 1406; https://doi.org/10.3390/electronics15071406 - 27 Mar 2026
Viewed by 667
Abstract
Against the global carbon neutrality target, predictive maintenance (PdM) of automotive engines represents a core technical strategy to advance the sustainable development of the automotive industry. Conventional single-modal diagnostic approaches for engine abnormal sound defects suffer from low accuracy and weak anti-interference capability. [...] Read more.
Against the global carbon neutrality target, predictive maintenance (PdM) of automotive engines represents a core technical strategy to advance the sustainable development of the automotive industry. Conventional single-modal diagnostic approaches for engine abnormal sound defects suffer from low accuracy and weak anti-interference capability. Existing multi-modal fusion methods fail to deeply mine the physical coupling between cross-modal features and often entail excessive model complexity, hindering deployment on resource-constrained on-board edge devices. To resolve these limitations, this study proposes a Physical Prior-Embedded Cross-Modal Attention (PPE-CMA) mechanism for lightweight multi-modal fusion diagnosis of engine abnormal sound defects. First, wavelet packet decomposition (WPD) and mel-frequency cepstral coefficients (MFCC) are integrated to extract time-frequency features from engine audio signals, while a channel-pruned ResNet18 is employed to extract spatial features from engine thermal imaging and vibration visualization images. Second, the PPE-CMA module is designed to adaptively assign attention weights to audio and image features by exploiting the physical coupling between engine fault acoustic and visual characteristics, enabling efficient cross-modal feature fusion with redundant information suppression. A rigorous theoretical derivation is provided to link cosine similarity with the physical correlation of engine fault acoustic-visual features, justifying the attention weight constraint (β = 1 − α) from the perspective of fault feature physical coupling. Third, an improved lightweight XGBoost classifier is constructed for fault classification, and a hybrid data augmentation strategy customized for engine multi-modal data is proposed to address the small-sample challenge in industrial applications. Ablation experiments on ResNet18 pruning ratios verify the optimal trade-off between diagnostic performance and computational efficiency, while feature distribution analysis validates the authenticity and effectiveness of the hybrid augmentation strategy. Experimental results on a self-constructed multi-modal dataset show that the proposed method achieves 98.7% diagnostic accuracy and a 98.2% F1-score, retaining 96.5% accuracy under 90 dB high-level environmental noise, with an end-to-end inference speed of 0.8 ms per sample (including preprocessing, feature extraction, and classification). Cross-engine and cross-domain validation on a 2.0T diesel engine small-sample dataset and the open-source SEMFault-2024 dataset yield average accuracies of 94.8% and 95.2%, respectively, demonstrating strong generalization. This method effectively enhances the accuracy and robustness of engine abnormal sound defect diagnosis, offering a lightweight technical solution for on-board real-time fault diagnosis and in-plant online quality inspection. By reducing engine fault-induced energy loss and spare parts waste, it further promotes energy conservation and emission reduction in the automotive industry. Quantified experimental data on fuel efficiency improvement and carbon emission reduction are provided to substantiate the ecological benefits of the proposed framework. Full article
Show Figures

Figure 1

33 pages, 88715 KB  
Article
A Co-Designed Framework Combining Dome-Aperture Imaging and Generative AI for Defect Detection on Non-Planar Metal Surfaces
by Zhongqing Jia, Zhaohui Yu, Chen Guan, Bing Zhao and Xiaofei Wang
Sensors 2026, 26(3), 1044; https://doi.org/10.3390/s26031044 - 5 Feb 2026
Viewed by 691
Abstract
Automated visual inspection of safety-critical metal assemblies such as automotive door lock strikes remains challenging due to their complex three-dimensional geometry, highly reflective surfaces, and scarcity of defect samples. While 3D sensing technologies are often constrained by cost and speed, traditional 2D optical [...] Read more.
Automated visual inspection of safety-critical metal assemblies such as automotive door lock strikes remains challenging due to their complex three-dimensional geometry, highly reflective surfaces, and scarcity of defect samples. While 3D sensing technologies are often constrained by cost and speed, traditional 2D optical methods struggle with severe imaging artifacts and poor generalization under few-shot conditions. This work constructs a complete system integrating defect imaging, generation, and detection. It proposes an integrated framework through the co-design of an image acquisition system and deep generative models to holistically enhance defect perception capability. First, we develop an imaging system using dome illumination and a small-aperture lens to acquire high-quality images of non-planar metal surfaces. Subsequently, we introduce a dual-stage generation strategy: stage one employs an improved FastGAN with Dynamic Multi-Granularity Fusion Skip-Layer Excitation (DMGF-SLE) and perceptual loss to efficiently generate high-quality local defect patches; stage two utilizes Poisson image editing and an optimized loss function to seamlessly fuse defect patches into specified locations of normal images. This strategy avoids modeling the complete complex background, concentrating computational resources on creating realistic defects. Experiments on a dedicated dataset demonstrate that our method can efficiently generate realistic defect samples under few-shot conditions, achieving 11–24% improvement in Fréchet Inception Distance (FID) scores over baseline models. The generated synthetic data significantly enhances downstream detection performance, increasing YOLOv8’s mAP@50:95 from 50.4% to 60.5%. Beyond proposing individual technical improvements, this research provides a complete, synergistic, and deployable system solution—from physical imaging to algorithmic generation—delivering a computationally efficient and practically viable technical pathway for defect detection in highly reflective, non-planar metal components. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

19 pages, 5679 KB  
Article
SDDNet: Two-Stage Network for Forgings Surface Defect Detection
by Shentao Wang, Depeng Gao, Byung-Won Min, Yue Hong, Tingting Xu and Zhongyue Xiong
Symmetry 2026, 18(1), 104; https://doi.org/10.3390/sym18010104 - 6 Jan 2026
Viewed by 750
Abstract
Detecting surface defects in forgings is crucial for ensuring the reliability of automotive components such as steering knuckles. In fluorescent magnetic particle inspection (FDMPI) images, normal forging surfaces generally exhibit locally symmetric texture patterns, whereas cracks and other flaws appear as locally asymmetric [...] Read more.
Detecting surface defects in forgings is crucial for ensuring the reliability of automotive components such as steering knuckles. In fluorescent magnetic particle inspection (FDMPI) images, normal forging surfaces generally exhibit locally symmetric texture patterns, whereas cracks and other flaws appear as locally asymmetric regions. Traditional FDMPI inspection relies on manual visual judgement, which is inefficient and error-prone. This paper introduces SDDNet, a symmetry-aware deep learning model for surface defect detection in FDMPI images. A dedicated FDMPI dataset is constructed and further expanded using a denoising diffusion probabilistic model (DDPM) to improve training robustness. To better separate symmetric background textures from asymmetric defect cues, SDDNet integrates a UPerNet-based segmentation layer for background suppression and a Scale-Variant Inception Module (SVIM) within an RTMDet framework for multi-scale feature extraction. Experiments show that SDDNet effectively suppresses background noise and significantly improves detection accuracy, achieving a mean average precision (mAP) of 45.5% on the FDMPI dataset, 19% higher than the baseline, and 71.5% mAP on the NEU-DET dataset, outperforming existing methods by up to 8.1%. Full article
(This article belongs to the Special Issue Symmetry/Asymmetry in Image Processing and Computer Vision)
Show Figures

Figure 1

16 pages, 1492 KB  
Article
TGDNet: A Multi-Scale Feature Fusion Defect Detection Method for Transparent Industrial Headlight Glass
by Zefan Zhang and Jin Tang
Sensors 2025, 25(24), 7437; https://doi.org/10.3390/s25247437 - 6 Dec 2025
Viewed by 1083
Abstract
In industrial production, defect detection for automotive headlight lenses is an essential yet challenging task. Transparent glass defect detection faces several difficulties, including a wide variety of defect shapes and sizes, as well as the challenge of identifying transparent surface defects. To enhance [...] Read more.
In industrial production, defect detection for automotive headlight lenses is an essential yet challenging task. Transparent glass defect detection faces several difficulties, including a wide variety of defect shapes and sizes, as well as the challenge of identifying transparent surface defects. To enhance the accuracy and efficiency of this process, we propose a computer vision-based inspection solution utilizing multi-angle lighting. For this task, we collected 2000 automotive headlight images to systematically categorize defects in transparent glass, with the primary defect types being spots, scratches, and abrasions. During data acquisition, we proposed a dataset augmentation method named SWAM to address class imbalance, ultimately generating the Lens Defect Dataset (LDD), which comprises 5532 images across these three main defect categories. Furthermore, we propose a defect detection network named the Transparent Glass Defect Network (TGDNet), designed based on common transparent glass defect types. Within the backbone of TGDNet, we introduced the TGFE module to adaptively extract local features for different defect categories and employed TGD, an improved SK attention mechanism, combined with a spatial attention mechanism to boost the network’s capability in multi-scale feature fusion. Experiments demonstrate that compared to other classical defect detection methods, TGDNet achieves superior performance on the LDD, improving the average detection precision by 6.7% in mAP and 8.9% in mAP50 over the highest-performing baseline algorithm. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

20 pages, 815 KB  
Article
Automotive Scratch Detection: A Lightweight Convolutional Network Approach Augmented by Generative Adversarial Learning
by Guojie Qu, Jiaying Liao, Kai Liu, Bin Xu and Yuwen Qian
Machines 2025, 13(12), 1107; https://doi.org/10.3390/machines13121107 - 29 Nov 2025
Viewed by 1024
Abstract
The growing demand for high-precision machining and inspection in modern manufacturing has positioned machine vision as a key technology for surface defect detection. However, identifying subtle surface scratches on automotive components remains a challenging task due to the stringent requirements on sensitivity, precision, [...] Read more.
The growing demand for high-precision machining and inspection in modern manufacturing has positioned machine vision as a key technology for surface defect detection. However, identifying subtle surface scratches on automotive components remains a challenging task due to the stringent requirements on sensitivity, precision, and robustness against complex background interference. In this paper, we propose an automated detection system with a Convolutional Neural Network (CNN) architecture. To address data scarcity, we construct a large-scale, high-quality dataset using both data augmentation and Generative Adversarial Network (GAN)-based synthesis. Furthermore, the proposed lightweight CNN replaces traditional fully connected layers with one-dimensional convolutional layers to reduce parameter complexity and model size, while a Dropout mechanism is incorporated to mitigate overfitting and enhance generalization. Experimental results demonstrate that the proposed model achieves superior detection accuracy and robustness across diverse imaging conditions. Moreover, the developed system effectively addresses the limitations of data insufficiency and model complexity, offering an efficient and automated solution for surface quality inspection in industrial manufacturing. Full article
(This article belongs to the Section Machines Testing and Maintenance)
Show Figures

Figure 1

21 pages, 3520 KB  
Article
SMPPALD—Segmentation Mask Post-Processing Algorithm for Improved Lane Detection
by Denis Vajak, Mario Vranješ, Ratko Grbić and Denis Vranješ
Sensors 2025, 25(19), 6057; https://doi.org/10.3390/s25196057 - 2 Oct 2025
Viewed by 1084
Abstract
As modern Advanced Driver Assistance Systems become increasingly prevalent in the automotive industry, Lane Detection (LD) solutions play a key role in enabling vehicles to drive autonomously or provide assistance to the driver. Many modern LD algorithms are based on neural networks, which [...] Read more.
As modern Advanced Driver Assistance Systems become increasingly prevalent in the automotive industry, Lane Detection (LD) solutions play a key role in enabling vehicles to drive autonomously or provide assistance to the driver. Many modern LD algorithms are based on neural networks, which estimate the locations of lane markings as segmentation masks in the input image. In this paper, we propose a novel algorithm, named SMPPALD (Segmentation Mask Post-Processing Algorithm for improved Lane Detection), designed to perform a set of post-processing operations on these segmentation masks to produce a list of points that define the lane markings. These operations follow geometric and contextual rules, taking into account the LD problem and improving detection accuracy. The algorithm was tested using the well-known and widely used Spatial Convolutional Neural Network (SCNN) on three different datasets (CULane, TuSimple, and LLAMAS). SMPPALD achieved a significant improvement in terms of F1 measure compared to SCNN on the TuSimple and LLAMAS datasets, while for the CULane dataset, it outperformed SCNN in most categories. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

16 pages, 1170 KB  
Article
LoRA-Tuned Multimodal RAG System for Technical Manual QA: A Case Study on Hyundai Staria
by Yerin Nam, Hansun Choi, Jonggeun Choi and Hyukjin Kwon
Appl. Sci. 2025, 15(15), 8387; https://doi.org/10.3390/app15158387 - 29 Jul 2025
Cited by 1 | Viewed by 4429
Abstract
This study develops a domain-adaptive multimodal RAG (Retrieval-Augmented Generation) system to improve the accuracy and efficiency of technical question answering based on large-scale structured manuals. Using Hyundai Staria maintenance documents as a case study, we extracted text and images from PDF manuals and [...] Read more.
This study develops a domain-adaptive multimodal RAG (Retrieval-Augmented Generation) system to improve the accuracy and efficiency of technical question answering based on large-scale structured manuals. Using Hyundai Staria maintenance documents as a case study, we extracted text and images from PDF manuals and constructed QA, RAG, and Multi-Turn datasets to reflect realistic troubleshooting scenarios. To overcome limitations of baseline RAG models, we proposed an enhanced architecture that incorporates sentence-level similarity annotations and parameter-efficient fine-tuning via LoRA (Low-Rank Adaptation) using the bLLossom-8B language model and BAAI-bge-m3 embedding model. Experimental results show that the proposed system achieved improvements of 3.0%p in BERTScore, 3.0%p in cosine similarity, and 18.0%p in ROUGE-L compared to existing RAG systems, with notable gains in image-guided response accuracy. A qualitative evaluation by 20 domain experts yielded an average satisfaction score of 4.4 out of 5. This study presents a practical and extensible AI framework for multimodal document understanding, with broad applicability across automotive, industrial, and defense-related technical documentation. Full article
(This article belongs to the Special Issue Innovations in Artificial Neural Network Applications)
Show Figures

Figure 1

27 pages, 27475 KB  
Article
LiGenCam: Reconstruction of Color Camera Images from Multimodal LiDAR Data for Autonomous Driving
by Minghao Xu, Yanlei Gu, Igor Goncharenko and Shunsuke Kamijo
Sensors 2025, 25(14), 4295; https://doi.org/10.3390/s25144295 - 10 Jul 2025
Viewed by 1704
Abstract
The automotive industry is advancing toward fully automated driving, where perception systems rely on complementary sensors such as LiDAR and cameras to interpret the vehicle’s surroundings. For Level 4 and higher vehicles, redundancy is vital to prevent safety-critical failures. One way to achieve [...] Read more.
The automotive industry is advancing toward fully automated driving, where perception systems rely on complementary sensors such as LiDAR and cameras to interpret the vehicle’s surroundings. For Level 4 and higher vehicles, redundancy is vital to prevent safety-critical failures. One way to achieve this is by using data from one sensor type to support another. While much research has focused on reconstructing LiDAR point cloud data using camera images, limited work has been conducted on the reverse process—reconstructing image data from LiDAR. This paper proposes a deep learning model, named LiDAR Generative Camera (LiGenCam), to fill this gap. The model reconstructs camera images by utilizing multimodal LiDAR data, including reflectance, ambient light, and range information. LiGenCam is developed based on the Generative Adversarial Network framework, incorporating pixel-wise loss and semantic segmentation loss to guide reconstruction, ensuring both pixel-level similarity and semantic coherence. Experiments on the DurLAR dataset demonstrate that multimodal LiDAR data enhances the realism and semantic consistency of reconstructed images, and adding segmentation loss further improves semantic consistency. Ablation studies confirm these findings. Full article
(This article belongs to the Special Issue Recent Advances in LiDAR Sensing Technology for Autonomous Vehicles)
Show Figures

Figure 1

29 pages, 4405 KB  
Article
Pupil Detection Algorithm Based on ViM
by Yu Zhang, Changyuan Wang, Pengbo Wang and Pengxiang Xue
Sensors 2025, 25(13), 3978; https://doi.org/10.3390/s25133978 - 26 Jun 2025
Cited by 1 | Viewed by 1917
Abstract
Pupil detection is a key technology in fields such as human–computer interaction, fatigue driving detection, and medical diagnosis. Existing pupil detection algorithms still face challenges in maintaining robustness under variable lighting conditions and occlusion scenarios. In this paper, we propose a novel pupil [...] Read more.
Pupil detection is a key technology in fields such as human–computer interaction, fatigue driving detection, and medical diagnosis. Existing pupil detection algorithms still face challenges in maintaining robustness under variable lighting conditions and occlusion scenarios. In this paper, we propose a novel pupil detection algorithm, ViMSA, based on the ViM model. This algorithm introduces weighted feature fusion, aiming to enable the model to adaptively learn the contribution of different feature patches to the pupil detection results; combines ViM with the MSA (multi-head self-attention) mechanism), aiming to integrate global features and improve the accuracy and robustness of pupil detection; and uses FFT (Fast Fourier Transform) to convert the time-domain vector outer product in MSA into a frequency–domain dot product, in order to reduce the computational complexity of the model and improve the detection efficiency of the model. ViMSA was trained and tested on nearly 135,000 pupil images from 30 different datasets, demonstrating exceptional generalization capability. The experimental results demonstrate that the proposed ViMSA achieves 99.6% detection accuracy at five pixels with an RMSE of 1.67 pixels and a processing speed exceeding 100 FPS, meeting real-time monitoring requirements for various applications including operation under variable and uneven lighting conditions, assistive technology (enabling communication with neuro-motor disorder patients through pupil recognition), computer gaming, and automotive industry applications (enhancing traffic safety by monitoring drivers’ cognitive states). Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

18 pages, 10317 KB  
Article
Advanced Thermal Imaging Processing and Deep Learning Integration for Enhanced Defect Detection in Carbon Fiber-Reinforced Polymer Laminates
by Renan Garcia Rosa, Bruno Pereira Barella, Iago Garcia Vargas, José Ricardo Tarpani, Hans-Georg Herrmann and Henrique Fernandes
Materials 2025, 18(7), 1448; https://doi.org/10.3390/ma18071448 - 25 Mar 2025
Cited by 16 | Viewed by 2985
Abstract
Carbon fiber-reinforced polymer (CFRP) laminates are widely used in aerospace, automotive, and infrastructure industries due to their high strength-to-weight ratio. However, defect detection in CFRP remains challenging, particularly in low signal-to-noise ratio (SNR) conditions. Conventional segmentation methods often struggle with noise interference and [...] Read more.
Carbon fiber-reinforced polymer (CFRP) laminates are widely used in aerospace, automotive, and infrastructure industries due to their high strength-to-weight ratio. However, defect detection in CFRP remains challenging, particularly in low signal-to-noise ratio (SNR) conditions. Conventional segmentation methods often struggle with noise interference and signal variations, leading to reduced detection accuracy. In this study, we evaluate the impact of thermal image preprocessing on improving defect segmentation in CFRP laminates inspected via pulsed thermography. Polynomial approximations and first- and second-order derivatives were applied to refine thermographic signals, enhancing defect visibility and SNR. The U-Net architecture was used to assess segmentation performance on datasets with and without preprocessing. The results demonstrated that preprocessing significantly improved defect detection, achieving an Intersection over Union (IoU) of 95% and an F1-Score of 99%, outperforming approaches without preprocessing. These findings emphasize the importance of preprocessing in enhancing segmentation accuracy and reliability, highlighting its potential for advancing non-destructive testing techniques across various industries. Full article
Show Figures

Figure 1

19 pages, 19125 KB  
Article
Automatic Segmentation of Gas Metal Arc Welding for Cleaner Productions
by Erwin M. Davila-Iniesta, José A. López-Islas, Yenny Villuendas-Rey and Oscar Camacho-Nieto
Appl. Sci. 2025, 15(6), 3280; https://doi.org/10.3390/app15063280 - 17 Mar 2025
Cited by 3 | Viewed by 1718
Abstract
In the industry, the robotic gas metal arc welding (GMAW) process has a huge range of applications, including in the automotive sector, construction companies, the shipping industry, and many more. Automatic quality inspection in robotic welding is crucial because it ensures the uniformity, [...] Read more.
In the industry, the robotic gas metal arc welding (GMAW) process has a huge range of applications, including in the automotive sector, construction companies, the shipping industry, and many more. Automatic quality inspection in robotic welding is crucial because it ensures the uniformity, strength, and safety of welded joints without the need for constant human intervention. Detecting defects in real time prevents defective products from reaching advanced production stages, reducing reprocessing costs. In addition, the use of materials is optimized by avoiding defective welds that require rework, contributing to cleaner production. This paper presents a novel dataset of robot GMAW images for experimental purposes, including human-expert segmentation and human knowledge labeling regarding the different errors that may appear in welding. In addition, it tests an automatic segmentation approach for robot GMAW quality assessment. The results presented confirm that automatic segmentation is comparable to human segmentation, guaranteeing a correct welding quality assessment to provide feedback on the robot welding process. Full article
(This article belongs to the Special Issue Sustainable Environmental Engineering)
Show Figures

Figure 1

Back to TopTop