Topic Editors

School of Electronics Engineering, IT College, Kyungpook National University, Daegu 41566, Republic of Korea
Department of Computer Science, University of Alicante, 03690 Alicante, Spain

Deep Visual Recognition: Methods, and Applications

Abstract submission deadline
closed (30 August 2026)
Manuscript submission deadline
30 October 2026
Viewed by
4380

Topic Information

Dear Colleagues,

Deep Visual Recognition focuses on enabling machines to automatically perceive and understand visual information from images and videos. Powered by deep learning, particularly Convolutional Neural Networks (CNNs) and more recently Vision Transformers (ViTs), this field has significantly advanced beyond traditional computer vision approaches that relied on hand-crafted features. By learning hierarchical and semantic representations directly from data, deep models achieve robust performance in complex visual recognition tasks. By learning hierarchical and semantic representations directly from data, deep models achieve robust performance on complex visual tasks, evolving from pattern recognition toward knowledge-driven and explainable intelligence that transforms visual data into interpretable, actionable insights for trustworthy decision-making.

Core methods in deep visual recognition address fundamental tasks such as image classification, object detection, semantic and instance segmentation, as well as face and human action recognition. Large-scale datasets (e.g., ImageNet, COCO) and high-performance computing have driven rapid improvements, while recent learning strategies—including self-supervised learning, transfer learning, and few-shot learning—have reduced dependence on extensive labeled data. Multimodal and foundation models further enhance generalization across diverse visual domains.

Deep visual recognition has become a key enabling technology in real-world applications such as autonomous driving, medical image analysis, intelligent surveillance, robotics, and industrial automation. Despite its success, challenges remain in handling domain shifts, occlusions, lighting variations, and real-time constraints. Future research increasingly emphasizes robustness, interpretability, and computational efficiency, positioning deep visual recognition as a central component of intelligent perception systems.

Prof. Dr. Min Young Kim
Dr. Francisco Gomez-Donoso
Topic Editors

Keywords

  • deep visual recognition
  • convolutional neural networks (CNNs)
  • vision transformers (ViTs)
  • object detection
  • semantic segmentation
  • self-supervised learning
  • multimodal learning
  • autonomous systems

Participating Journals

Journal Name Impact Factor CiteScore Launched Year First Decision (median) APC
AI
ai
6.5 7.3 2020 20.4 Days CHF 1800 Submit
Electronics
electronics
2.9 7.0 2012 14.8 Days CHF 2400 Submit
Machine Learning and Knowledge Extraction
make
8.4 12.7 2019 18.7 Days CHF 1800 Submit
Robotics
robotics
3.6 7.4 2012 20 Days CHF 1800 Submit
Sensors
sensors
4.0 9.4 2001 17.8 Days CHF 2600 Submit

Preprints.org is a multidisciplinary platform offering a preprint service designed to facilitate the early sharing of your research. It supports and empowers your research journey from the very beginning.

MDPI Topics is collaborating with Preprints.org and has established a direct connection between MDPI journals and the platform. Authors are encouraged to take advantage of this opportunity by posting their preprints at Preprints.org prior to publication:

  1. Share your research immediately: disseminate your ideas prior to publication and establish priority for your work.
  2. Safeguard your intellectual contribution: Protect your ideas with a time-stamped preprint that serves as proof of your research timeline.
  3. Boost visibility and impact: Increase the reach and influence of your research by making it accessible to a global audience.
  4. Gain early feedback: Receive valuable input and insights from peers before submitting to a journal.
  5. Ensure broad indexing: Web of Science (Preprint Citation Index), Google Scholar, Crossref, SHARE, PrePubMed, Scilit and Europe PMC.

Published Papers (5 papers)

Order results
Result details
Journals
Select all
Export citation of selected articles as:
22 pages, 35097 KB  
Article
Spot-Weld Defect Detection with YOLOv8n Integrating Multi-Receptive-Field Attention and Structural Re-Parameterization
by Yuxuan Zhou, Shudong Zhuang, Ao Sheng, Yizheng Ge, Jiarui Zhu, Zhizhou Wang, Yuxian Lei and Xinyan Cao
AI 2026, 7(9), 379; https://doi.org/10.3390/ai7090379 - 19 Sep 2026
Viewed by 21
Abstract
The reliable detection of spot-weld defects in automotive structural components is challenged by large variations in defect scale, severe background interference and limited detection accuracy. Here, we propose YOLOv8-RFA-iEMA-RH, an improved YOLOv8n-based detector for spot-weld defects. A receptive field attention convolution module (RFACM) [...] Read more.
The reliable detection of spot-weld defects in automotive structural components is challenged by large variations in defect scale, severe background interference and limited detection accuracy. Here, we propose YOLOv8-RFA-iEMA-RH, an improved YOLOv8n-based detector for spot-weld defects. A receptive field attention convolution module (RFACM) is introduced into the backbone to strengthen local texture representation through multi-receptive-field feature modelling. An improved Efficient Multi-scale Attention module (iEMA) is incorporated into the neck to enhance global context modelling and suppress background interference. In addition, a structurally re-parameterized RepHead is integrated into the detection head to enhance feature learning during training while maintaining a simplified single-branch structure for inference. On the self-built spot-weld defect dataset, the proposed model achieves 92.7% Recall, 91.2% F1, 97.9% mAP@0.5 and 71.9% mAP@0.5:0.95, improving on the YOLOv8n baseline by 2.3, 1.3, 2.5 and 3.1 percentage points, respectively. Cross-dataset evaluation on NEU-DET further yields 78.8% mAP@0.5 and 48.7% mAP@0.5:0.95. These results demonstrate improved detection accuracy and cross-dataset adaptability; actual inference speed and memory consumption require further validation on specific deployment hardware. Full article
(This article belongs to the Topic Deep Visual Recognition: Methods, and Applications)
Show Figures

Figure 1

14 pages, 8179 KB  
Article
Effects of Cervical Motion Restriction on Wheelchair Propulsion Velocity and Upper-Limb Kinematics in Healthy Individuals
by Tomoe Sobu, Yusuke Sekiguchi, Ziheng Wang, Keita Honda, Kaoru Kuroki, Naoki Suzuki, Ryoichi Nagatomi and Satoru Ebihara
Sensors 2026, 26(18), 5829; https://doi.org/10.3390/s26185829 - 14 Sep 2026
Viewed by 248
Abstract
Forward head flexion during wheelchair propulsion has been suggested to compensate for reduced trunk muscle strength in individuals with spinal cord injury, but the influence of restricting head–neck movement on propulsion remains unclear. This study examined wheelchair propulsion under cervical fixation and non-fixation [...] Read more.
Forward head flexion during wheelchair propulsion has been suggested to compensate for reduced trunk muscle strength in individuals with spinal cord injury, but the influence of restricting head–neck movement on propulsion remains unclear. This study examined wheelchair propulsion under cervical fixation and non-fixation conditions. Fourteen healthy young men (20.7 ± 0.5 years) performed 5 m wheelchair propulsion under both conditions. Propulsion was recorded with a smartphone camera, and coordinates of the left ear, shoulder, elbow, and wrist were extracted using You Only Look Once version 7 (YOLO v7) to calculate upper-arm segment and elbow angles and angular accelerations. Analysis focused on the first three propulsion cycles after movement onset. Propulsion velocity showed a significant main effect of measurement point, but neither the main effect of cervical fixation nor the condition-by-point interaction was significant. Velocity increment differed significantly among measurement intervals, whereas the condition effect and interaction were not significant. Cervical fixation was associated with a smaller minimum upper-arm segment angle and greater upper-arm segment excursion, maximum elbow angle, and elbow range of motion. A significant condition-by-cycle interaction was found for minimum elbow angle, with a smaller angle under fixation during the first cycle. Minimum upper-arm segment angular acceleration in the flexion direction also differed among cycles. Thus, cervical fixation was associated with changes in upper-limb kinematics, but a statistically significant change in propulsion velocity was not demonstrated. These preliminary findings provide a basis for future studies in wheelchair users and athletes. Full article
(This article belongs to the Topic Deep Visual Recognition: Methods, and Applications)
Show Figures

Figure 1

34 pages, 42016 KB  
Article
Interactive Playback Visualizer to Analyze Joint-Angle Co-Modulation with a Wavelet Approach: Application to Pose-Voice Relationships During Spontaneous Conversation
by Miguel A. Zamora-Ursulo, Amira Flores and Elias Manjarrez
Mach. Learn. Knowl. Extr. 2026, 8(9), 265; https://doi.org/10.3390/make8090265 - 28 Aug 2026
Viewed by 415
Abstract
Deep visual recognition can turn ordinary video into interpretable motor knowledge, yet coordination among the joints of a single body during social interaction remains largely unexplored. We present an interactive playback visualizer that couples markerless pose estimation with the cross-wavelet transform to quantify [...] Read more.
Deep visual recognition can turn ordinary video into interpretable motor knowledge, yet coordination among the joints of a single body during social interaction remains largely unexplored. We present an interactive playback visualizer that couples markerless pose estimation with the cross-wavelet transform to quantify amplitude co-modulation between all joint pairs. Amplitude co-modulation proved anatomically structured: bilateral homologous pairs exceeded cross-limb and head–body pairs even among pairs sharing no keypoint, where correlated tracking error cannot produce it. Within-limb pairs also scored high but share keypoints, so their elevation is confounded with measurement error and not treated as established. Co-modulation concentrated at low postural frequencies and showed no systematic temporal trend; each profile remained temporally consistent within the session, and equivalence to a flat trend was not established. Vocal activity correlated with upper-limb movement at gesture frequencies. Each participant’s 78-dimensional profile was individually distinctive: split-half identification reached 28.6% (95% CI 19.3–40.1%) against 1.4% chance, demonstrating within-session identifiability rather than a cross-session trait. No sex differences were detected in overall or category-level co-modulation, and equivalence was not established for any measure; women exceeded men in the micro-movement band, the narrowest and most attenuated by preprocessing, so that difference is reported but not interpreted. Full article
(This article belongs to the Topic Deep Visual Recognition: Methods, and Applications)
Show Figures

Figure 1

27 pages, 7514 KB  
Article
SASA-CLIP: Structure-Aware Alignment with a Gaussian Prior for Fine-Grained Video Action Recognition
by Xiaowei Han, Wenbao Si, Honghui Zhang, Maolin Yang, Lin Ma and Yibo Feng
Sensors 2026, 26(15), 4865; https://doi.org/10.3390/s26154865 - 2 Aug 2026
Viewed by 535
Abstract
Fine-grained video action recognition remains challenging because action categories often differ only in subtle inter-class variations and complex temporal dynamics. Recent Contrastive Language–Image Pre-training (CLIP)-based extensions perform well on general action recognition, but they typically rely on early global pooling of video features. [...] Read more.
Fine-grained video action recognition remains challenging because action categories often differ only in subtle inter-class variations and complex temporal dynamics. Recent Contrastive Language–Image Pre-training (CLIP)-based extensions perform well on general action recognition, but they typically rely on early global pooling of video features. Such coarse representations discard the fine temporal cues that distinguish subtle actions, causing a granularity mismatch in cross-modal alignment. To address this, we propose Structure-Aware Semantic-Adaptive (SASA)-CLIP, a framework for multi-granular cross-modal alignment. SASA-CLIP adopts a dual-branch design: a coarse-grained branch captures the global context, while a fine-grained branch matches descriptors against individual frames before aggregation, rather than pooling features early. To keep this alignment temporally coherent, we introduce a Gaussian prior as a temporal structural constraint, encoding the inductive bias of local temporal continuity into the attention matrix to guide an ordered alignment of key action segments along the temporal axis. On Kinetics-400 (ViT-B/32), SASA-CLIP reaches a Top-1 accuracy of 81.37%, improving over the X-CLIP baseline by 0.97%; on HMDB-51 and UCF-101 (ViT-B/16), it reaches 74.0% and 96.81%, improving by 3.25% and 2.61%, respectively. It also transfers to the zero-shot setting, improving over the baseline on HMDB-51 and UCF-101. These results show that combining multi-granular representations with a temporal structural prior benefits fine-grained recognition, suggesting that SASA-CLIP is a practical option for real-world visual sensing applications such as intelligent surveillance and wearable activity monitoring. Full article
(This article belongs to the Topic Deep Visual Recognition: Methods, and Applications)
Show Figures

Figure 1

34 pages, 36077 KB  
Article
Modular Multi-Attribute Vehicle Analysis by Color, License Plate, Make and Sub-Model Using YOLO and OCR: A Benchmark Across YOLO Versions
by Cristian Japhet Islas-Yañez, Viridiana Hernández-Herrera and Moisés Márquez-Olivera
Sensors 2026, 26(9), 2785; https://doi.org/10.3390/s26092785 - 29 Apr 2026
Viewed by 1583
Abstract
We present a modular multi-attribute vehicle analysis pipeline that integrates YOLO-based models and an OCR engine into a single workflow. The system detects vehicles, classifies color, recognizes make and sub-model, detects license plates, and extracts plate characters to generate a structured vehicle record. [...] Read more.
We present a modular multi-attribute vehicle analysis pipeline that integrates YOLO-based models and an OCR engine into a single workflow. The system detects vehicles, classifies color, recognizes make and sub-model, detects license plates, and extracts plate characters to generate a structured vehicle record. Vehicle detection is reported with standard metrics (precision, recall, and mAP@0.5), while license plate detection is reported at IoU = 0.3 to reflect the small-object nature of plates and downstream OCR usability. Among the evaluated versions, YOLOv8 provides the most balanced overall performance across modules, while maintaining real-time-equivalent throughput of approximately 18–22 FPS for the full pipeline on recorded traffic videos, depending on scene complexity. We emphasize module-level evaluation and runtime benchmarking; instance-level end-to-end identification across unique vehicles is defined as future work once track-based ground truth becomes available. Full article
(This article belongs to the Topic Deep Visual Recognition: Methods, and Applications)
Show Figures

Figure 1

Back to TopTop