electronics-logo

Journal Browser

Journal Browser

Image/Video Processing and Computer Vision

A Special Issue of Electronics (ISSN 2079-9292) belonging to the section "Computer Science & Engineering".

Deadline for manuscript submissions: 28 February 2027 | Viewed by 3761

Editor


E-Mail Website
Guest Editor
School of Automation, Wuhan University of Technology, Wuhan 430070, China
Interests: computer vision; medical image analysis; intelligent medical diagnosis; intelligent navigation; surgical robot system development

Special Issue Information

Dear Colleagues,

Image and video processing, together with computer vision, is advancing rapidly thanks to breakthroughs in deep learning, multimodal modeling, and real-time inference. These technologies enable machines to interpret visual content with high accuracy and drive transformative applications in autonomous systems, medical imaging, industrial inspection, augmented reality, and many other fields. In this Special Issue, we present cutting-edge research that bridges algorithmic innovation with the real-world impact of visual data understanding.

We invite you to submit original contributions to this Special Issue entitled Image/Video Processing and Computer Vision, with a focus on the following topics:

  1. Novel architectures and learning paradigms for image and video analysis;
  2. Techniques for low-level vision tasks such as denoising, super-resolution, enhancement, and restoration;
  3. High-level vision problems including object detection, segmentation, pose estimation, scene understanding, and 3D reconstruction;
  4. Multimodal fusion involving vision, language, audio, or sensor data;
  5. Real-world applications in healthcare, robotics, smart cities, manufacturing, and other fields.

Dr. Quan Zhou
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • computer vision
  • image processing
  • multimodal information fusion
  • intelligent navigation
  • smart healthcare

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

29 pages, 19630 KB  
Article
Single-Image 3D Mesh Reconstruction for Stylized Side-Face Characters via Prompt-Driven Multi-View Diffusion and Consistency Optimization
by Ke Zhang, Jiayi Lin, Zhixiang Zhang and Junghyun Heo
Electronics 2026, 15(13), 2963; https://doi.org/10.3390/electronics15132963 - 6 Jul 2026
Viewed by 795
Abstract
Single-image 3D reconstruction of stylized side-face characters remains challenging because profile-view inputs contain severe self-occlusion, missing frontal geometry, and stylized appearance cues that differ from the assumptions of generic reconstruction models. Because the unseen facial geometry cannot be uniquely determined from a single [...] Read more.
Single-image 3D reconstruction of stylized side-face characters remains challenging because profile-view inputs contain severe self-occlusion, missing frontal geometry, and stylized appearance cues that differ from the assumptions of generic reconstruction models. Because the unseen facial geometry cannot be uniquely determined from a single profile-view input, this study focuses on generating plausible and visually consistent 3D completions rather than uniquely recovering the unobserved geometry. When CRM is directly applied to stylized profile inputs, the outputs often exhibit unstable facial completion, local mesh collapse, UV misalignment, texture discontinuity, and other reconstruction artifacts. Rather than introducing a new reconstruction backbone, this study first diagnoses the task-specific limitations of CRM in this setting. We identify eight characteristic failure modes that occur when CRM is directly applied to stylized profile inputs and use this diagnosis to guide a retraining-free inference-time intervention strategy. The proposed strategy combines reconstruction-compatible auxiliary-view generation with failure-mode-oriented CRM refinement, including candidate verification, adaptive facial cropping, geometric stabilization, local detail enhancement, normal correction, UV repair, and texture continuity improvement. Experiments on a rendered stylized-character dataset and a cross-style adaptation set show that the proposed intervention improves frontal-view plausibility, mesh usability, texture continuity, and rendered appearance compared with direct reconstruction baselines. The seven-configuration progressive ablation and parameter sensitivity analyses further support the complementary role of the main intervention stages and the stability of the selected settings. These findings suggest that systematic failure-mode diagnosis, followed by task-specific inference-time intervention, provides a practical way to extend public image-to-3D models to stylized profile reconstruction scenarios, within the scope of the evaluated stylized-character datasets, while extreme viewpoints and highly abstract styles remain challenging. Full article
(This article belongs to the Special Issue Image/Video Processing and Computer Vision)
Show Figures

Figure 1

17 pages, 3855 KB  
Article
Learning Depth from Focus with Multi-Candidate Estimation and Proximal Refinement
by Muhammad Tariq Mahmood
Electronics 2026, 15(12), 2548; https://doi.org/10.3390/electronics15122548 - 9 Jun 2026
Viewed by 422
Abstract
In this paper, we propose a novel Depth from Focus (DFF) framework that formulates depth estimation as an energy minimization problem and unrolls the corresponding iterative optimization into a trainable neural architecture. Given a focal stack, a deep feature extractor constructs a learned [...] Read more.
In this paper, we propose a novel Depth from Focus (DFF) framework that formulates depth estimation as an energy minimization problem and unrolls the corresponding iterative optimization into a trainable neural architecture. Given a focal stack, a deep feature extractor constructs a learned focus volume that encodes defocus and structural cues. Based on this representation, multiple candidate depth maps are generated using a plane-based probabilistic formulation, while an attention mechanism adaptively assigns pixel-wise confidence weights to each candidate. The depth estimation is performed through an iterative refinement process, where each stage corresponds to a learned proximal update implemented via lightweight conditional networks. These updates incorporate focus consistency, adaptive step sizes, and learned regularization priors, enabling effective integration of physical imaging constraints with data-driven modeling. A final refinement module further enhances prediction accuracy by fusing the refined depth, focus volume features, and candidate hypotheses to estimate residual corrections. The entire framework is trained end-to-end, ensuring coherent optimization across all components. Experimental results demonstrate that the proposed method achieves improved robustness and accuracy, particularly in low-texture and noisy regions, while preserving interpretability through its unfolding-based design. Full article
(This article belongs to the Special Issue Image/Video Processing and Computer Vision)
Show Figures

Figure 1

22 pages, 7033 KB  
Article
WSNet: Person Re-Identification Based on Wavelet Convolution and Assisted by Image Generation at Inference Time
by Honggang Xie, Jinyang Huang, Xinxin Yi, Zhiwei Chen, Wei Xiong, Yuan Yao, Yongsheng Bai and Xiuyuan Meng
Electronics 2026, 15(8), 1645; https://doi.org/10.3390/electronics15081645 - 15 Apr 2026
Viewed by 690
Abstract
In pedestrian re-identification (ReID) tasks, existing models face dual challenges: first, surveillance cameras capture images at long distances with low resolution and blurriness; second, image data suffers from insufficient samples, limited poses, and cross-domain adaptation issues. To address these issues, we propose a [...] Read more.
In pedestrian re-identification (ReID) tasks, existing models face dual challenges: first, surveillance cameras capture images at long distances with low resolution and blurriness; second, image data suffers from insufficient samples, limited poses, and cross-domain adaptation issues. To address these issues, we propose a wavelet-convolution-based person re-identification framework assisted by a Stable Diffusion-based identity-preserving image generation module used only at inference time. This approach employs a dual-channel wavelet convolutional neural network for multi-scale feature extraction of pedestrian images, combined with cross-attention and gating mechanisms for dynamic data fusion. Additionally, we incorporate a pre-trained Pose2ID-based auxiliary generation branch that synthesizes identity-preserving pedestrian views with diverse poses under human keypoint guidance. These generated views are used only at inference time, where their WSNet features are fused with the feature of the original image to provide pose-complementary representation enhancement. Experiments on the Market-1501 and MSMT17 benchmark datasets demonstrate that our method achieves an mAP of 92.1% and a Rank-1 accuracy of 96.5% on Market-1501, and an mAP of 60.1% and a Rank-1 accuracy of 81.2% on MSMT17, with a WSNet backbone of 2.66 M parameters. Compared with the baseline models, the proposed method improves mAP by 5.1 and 7.6 percentage points on Market-1501 and MSMT17, respectively. Full article
(This article belongs to the Special Issue Image/Video Processing and Computer Vision)
Show Figures

Figure 1

31 pages, 9056 KB  
Article
Edge-Based Artificial Intelligence Analysis for Real-Time Content Classification and Knowledge Graph Construction of Movie Archives
by Peixuan Qi and Weidong Zhu
Electronics 2026, 15(5), 1011; https://doi.org/10.3390/electronics15051011 - 28 Feb 2026
Viewed by 829
Abstract
Movie archives still rely on manual cataloging and sparse metadata, limiting fine-grained retrieval, relationship tracing, and reuse under privacy-constrained edge settings. We propose EdgeCineTag-KG, an edge framework using a single video foundation encoder and knowledge-constrained multi-label learning to produce consistent labels and build [...] Read more.
Movie archives still rely on manual cataloging and sparse metadata, limiting fine-grained retrieval, relationship tracing, and reuse under privacy-constrained edge settings. We propose EdgeCineTag-KG, an edge framework using a single video foundation encoder and knowledge-constrained multi-label learning to produce consistent labels and build a queryable movie-archive knowledge graph. The objective jointly models label co-occurrence, mutual exclusion, hierarchy, and temporal consistency to reduce semantic contradictions and label jitter. For deployment, an uncertainty-driven adaptive computation strategy meets real-time constraints with controlled quality loss. Across MovieNet, Condensed Movies, Trailers12k, MMTF-14K, and TVQA, performance improves from 47.8 to 55.6 mAP and from 38.2 to 44.9 Macro-F1 on MovieNet, from 42.1 to 49.3 mAP on Condensed Movies, and from 71.2 to 75.4 mAP on Trailers12k. Knowledge graph quality also improves, with rule violation rate dropping from 6.8% to 2.4% and link prediction MRR rising from 0.248 to 0.312. Under INT8 adaptive inference, the system reaches 5.3 Clip-FPS, 182 ms P95 latency, and 1.9 GB peak memory. This combination improves consistency and retrieval usability without relying on multiple stacked foundation models. These results support reliable, interpretable, and edge-deployable movie archive understanding. Full article
(This article belongs to the Special Issue Image/Video Processing and Computer Vision)
Show Figures

Figure 1

Back to TopTop