Advances of Artificial Intelligence and Vision Applications: Third Edition

A special issue of Electronics (ISSN 2079-9292). This special issue belongs to the section "Artificial Intelligence".

Deadline for manuscript submissions: 16 October 2026 | Viewed by 3294

Editors

Special Issue Information

Dear Colleagues,

We are pleased to announce the third volume of this Special Issue of Electronics, “Advances of Artificial Intelligence and Vision Applications: Third Edition”, following its prior success. Artificial intelligence technologies represented by deep learning and convolutional neural networks have greatly promoted the research and development of computer vision in the last decade. Simultaneously, advances in software and hardware also enable engineers to implement their elaborated computer vision algorithms onto powerful platforms. These advancements have enabled computer vision to attain enormous success across every aspect of modern society, including agriculture, retail, insurance, manufacturing, logistics, smart city, healthcare, pharmaceutical, construction, etc. The performance of an AI-based computer vision system is still constrained by the quality and quantity of training data and the hardware platforms' computing power and processing speed. This Special Issue aims to compile the advances and contributions of related research to the design, optimization, and implementation of artificial intelligence and computer vision applications.

General topics covered in this Special Issue include, but are not limited to, the following:

  • Image interpretation;
  • Object recognition and tracking;
  • Shape analysis, monitoring, and surveillance;
  • Biologically inspired computer vision;
  • Motion analysis;
  • Document image understanding;
  • Face and gesture recognition;
  • Vision-based human–computer interaction;
  • Human activity and behavior understanding;
  • Emotion recognition.

Dr. Dong Zhang
Prof. Dr. Dah-Jye Lee
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • artificial intelligence
  • computer vision
  • deep learning
  • convolutional neural networks
  • affective computing

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Related Special Issue

Published Papers (6 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

24 pages, 16370 KB  
Article
Unifying Inconsistent Emotion Labeling Criteria Across Datasets via Prototype-Guided Multimodal Alignment for Facial Expression Recognition
by Junjie Liu, Yufei Xie, Dong Zhang and Dah-Jye Lee
Electronics 2026, 15(15), 3279; https://doi.org/10.3390/electronics15153279 (registering DOI) - 25 Jul 2026
Abstract
Facial expression recognition (FER) plays an important role in human–computer interaction and affective computing. Although combining multiple FER datasets in joint training can potentially improve model generalization, it is hindered by inconsistent emotion labeling criteria across datasets. To address this issue, we propose [...] Read more.
Facial expression recognition (FER) plays an important role in human–computer interaction and affective computing. Although combining multiple FER datasets in joint training can potentially improve model generalization, it is hindered by inconsistent emotion labeling criteria across datasets. To address this issue, we propose a Prototype-Guided Multimodal Alignment Joint Training framework for multi-dataset FER. The core idea is to leverage the image–text alignment knowledge learned by Vision–Language Models from the target dataset as a unified emotion labeling criterion across datasets, while deriving prototypes from visual features to serve as emotion anchors for reliable feature alignment in the latent space. Based on the consistency among semantic predictions, prototype-distance predictions, and auxiliary labels, semantically consistent auxiliary samples are selected for joint training under a unified labeling criterion. Extensive experiments show that the proposed framework, with the Amending Representation Module as the backbone network, achieves 93.74% accuracy on the RAF-DB dataset and 98.31% on the CAER-S dataset, attaining state-of-the-art performance. The experimental results demonstrate that establishing semantically consistent labeling criteria across datasets is an effective strategy for multi-dataset FER learning. Full article
Show Figures

Figure 1

27 pages, 6205 KB  
Article
Low-Latency Machine Vision Based on a Neuromorphic Vision Sensor
by Paul K. J. Park, Junseok Kim, Juhyun Ko and Yeoungjin Chang
Electronics 2026, 15(13), 2828; https://doi.org/10.3390/electronics15132828 - 27 Jun 2026
Viewed by 429
Abstract
Low-latency visual perception is essential for interactive machine vision on edge AI devices, but conventional frame-based image sensors impose frame period delays and generate dense image data that increase memory bandwidth and processing latency. Although Dynamic Vision Sensors (DVSs) are known to provide [...] Read more.
Low-latency visual perception is essential for interactive machine vision on edge AI devices, but conventional frame-based image sensors impose frame period delays and generate dense image data that increase memory bandwidth and processing latency. Although Dynamic Vision Sensors (DVSs) are known to provide low latency, sparse output, and high dynamic range, these sensor-level properties do not automatically translate into practical application-level latency reduction on resource-constrained edge platforms. This paper presents a latency-driven sensing algorithm co-design approach for DVS-based low-latency machine vision. The main objective is to connect DVS sensor-level characteristics, event representations, task-dependent processing flows, and measured response times on mobile application processors. We first analyze latency requirements for three representative edge AI applications (i.e., person detection, gesture recognition, and Simultaneous Localization and Mapping (SLAM)), which correspond to different latency regimes and processing structures. We then describe the DVS operating principle, pixel-level event latency, and readout latency, showing how asynchronous event generation reduces sensing delay and suppresses redundant static background information before algorithmic processing. In contrast to prior event camera studies that mainly optimize a single task or a specific event representation, this work evaluates three task-specific event processing systems on mobile processors. Person detection achieves 92 ms processing latency on Exynos 7570, gesture recognition based on event-driven 4-DoF motion estimation achieves 20 ms latency on Exynos 5422, and SLAM achieves 15.9 ms latency on Snapdragon 845. These results satisfy the practical latency targets of the corresponding applications and demonstrate that DVS-based sensing can provide not only sensor-level speed advantages but also system-level latency benefits for AIoT, mobile, robotics, and AR/VR machine vision systems. Full article
Show Figures

Figure 1

30 pages, 6273 KB  
Article
Benchmarking Large Language Model Inference on Limited-Resource Edge Systems
by Henrikas Giedra, Dalius Matuzevičius, Tomyslav Sledevič, Giga Shubitidze and Artūras Serackis
Electronics 2026, 15(11), 2451; https://doi.org/10.3390/electronics15112451 - 3 Jun 2026
Viewed by 944
Abstract
Large language models (LLMs) are increasingly considered for deployment on edge and limited-resource systems, where local inference can reduce latency, improve privacy, and decrease dependence on cloud infrastructure. While prior studies have evaluated either task accuracy or hardware efficiency in isolation, few benchmarks [...] Read more.
Large language models (LLMs) are increasingly considered for deployment on edge and limited-resource systems, where local inference can reduce latency, improve privacy, and decrease dependence on cloud infrastructure. While prior studies have evaluated either task accuracy or hardware efficiency in isolation, few benchmarks combine generation-based response-quality evaluation with real-device power measurements on a representative limited-resource platform. This study addresses that gap by benchmarking twelve compact and mid-scale open-weight LLMs (sub-1B to 8B parameters), evaluating generation-based accuracy on a desktop platform and measuring deployment efficiency—throughput, power consumption, and energy use—on an NVIDIA Jetson Orin Nano Super; the accuracy–efficiency trade-off is thus established at the model-configuration level. Unlike prior Jetson-based evaluations relying solely on internal telemetry, this work pairs generation-compatible lm-eval accuracy tasks with a dual power-measurement setup that combines internal tegrastats rail readings with external board-level input power measured using a digital multimeter and explicitly compares GPU-accelerated and CPU-only inference modes. GPU-accelerated inference provided a clear advantage, increasing median throughput from 7.12 to 18.13 tok/s and improving external-meter energy efficiency from 0.453 to 0.823 tok/J, despite higher mean input power. Sub-1B models offered the best throughput and energy efficiency, whereas 7–8B models achieved stronger accuracy at a substantially higher energy cost per generated token. These results demonstrate that edge LLM deployment requires multi-objective evaluation balancing accuracy, throughput, power consumption, and energy efficiency. Full article
Show Figures

Figure 1

22 pages, 19167 KB  
Article
RGB Ensemble Strategies for Unsupervised Industrial Anomaly Detection on the AutoVI Dataset
by Sergio Villanueva López, Emilio Soria-Olivas and Manuel Sánchez-Montañés
Electronics 2026, 15(10), 2077; https://doi.org/10.3390/electronics15102077 - 13 May 2026
Viewed by 393
Abstract
Real automotive inspection lines need robust defect detection under cluttered backgrounds, fluctuating illumination, and operator-introduced clutter, conditions under which fully supervised pipelines are rarely feasible because defective samples are scarce and heterogeneous. We address this gap with a deployment-oriented study of unsupervised anomaly [...] Read more.
Real automotive inspection lines need robust defect detection under cluttered backgrounds, fluctuating illumination, and operator-introduced clutter, conditions under which fully supervised pipelines are rarely feasible because defective samples are scarce and heterogeneous. We address this gap with a deployment-oriented study of unsupervised anomaly detection (UAD) on AutoVI, a public automotive benchmark, and we go beyond running existing detectors in three ways. First, we establish unified RGB and pseudo-depth baselines for seven UAD models under a single calibration and evaluation policy that combines threshold-agnostic metrics (AUROC, AP), operational metrics (TPR at fixed TNR), and pixel-level sPRO/AUsPRO under a 5% false-positive budget. Second, we show that a plug-and-play late fusion of independently calibrated RGB detectors consistently recovers pixel-level localization that no single model achieves, with no extra training and no architectural change; effects are large (Cohen’s d>2 on every flagged improvement) and statistically significant across three seeds. Third, we report an actionable negative result: combining RGB with monocular pseudo-depth through the same fusion scheme degrades rather than improves localization, and we trace the failure to the relative nature of estimated depth interacting with a parameter-free aggregator. Ablations on the fusion operator, ensemble size, and calibration support these findings, and the released calibrated artifacts make the comparison reproducible on other MVTec-style benchmarks. Full article
Show Figures

Figure 1

15 pages, 913 KB  
Article
Task-Aware Preprocessing Selection for Underwater Sparse 3D Reconstruction via Lightweight Machine Learning Under Grouped Evaluation Protocol
by Ning Hu and Senhao Cao
Electronics 2026, 15(9), 1923; https://doi.org/10.3390/electronics15091923 - 1 May 2026
Viewed by 389
Abstract
Underwater image enhancement has been widely studied to improve visual quality; however, its impact on downstream geometric tasks such as sparse 3D reconstruction remains insufficiently understood. In particular, visually enhanced images do not necessarily lead to improved feature matching or reconstruction performance. This [...] Read more.
Underwater image enhancement has been widely studied to improve visual quality; however, its impact on downstream geometric tasks such as sparse 3D reconstruction remains insufficiently understood. In particular, visually enhanced images do not necessarily lead to improved feature matching or reconstruction performance. This work addresses the problem of selecting appropriate preprocessing strategies for underwater Structure-from-Motion (SfM) pipelines from a task-oriented perspective. We propose a lightweight machine-learning-based preprocessing selector that predicts reconstruction performance from image statistics and recommends suitable enhancement strategies for each input sequence. To ensure reliable evaluation, we introduce a grouped leave-one-parent-sequence-out protocol that avoids overlap-induced bias common in clip-wise splitting. Experiments are conducted on challenging underwater datasets derived from the Real-world Underwater Image Enhancement (RUIE) benchmark, with the primary comparison variable defined as the number of reconstructed sparse 3D points. Supporting geometric variables, including the number of registered images, mean track length, and mean reprojection error, are recorded for interpretation. Results show that preprocessing choices significantly affect reconstruction outcomes and that the optimal strategy is scene-dependent. The proposed selector consistently improved over raw input on the evaluated grouped subset and remained competitive with a strong fixed preprocessing baseline. The grouped leave-one-parent-sequence-out protocol is intended to reduce overlap-induced bias common in clip-wise splitting and to provide a more conservative estimate of generalization. This work highlights the importance of task-aware preprocessing and reliable evaluation in underwater vision systems, offering practical insights for deploying enhancement strategies in real-world 3D reconstruction pipelines. Full article
Show Figures

Figure 1

20 pages, 1281 KB  
Article
HGRN2-Based Personal Voice Activity Detection: A Lightweight Recurrent Framework for Inference and Training
by Tzu-Wei Wang, Tai-You Chen, Chien-Chia Chiu, Berlin Chen and Jeih-Weih Hung
Electronics 2026, 15(8), 1561; https://doi.org/10.3390/electronics15081561 - 8 Apr 2026
Viewed by 561
Abstract
This study presents HGRN2-based Flexible Dynamic Encoder Personal VAD (FDE-HGRN2), a recurrent framework for personal voice activity detection (PVAD). Building on the original LSTM-based FDE-RNN backbone, we replace all recurrent modules with the recently introduced HGRN2 gated linear RNN and adopt a cosine-annealing [...] Read more.
This study presents HGRN2-based Flexible Dynamic Encoder Personal VAD (FDE-HGRN2), a recurrent framework for personal voice activity detection (PVAD). Building on the original LSTM-based FDE-RNN backbone, we replace all recurrent modules with the recently introduced HGRN2 gated linear RNN and adopt a cosine-annealing learning rate schedule to improve both detection accuracy and efficiency. HGRN2 uses gated linear recurrence with non-parametric state expansion, enlarging the recurrent state without increasing the number of trainable parameters and enabling more expressive long-range temporal modeling than conventional LSTMs. We evaluate FDE-HGRN2 on a LibriSpeech-derived PVAD benchmark, where multi-speaker mixtures are constructed by concatenating one to three speakers per utterance and randomly designating a target speaker, following established PVAD data construction practices to ensure direct comparability with prior work. The system uses 40-dimensional Mel-filterbank features as acoustic inputs and conditions the detector on 256-dimensional d-vector embeddings extracted from a pretrained speaker verification network. Experimental results show that FDE-HGRN2 consistently outperforms the original FDE-RNN baseline and several state-of-the-art PVAD models in terms of mean Average Precision and frame-level accuracy, while reducing the parameter count of the recurrent backbone by roughly 15% and yielding substantially smaller models than many competing systems. These findings indicate that HGRN2 provides a more temporally expressive and parameter-efficient alternative to LSTM for PVAD, offering a favorable accuracy–efficiency trade-off for real-world, deployment-oriented personalized speech interfaces. Full article
Show Figures

Figure 1

Back to TopTop