electronics-logo

Journal Browser

► Journal Browser

Advances in Deep Learning for Open-World Computer Vision and Pattern Recognition

A Special Issue of Electronics (ISSN 2079-9292) belonging to the section "Artificial Intelligence".

Deadline for manuscript submissions: 15 November 2026 | Viewed by 2591

Editors


E-Mail Website
Guest Editor
School of Software, Shandong University, Jinan 250101, China
Interests: computer vision; multimedia computing and information retrieval; explainable AI

E-Mail Website
Guest Editor
School of Electronics and Information Engineering, Harbin Institute of Technology, Harbin 150001, China
Interests: computer vision; embedded intelligent computing; perception and decision-making for unmanned systems

E-Mail Website
Guest Editor
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
Interests: computer vision; visual salient object detection; intelligent perception and deep learning; visual navigation

Special Issue Information

Dear Colleagues,

In recent years, deep learning has significantly advanced the fields of computer vision and pattern recognition, and its applications have been widely integrated into various aspects of daily life and industrial production. While encouraging progress has been achieved in relatively simple or closed scenarios, performance in open-world environments remains unsatisfactory. Challenges stem not only from the openness of visual scenes (such as illumination variations, multi-scale objects, rainy or foggy conditions, and occlusion), but also from the openness of training and learning paradigms, including few-shot and zero-shot conditions where annotated data is scarce or unavailable. These challenges underscore the urgent need to investigate advanced deep learning techniques, encompassing diverse neural architectures (such as convolutional networks, Transformers, graph convolutional networks, and Mamba) as well as innovative learning strategies (such as few-shot or zero-shot learning), to enhance the robustness, adaptability, and generalization capability of computer vision and pattern recognition systems in open-world settings.

We are pleased to invite you to contribute to this Special Issue on “Advances in Deep Learning for Open-World Computer Vision and Pattern Recognition”. The aim of this Special Issue is to bring together cutting-edge research efforts that leverage the latest advances in deep learning to tackle open-world challenges in computer vision and pattern recognition.

This Special Issue seeks to provide a platform for researchers and practitioners to share original contributions, novel methodologies, and comprehensive reviews that address both theoretical and practical aspects. Submissions exploring innovative learning paradigms, robust model architectures, and application-driven solutions are highly encouraged.

In this Special Issue, original research articles and reviews are welcome. Research areas may include (but are not limited to) the following:

  • Multimodal computer vision and pattern recognition;
  • Open-world image recognition and understanding (including detection, classification, and segmentation, and enhancement);
  • Advanced neural network architectures for visual representation;
  • Few-shot, zero-shot, and other data-efficient learning strategies;
  • Generative and self-supervised methods for robust visual understanding;
  • Novel benchmarks, datasets, applications, and evaluation protocols for open-world vision tasks.

We look forward to receiving your valuable contributions.

Dr. Mingzhu Xu
Prof. Dr. Bing Liu
Dr. Lina Gao
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • deep learning
  • open-world computer vision
  • pattern recognition
  • multimodal learning
  • neural network architectures
  • few-shot learning
  • zero-shot learning
  • self-supervised learning
  • generative models
  • benchmarks and evaluation protocols

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (3 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

Jump to: Review

29 pages, 4146 KB  
Article
Anchor-Guided Discriminative Semantic Expansion for Point-Supervised Video Moment Localization
by Zhaoliang Zhou, Longqiang Pang, Zhen Li, Zhongsheng He, Jiqian Zhang, Jiecai Zheng and Xueqing Li
Electronics 2026, 15(14), 3011; https://doi.org/10.3390/electronics15143011 - 9 Jul 2026
Viewed by 416
Abstract
With the rapid growth of real-world untrimmed videos, video moment localization (VML) has become a fundamental task in query-guided video understanding, aiming to identify the temporal moment that semantically matches a natural language query. Although fully supervised methods have achieved promising performance, their [...] Read more.
With the rapid growth of real-world untrimmed videos, video moment localization (VML) has become a fundamental task in query-guided video understanding, aiming to identify the temporal moment that semantically matches a natural language query. Although fully supervised methods have achieved promising performance, their reliance on precise start–end annotations makes them difficult to scale to open-domain video scenarios, where visual contents are diverse, distracting contexts are common, and annotation resources are limited. Point-supervised VML provides a more data-efficient alternative by requiring only a single annotated point inside the target moment. However, such sparse supervision makes it challenging to infer the complete query-relevant interval and to suppress semantically similar distractors. To address these challenges, we propose an anchor-guided discriminative semantic expansion (ADSE) framework. ADSE treats the annotated point as a reliable semantic anchor, learns anchor-centered cross-modal alignment to generate a temporal relevance curve, and adaptively expands the anchor into a coherent target moment by integrating query relevance and temporal semantic continuity. Meanwhile, an anchor-guided discriminative learning strategy mines high-confidence anchor-excluded intervals as hard negatives, and an inside–outside separation objective further distinguishes target moments from surrounding contexts. Extensive experiments on public benchmarks demonstrate the effectiveness of ADSE and show consistent improvements over existing point-supervised methods under sparse point-level supervision. Full article
►▼ Show Figures

Figure 1

15 pages, 978 KB  
Article
SpectTrans: Joint Spectral–Temporal Modeling for Polyphonic Piano Transcription via Spectral Gating Networks
by Rui Cao, Yan Liang, Lei Feng and Yuanzi Li
Electronics 2026, 15(3), 665; https://doi.org/10.3390/electronics15030665 - 3 Feb 2026
Viewed by 881
Abstract
Automatic Music Transcription (AMT) plays a fundamental role in Music Information Retrieval (MIR) by converting raw audio signals into symbolic representations such as MIDI or musical scores. Despite advances in deep learning, accurately transcribing piano performances remains challenging due to dense polyphony, wide [...] Read more.
Automatic Music Transcription (AMT) plays a fundamental role in Music Information Retrieval (MIR) by converting raw audio signals into symbolic representations such as MIDI or musical scores. Despite advances in deep learning, accurately transcribing piano performances remains challenging due to dense polyphony, wide dynamic range, sustain pedal effects, and harmonic interactions between simultaneous notes. Existing approaches using convolutional and recurrent architectures, or autoregressive models, often fail to capture long-range temporal dependencies and global harmonic structures, while conventional Vision Transformers overlook the anisotropic characteristics of audio spectrograms, leading to harmonic neglect. In this work, we propose SpectTrans, a novel piano transcription framework that integrates a Spectral Gating Network with a multi-head self-attention Transformer to jointly model spectral and temporal dependencies. Latent CNN features are projected into the frequency domain via a Real Fast Fourier Transform, enabling adaptive filtering of overlapping harmonics and suppression of non-stationary noise, while deeper layers capture long-term melodic and chordal relationships. Experimental evaluation on polyphonic piano datasets demonstrates that this architecture produces acoustically coherent representations, improving the robustness and precision of transcription under complex performance conditions. These results suggest that combining frequency-domain refinement with global temporal modeling provides an effective strategy for high-fidelity AMT. Full article
►▼ Show Figures

Figure 1

Review

Jump to: Research

37 pages, 35970 KB  
Review
A Survey on Action Recognition: Multimodal Approaches, Ethical Considerations, and Feedback Mechanisms
by Bilal AbdulRahman, Zhigang Zhu and Alison Conway
Electronics 2026, 15(14), 3139; https://doi.org/10.3390/electronics15143139 - 16 Jul 2026
Viewed by 529
Abstract
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of [...] Read more.
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of human actions holds immense potential to improve system design, predictive modeling, data-informed decision-making, and real-time operational improvements across a wide variety of domains. Some examples of applications range from surveillance and real-time management of public spaces and infrastructure systems, to development of predictive modeling and robotic systems for individualized healthcare interventions, to implementing effective human–computer interaction in both professional and recreational settings. This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance. We delve into the evolution of action recognition, from early feature-based approaches to the deep learning revolution, emphasizing how the incorporation of multiple sensory modalities—such as visual, audio, and depth data as well as other cues—has advanced the field. Furthermore, we examine the ethical challenges associated with deploying these technologies in the public domain, particularly regarding privacy, bias, and societal impact, and discuss the need for responsible development and regulation. The third focus of the paper is the use of top-down and bottom-up feedback mechanisms within deep learning architectures, exploring how these strategies can mimic human cognitive processes to improve accuracy and reliability in action recognition systems. By identifying current gaps and proposing future research directions, this paper aims to inspire continued innovation in this dynamic and impactful field for intelligent systems. Full article
►▼ Show Figures

Figure 1

Back to TopTop