Computational Intelligence, Computer Vision and Pattern Recognition

A Special Issue of Mathematics (ISSN 2227-7390) belonging to the section "E1: Mathematics and Computer Science".

Deadline for manuscript submissions: 15 September 2026 | Viewed by 4697

Editor


E-Mail Website
Guest Editor
1. School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
2. Guangdong Province Key Laboratory of Information Security Technology, Guangzhou, China
3. Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, Beijing, China
Interests: computer vision; 3D human video prediction and generation; multimodal video understanding; multimodal large models; ocean large models; trajectory prediction
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

This Special Issue titled “Computational Intelligence, Computer Vision and Pattern Recognition” aims to provide a comprehensive overview of the latest research and innovations in the application of computer vision and pattern recognition techniques. It seeks to highlight how these technologies are being used to address real-world challenges across various domains such as healthcare, autonomous systems, security, and robotics. We invite authors to submit original research that explores both theoretical advancements and practical applications, with a focus on areas such as image and video analysis, object detection, recognition systems, deep learning methods, and multimodal learning. The goal of this Special Issue is to offer a platform for researchers to present novel solutions, discuss emerging trends, and showcase the transformative potential of computer vision and pattern recognition in solving complex, real-world problems.

Dr. Jian-Fang Hu
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Mathematics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2600 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • computer vision
  • pattern recognition
  • deep learning
  • image analysis
  • object detection
  • recognition systems
  • multimodal learning
  • video analysis

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

36 pages, 1705 KB  
Article
EATMamba: Evolutionary Token-Refined Vision Mamba for Tomato Leaf Disease and Pest Classification
by Yingbiao Hu, Huinian Li, Yu He, Zhenfu Pan, Ningxia Chen, Chengcheng Yang and Wei Ke
Mathematics 2026, 14(15), 2813; https://doi.org/10.3390/math14152813 - 5 Aug 2026
Viewed by 263
Abstract
Accurate and efficient recognition of tomato leaf diseases and pests is essential for precision agriculture, yet practical deployment remains challenging due to complex backgrounds, domain shifts, and subtle inter-class visual differences. From a broader mathematical perspective, this task can be viewed as token-level [...] Read more.
Accurate and efficient recognition of tomato leaf diseases and pests is essential for precision agriculture, yet practical deployment remains challenging due to complex backgrounds, domain shifts, and subtle inter-class visual differences. From a broader mathematical perspective, this task can be viewed as token-level representation refinement under noise, ambiguity, and distribution shift. Recent state-space-model-based vision backbones provide favorable efficiency for high-resolution imagery but lack explicit mechanisms for adaptive feature refinement under noisy conditions. To address these issues, we propose EATMamba, an evolution-inspired Vision Mamba framework for tomato leaf disease and pest classification. Rather than being limited to a task-specific classifier, EATMamba is formulated as an evolution-inspired differentiable token-refinement mechanism designed to be compatible with state-space visual recognition backbones. EATMamba introduces two lightweight and fully differentiable modules—Evolutionary Crossover–Interaction and Knowledge-Guided Mutation–Selection—which perform token-level recombination and selective refinement to emphasize discriminative disease cues while suppressing irrelevant background information. These modules are inspired by crossover, mutation, and selection concepts, but are implemented as trainable differentiable operations over visual tokens, forming a generate–recombine–select style mechanism for representation refinement. The scope of this study is low-cost RGB-based visible-symptom disease and pest classification, rather than pre-symptomatic early disease detection. Extensive experiments on two complementary tomato datasets, including a controlled high-resolution dataset and an in-the-wild farm dataset, demonstrate that EATMamba consistently outperforms representative CNN-, Transformer-, and state-space-model-based baselines. Ablation studies and visualization analyses further confirm the complementary contributions of the proposed modules. Overall, EATMamba provides an effective and efficient framework for fine-grained plant disease recognition and illustrates how evolution-inspired principles can be incorporated into modern vision backbones for robust agricultural image analysis. Full article
(This article belongs to the Special Issue Computational Intelligence, Computer Vision and Pattern Recognition)
Show Figures

Figure 1

12 pages, 2417 KB  
Article
An Air Quality Forecasting Model Based on Second-Order Recurrent Neural Network
by Leqian Zhang, Yujian Li, Shiyao Zhou and Yabing Wang
Mathematics 2026, 14(11), 1833; https://doi.org/10.3390/math14111833 - 25 May 2026
Viewed by 316
Abstract
Air quality forecasting plays an important role in environmental protection and public health. Nevertheless, existing deep learning models remain limited in capturing long-term spatiotemporal dependencies. To overcome this limitation, a second-order recurrent neural network (SndRNN) is proposed in this study. By incorporating the [...] Read more.
Air quality forecasting plays an important role in environmental protection and public health. Nevertheless, existing deep learning models remain limited in capturing long-term spatiotemporal dependencies. To overcome this limitation, a second-order recurrent neural network (SndRNN) is proposed in this study. By incorporating the two historical states of the memory unit together with the current input into the update of the current memory state, the proposed model enhances historical information representation and improves long-term dependency modeling. Extensive experiments are conducted on six air quality datasets collected from multiple monitoring stations in China, where the proposed model is compared against five related baseline models. The results demonstrate that SndRNN achieves better MSE and RMSE values in 90% and 97% of the experimental cases, respectively, and overall outperforms traditional RNN-based approaches and other representative models. Full article
(This article belongs to the Special Issue Computational Intelligence, Computer Vision and Pattern Recognition)
Show Figures

Figure 1

18 pages, 4244 KB  
Article
Dual-Modal Contrastive Learning for Continual Generalized Category Discovery
by Wei Jin, Nannan Li, Chengcheng Yang, Huanqiang Hu and Kuo Li
Mathematics 2026, 14(2), 365; https://doi.org/10.3390/math14020365 - 21 Jan 2026
Viewed by 1002
Abstract
Continual Generalized Category Discovery (C-GCD) is an emerging research direction in Open-World Learning. The model aims to incrementally discover novel classes from unlabeled data while maintaining recognition of previously learned classes, without accessing historical samples. The absence of supervision signal in incremental sessions [...] Read more.
Continual Generalized Category Discovery (C-GCD) is an emerging research direction in Open-World Learning. The model aims to incrementally discover novel classes from unlabeled data while maintaining recognition of previously learned classes, without accessing historical samples. The absence of supervision signal in incremental sessions makes catastrophic forgetting more severe than in traditional incremental learning. Existing methods primarily enhance generalization through single-modality contrastive learning, overlooking the natural advantages of textual information. Visual features capture perceptual details such as shapes and textures, while textual information helps distinguish visually similar but semantically distinct categories, offering complementary benefits. However, directly obtaining category descriptions for unlabeled data in C-GCD is challenging. To address this, we introduce a conditional prompt learning mechanism to generate pseudo-prompts as textual information for unlabeled samples. Additionally, we propose a dual-modal contrastive learning strategy to enhance vision-text alignment and exploit CLIP’s multimodal potential. Extensive experiments on four benchmark datasets demonstrate that our method achieves competitive performance. We hope this work provides new insights for future research. Full article
(This article belongs to the Special Issue Computational Intelligence, Computer Vision and Pattern Recognition)
Show Figures

Figure 1

15 pages, 4930 KB  
Article
A Lightweight Hybrid CNN-ViT Network for Weed Recognition in Paddy Fields
by Tonglai Liu, Yixuan Wang, Chengcheng Yang, Youliu Zhang and Wanzhen Zhang
Mathematics 2025, 13(17), 2899; https://doi.org/10.3390/math13172899 - 8 Sep 2025
Cited by 4 | Viewed by 2201
Abstract
Accurate identification of weed species is a fundamental task for promoting efficient farmland management. Existing recognition approaches are typically based on either conventional Convolutional Neural Networks (CNNs) or the more recent Vision Transformers (ViTs). CNNs demonstrate strong capability in capturing local spatial patterns, [...] Read more.
Accurate identification of weed species is a fundamental task for promoting efficient farmland management. Existing recognition approaches are typically based on either conventional Convolutional Neural Networks (CNNs) or the more recent Vision Transformers (ViTs). CNNs demonstrate strong capability in capturing local spatial patterns, yet they are often limited in modeling long-range dependencies. In contrast, ViTs can effectively capture global contextual information through self-attention, but they may neglect fine-grained local features. These inherent shortcomings restrict the recognition performance of current models. To overcome these limitations, we propose a lightweight hybrid architecture, termed RepEfficientViT,which integrates convolutional operations with Transformer-based self-attention. This design enables the simultaneous aggregation of both local details and global dependencies. Furthermore, we employ a structural re-parameterization strategy to enhance the representational capacity of convolutional layers without introducing additional parameters or computational overhead. Experimental evaluations reveal that RepEfficientViT consistently surpasses state-of-the-art CNN and Transformer baselines. Specifically, the model achieves an accuracy of 94.77%, a precision of 94.75%, a recall of 94.93%, and an F1-score of 94.84%. In terms of efficiency, RepEfficientViT requires only 223.54 M FLOPs and 1.34 M parameters, while attaining an inference latency of merely 25.13 ms on CPU devices. These results demonstrate that the proposed model is well-suited for deployment in edge-computing scenarios subject to stringent computational and storage constraints. Full article
(This article belongs to the Special Issue Computational Intelligence, Computer Vision and Pattern Recognition)
Show Figures

Figure 1

Back to TopTop