sensors-logo

Journal Browser

Journal Browser

Advanced Pattern Recognition: Intelligent Sensing and Imaging

A Special Issue of Sensors (ISSN 1424-8220) belonging to the section "Sensing and Imaging".

Deadline for manuscript submissions: 20 November 2026 | Viewed by 7650

Editors


E-Mail Website
Guest Editor
School of Aeronautics and Astronautics, Zhejiang University, Hangzhou 310027, China
Interests: image/video processing and analysis; deep learning; data mining; information security
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
School of Aeronautics and Astronautics, Zhejiang University, Hangzhou 310027, China
Interests: machine self-learning and evolution; image and video analysis and processing; deep neural networks and machine behavior

E-Mail Website
Guest Editor Assistant
School of Information Science and Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China
Interests: image/video processing and analysis; signal processing; information security
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues, 

Deep learning and a new round of artificial intelligence development have greatly promoted the development of pattern recognition in computer vision and intelligent sensing, e.g., human action pattern recognition based on acceleration sensors has become an emerging research direction in the field of pattern recognition.

This Special Issue is oriented towards intelligent algorithms and technologies in pattern recognition and sensing fields. The aim is to share the latest theoretical and technological achievements in intelligent sensing and pattern recognition, and to encourage scientists to publish their experimental and theoretical results in these fields, mainly those based on deep learning. The related application areas include the following: advanced pattern recognition; image and video analysis and processing; intelligent sensors; intelligent video surveillance; intelligent visual inspection; and security and privacy problems in sensing.

This Special Issue warmly welcomes the submission of studies related to the following research topics: vision research under new imaging conditions; biologically inspired computer vision research; multi-sensor fusion 3D vision research; visual scene understanding under high dynamic complex scenes; small-sample target recognition and understanding; and complex behavior semantic understanding. Electronic files and software providing full details of calculation and experimental procedures can be deposited as Supplementary Material. 

We look forward to receiving your submissions. 

Prof. Dr. Zhe-Ming Lu
Dr. Yangming Zheng
Dr. Hao Luo
Prof. Dr. Junbao Li
Guest Editors

Prof. Dr. Yijia Zhang
Guest Editor Assistant

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Sensors is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2600 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • computer vision
  • pattern recognition
  • intelligent sensing
  • deep learning
  • image and video analysis and processing
  • intelligent sensors
  • intelligent video surveillance
  • intelligent visual inspection
  • security and privacy in sensing

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (7 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

28 pages, 4789 KB  
Article
Geometry-Constrained Multi-Frame Character Association for License Plate Recognition on Moving Cameras
by Ufuk Asil and İlker Yoncacı
Sensors 2026, 26(18), 5704; https://doi.org/10.3390/s26185704 - 8 Sep 2026
Abstract
Multi-frame fusion is standard for converting frame-by-frame license plate character detections into stable, reliable readings. Character Time-Series Matching (CTM), a leading approach, associates characters across frames using the Hungarian algorithm with a fixed Euclidean distance threshold and a translation-only motion model, reporting 96.7% [...] Read more.
Multi-frame fusion is standard for converting frame-by-frame license plate character detections into stable, reliable readings. Character Time-Series Matching (CTM), a leading approach, associates characters across frames using the Hungarian algorithm with a fixed Euclidean distance threshold and a translation-only motion model, reporting 96.7% accuracy on the UFPR-ALPR dataset. In this work, we demonstrate that this high performance is protocol-dependent: when ground-truth static plate crops and pre-segmented tracks are used, CTM performs strongly. However, in real-world scenarios involving moving cameras (such as drone-mounted cameras, helmet-mounted cameras, and mobile platforms) where inter-frame geometry changes dynamically, baseline multi-frame association frameworks that combine fixed spatial gates with unconstrained translation propagation fail. In these cases, temporal fusion provides no benefit and degrades plate recognition performance below the single-frame baseline. Indeed, under its own Intersection over Union (IoU) tracker, this literature method correctly reads only 15.1% of plates in traffic videos recorded with a real moving camera (86 human-verified tracks). To address this vulnerability, we propose Geo-CTM (Geometry-Constrained CTM), an association pipeline integrating height-scaled adaptive matching gates, inter-frame similarity estimation via Random Sample Consensus (RANSAC), transform-guided character coasting, and co-occurrence-constrained duplicate track elimination. Systematic motion-model ablation demonstrates that while the complete association pipeline provides the primary foundation for robustness (raising mean accuracy from 85.40% to over 91.5%), estimating a similarity transform (91.82%) delivers the most physically grounded and identifiable representation on planar plates without estimation degeneration. While our method performs comparably to CTM on ideal data when using the same detector and detections, it minimizes performance loss under geometric distortion conditions where CTM is inadequate. For instance, a statistically significant improvement is achieved under a 0 → 60° perspective change; in real traffic videos, with the tracker held fixed so that the fusion layer is the only variable, performance rises from 15.1% to 26.7% under the IoU tracker of the original system and from 16.3% to 29.1% under ByteTrack (+11.6 and +12.8 points; exact McNemar p=0.021 and p=0.013), whereas changing the tracker alone while holding the fusion layer fixed moves accuracy by only 1–2 points and is not statistically significant. Finally, our error taxonomy analysis demonstrates that on the undistorted benchmark all residual errors correspond to zero-evidence cases beyond the reach of decision-level fusion, while under dynamic perspective distortion errors are dominated by association misalignment, highlighting the specific development areas that future performance improvements must target. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

15 pages, 606 KB  
Article
SPA-DETR: An Enhanced RT-DETR with Spatial-Preserving Attention and Adaptive Loss for UAV Spectrogram Signal Detection
by Conghao Fu, Lu Xu and Yijia Zhang
Sensors 2026, 26(15), 4846; https://doi.org/10.3390/s26154846 - 1 Aug 2026
Viewed by 486
Abstract
Rapid detection of unauthorized unmanned aerial vehicles (UAVs) via radio frequency (RF) spectrograms is critical for low-altitude security. However, standard object detectors struggle to locate transient, frequency-hopping UAV signals because their microscopic spatial footprints are easily discarded by conventional lossy downsampling and overwhelmed [...] Read more.
Rapid detection of unauthorized unmanned aerial vehicles (UAVs) via radio frequency (RF) spectrograms is critical for low-altitude security. However, standard object detectors struggle to locate transient, frequency-hopping UAV signals because their microscopic spatial footprints are easily discarded by conventional lossy downsampling and overwhelmed by complex background noise. To overcome this limitation, we propose SPA-DETR, a custom architecture based on the RT-DETR framework. The core of our design is the Spatial-Preserving Attention (SPA) block, which integrates Space-to-Depth Convolution (SPDConv) with a Parallel Patch-Aware Attention (PPA) module. By replacing traditional pooling mechanisms, the SPA block preserves the spatial details of weak signals without information loss, while the PPA module concurrently filters out ambient background interference. Furthermore, to address the severe foreground–background imbalance in RF spectrograms, we introduce an Adaptive Threshold Focal Loss (ATFL). Operating exclusively during training, ATFL prevents background noise gradients from dominating the learning process, forcing the network to focus on hard-to-detect signal patches without adding computational overhead during inference. Experiments on our public RFUAV dataset validate the approach. SPA-DETR achieves an mAP50:95 of 86.8% and an APS of 85.6%, improving upon the baseline RT-DETR-R18 by 4.9% and 5.2%, respectively. Operating at 235.2 FPS with only 23.74 M parameters, SPA-DETR outperforms contemporary detectors such as YOLOv10m, as well as heavier models like YOLOv8m and RT-DETR-R50, highlighting its efficiency and practical value for real-time low-altitude security applications. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

22 pages, 7374 KB  
Article
Cosine Similarity Distillation Vision Mixture-of-Experts for Intelligent Housing-Dimensional Urban Physical Examinations
by Kun Zhao, Helei Ren, Wenbin He, Yuhong Zhao, Jinming Jiang, Wanxiang Yao, Weijun Gao and Qichao Ban
Sensors 2026, 26(11), 3473; https://doi.org/10.3390/s26113473 - 31 May 2026
Viewed by 1249
Abstract
Intelligent housing-dimensional urban physical examination requires evaluating complex visual scenes in aging communities. Existing methods and datasets are insufficient for these heterogeneous tasks and severe class imbalances. To address this, we introduce the Housing-dimensiOnal visUal inSpection [...] Read more.
Intelligent housing-dimensional urban physical examination requires evaluating complex visual scenes in aging communities. Existing methods and datasets are insufficient for these heterogeneous tasks and severe class imbalances. To address this, we introduce the Housing-dimensiOnal visUal inSpection imagE Dataset (HOUSED) with a hierarchical labeling scheme, and propose a hierarchical Vision Mixture of Experts (VMoE) framework. At its core, the proposed CS-DisVMoE module utilizes a CS-Soft routing mechanism to capture spatial feature correlations, optimizing expert assignment and reducing inference overhead. Additionally, a FENNEL-based non-linear graph partitioning mechanism converts pre-trained dense weights into semantically coherent expert initializations, accelerating convergence while preserving localized visual clustering. To address the hierarchical labels, we design a composite loss function: a Supervised Contrastive Loss acts as a parent-category soft constraint to accelerate convergence, while Focal Loss mitigates data imbalance and handles fine-grained subcategory classification via hard sample mining. Across evaluated datasets, the full proposed framework improves accuracy by an average of 4.3% over the ViT-Tiny baseline and 1.81% over the best-performing VMoE baseline. Furthermore, it achieves these improvements with lower computational costs. Further tests on mixed public vision datasets verify its generalizability and competitive performance for complex-scene applications. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

14 pages, 6407 KB  
Article
CoRe: Joint Optimization with Contrastive Learning for Medical Image Registration
by Eytan Kats, Christoph Grossbroehmer, Ziad Al-Haj Hemidi, Fenja Falta, Wiebke Heyer and Mattias P. Heinrich
Sensors 2026, 26(11), 3425; https://doi.org/10.3390/s26113425 - 28 May 2026
Viewed by 610
Abstract
Medical image registration is a fundamental task in medical image analysis, enabling the alignment of images from different modalities or time points. However, intensity inconsistencies and nonlinear tissue deformations pose significant challenges to the robustness of registration methods. Recent approaches leveraging self-supervised representation [...] Read more.
Medical image registration is a fundamental task in medical image analysis, enabling the alignment of images from different modalities or time points. However, intensity inconsistencies and nonlinear tissue deformations pose significant challenges to the robustness of registration methods. Recent approaches leveraging self-supervised representation learning show promise by pre-training feature extractors to generate robust anatomical embeddings, that further used for the registration. In this work, we propose a novel framework that integrates equivariant contrastive learning directly into the registration model. Our approach leverages the power of contrastive learning to learn robust feature representations that are invariant to tissue deformations. By jointly optimizing the contrastive and registration objectives, we ensure that the learned representations are not only informative but also suitable for the registration task. We evaluate our method on abdominal and thoracic image registration tasks, including both intra-patient and inter-patient scenarios. Experimental results demonstrate that the integration of contrastive learning directly into the registration framework significantly improves performance, surpassing strong baseline methods. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

18 pages, 11374 KB  
Article
CSGL-Former: Cross-Stripes Global–Local Fusion Transformer for Remote Sensing Image Dehazing
by Shuyi Feng, Xiran Zhang, Jie Yuan and Youwen Zhu
Sensors 2026, 26(7), 2102; https://doi.org/10.3390/s26072102 - 28 Mar 2026
Viewed by 596
Abstract
Remote sensing (RS) images are often degraded by atmospheric haze, which compromises both visual interpretation and downstream applications. To address this, we introduce CSGL-Former, a novel Cross-Stripes Global–Local Fusion Transformer for RS image dehazing. Our model efficiently captures anisotropic long-range dependencies using cross-stripes [...] Read more.
Remote sensing (RS) images are often degraded by atmospheric haze, which compromises both visual interpretation and downstream applications. To address this, we introduce CSGL-Former, a novel Cross-Stripes Global–Local Fusion Transformer for RS image dehazing. Our model efficiently captures anisotropic long-range dependencies using cross-stripes attention (CSA) and aggregates hierarchical global semantics via a Multi-Layer Global Aggregation (MLGA) module. In the decoder, global context is adaptively blended with fine-grained local features to restore intricate textures. Finally, inspired by the atmospheric scattering model, a soft reconstruction head restores the clear image by predicting spatially varying affine parameters, strictly preserving content fidelity while effectively removing haze. Trained end-to-end, CSGL-Former demonstrates a compelling balance of accuracy and efficiency. Extensive experiments on the RRSHID and SateHaze1K benchmarks show that our model achieves state-of-the-art or highly competitive performance against representative baselines. Ablation studies further validate the effectiveness of each proposed component. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

21 pages, 8478 KB  
Article
ClearSight-RS: A YOLOv5-Based Network with Dynamic Enhancement for Remote Sensing Small Target Detection
by Jie Yuan, Shuyi Feng and Hao Han
Sensors 2026, 26(1), 117; https://doi.org/10.3390/s26010117 - 24 Dec 2025
Cited by 2 | Viewed by 1058
Abstract
Small target detection in remote sensing images faces challenges due to complex backgrounds, weak features, and large scale differences. This paper proposes an improved YOLOv5-based network, termed ClearSight-RS, with the full name “Clear and Accurate Small-target Insight for Remote Sensing”. As the name [...] Read more.
Small target detection in remote sensing images faces challenges due to complex backgrounds, weak features, and large scale differences. This paper proposes an improved YOLOv5-based network, termed ClearSight-RS, with the full name “Clear and Accurate Small-target Insight for Remote Sensing”. As the name implies, the network is dedicated to achieving clear feature perception and accurate target localization for small targets in remote sensing images. The improvements focus on three aspects: integrating an improved Dynamic Snake Convolution (DSConv) module into the backbone network to strengthen the extraction of small target boundaries and geometric features, as well as the expression of weak textures; embedding a Bi-Level Routing Attention (BRA) module in the Neck part to enhance target focusing and suppress background interference; and optimizing the detection head by retaining only shallow high-resolution feature layers for prediction, reducing feature loss and redundant computations. Experimental results show that, based on the VEDAI dataset, ClearSight-RS achieves the highest mAP for all 8 vehicle categories; based on the NWPU VHR-10 dataset, its overall mAP reaches 93.8%, significantly outperforming algorithms such as Faster RCNN and YOLOv5l; based on the DOTA dataset, the capability of the proposed BRA module in suppressing background interference and capturing small target features is demonstrated. The network balances accuracy and efficiency, performing prominently in detecting vehicles and multi-category small targets in complex backgrounds, verifying its effectiveness. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

16 pages, 17338 KB  
Article
MSRS-DETR: End-to-End Object Detection for Multi-Scale Remote Sensing
by Jie Yuan, Shuyi Feng and Hao Han
Sensors 2025, 25(18), 5734; https://doi.org/10.3390/s25185734 - 14 Sep 2025
Cited by 5 | Viewed by 2894
Abstract
Remote sensing imagery (RSI) object detection is critical to many applications, yet mainstream detectors analyse only spatial features and, because of spectral bias, fail to learn high-frequency information adequately, resulting in performance bottlenecks under cluttered backgrounds, distractors, and multi-scale targets, especially small ones. [...] Read more.
Remote sensing imagery (RSI) object detection is critical to many applications, yet mainstream detectors analyse only spatial features and, because of spectral bias, fail to learn high-frequency information adequately, resulting in performance bottlenecks under cluttered backgrounds, distractors, and multi-scale targets, especially small ones. To break these limitations, we propose MSRS-DETR, an end-to-end framework that deeply fuses spatial and frequency cues. The approach introduces three key innovations: (1) C2fFATNET, a frequency-attention-enhanced lightweight residual backbone that provides richer dual-domain features with fewer parameters; (2) an Entanglement Transformer Block (ETB) in the encoder that refines deep semantics via cross-domain frequency–spatial interaction and suppresses background interference; and (3) S2-CCFF, a shallow-feature-extended bidirectional fusion path that markedly improves the retention and utilisation of fine details for small objects. Experiments on HRSC2016 and ShipRSImageNet demonstrate the effectiveness and generalisation of this spatial–frequency paradigm: relative to the baseline, MSRS-DETR reduces parameters by 29.1%, boosts inference speed by 12.4% and 8.4%, and raises mAP50-95 by 1.69% and 2.16%, respectively. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
Show Figures

Figure 1

Back to TopTop