Advanced Methods and Applications with Deep Learning in Object Recognition, 2nd Edition
A Special Issue of Mathematics (ISSN 2227-7390) belonging to the section "E1: Mathematics and Computer Science".
Deadline for manuscript submissions: 27 February 2027 | Viewed by 210
Editors
Interests: information fusion; artificial intelligence; machine vision; autonomous vehicles
Special Issues, Collections and Topics in MDPI journals
Interests: computer vision; multi-object tracking; image/video segmentation; foundation image models; search and rescue; technologies supporting drone operations; information fusion systems; machine/deep learning; artificial intelligence
Special Issue Information
Dear Colleagues,
We are pleased to announce a Second Edition of the Special Issue of Mathematics entitled “Advanced Methods and Applications with Deep Learning in Object Recognition”.
This second edition keeps the focus on object detection and recognition as central tasks in computer vision, essential in most applications of this technology such as video surveillance, warehouse logistics, search and rescue missions or video monitoring using UAVs. The detection conditions may differ across different situations such as low resolution or varying perspectives in aerial images, making it challenging to achieve general solutions requiring minimum fine-tuning to adapt them in new scenarios.
Deep learning has become the reference technique in the field of computer vision for object detection, with a wide range of detection models, generally classified into three main categories: two-stage detectors, one-stage detectors and detectors based on attention mechanisms (transformers).
The first two are the classical detectors based on deep learning, specifically on CNN architectures. Two-stage methods first generate a set of region proposals which may contain objects and then perform classification on these proposals, looking for high accuracy at the cost of increased computational requirements. Some representative solutions of this class are RCNN (region-based convolutional neural network), Fast-RCNN and Faster-RCNN. In contrast, one-stage methods apply a predefined grid to the image and directly compute predictions for each grid cell. Models like Single-Shot Detector (SSD) and You Only Look Once (YOLO) are stand out as one-stage detectors, achieving a good trade-off between accuracy and computational efficiency.
Regarding the transformer-based models, they are based on attention-based architectures which have been imported from natural language processing domain (NLP) and adapted to computer vision tasks. ViT (Vision Transformer) was the first approach designed for image classification, and thereafter other alternatives appeared as the Detection Transformer (DETR), or You Only Look One Sequence (YOLOS), including attentional mechanisms for object detection and classification as alternative to region proposals generation and postprocessing. ViTs usually make use of CNNs as a backbone for relevant feature extraction, and use transformer layers to learn contextualized representations.
More recently, a new generation of object detection models has emerged under the paradigm of Open-Vocabulary Object Detection (OVD). Unlike conventional detectors, which are trained on predefined categories and are therefore commonly referred to as closed-vocabulary models, open-vocabulary approaches leverage the semantic knowledge encoded in large vision-language foundation models to recognize novel object categories that were not explicitly annotated during training. Representative examples include Grounding DINO, GLIP, OWL-ViT, and YOLO-World, which establish semantic correspondences between textual descriptions and image regions, enabling the detection of previously unseen objects through natural language prompts. These models are particularly attractive in scenarios where collecting and annotating large-scale task-specific datasets is impractical, or where the set of target objects evolves over time. Furthermore, they facilitate fine-grained instance discrimination and retrieval through semantic descriptions, opening new possibilities for human-centered and interactive vision systems. Nevertheless, in addition to the classical challenges of object detection, open-vocabulary methods must address issues related to multimodal representation learning, semantic ambiguity, language bias, prompt sensitivity, and the alignment between visual and linguistic concepts, which remain active research topics.
Furthermore, object recognition can be addressed at different levels of granularity, from bounding-box-based object detection to pixel-level semantic and instance segmentation, depending on the requirements of the application and the desired level of scene understanding.
Recent advances in foundation models and vision-language architectures, such as Segment Anything Model (SAM), Grounding DINO, OWL-ViT, and related approaches, have also contributed to bridging object detection and segmentation tasks, enabling more accurate and flexible scene understanding in complex environments.
Most classical detectors have been well studied and highly optimized, with very competitive performance compared with the newly developed models like DETR, their variants and OVD.
Therefore, evaluation is a fundamental task, with different metrics like IoU o mAP to develop fair comparisons among different solutions, considering the balance between accuracy and speed, the resolution of the input images, the configuration of the evaluation parameters, etc. Some relevant challenges are learning models for imbalanced situations, avoiding biases towards minority classes, or the detection of small targets like the ones present in images captured by UAVs, which typically cover big areas with small targets, sometimes also with high density of objects like in crowds or heavy traffic scenarios.
Finally, object detection is closely related with other open challenges in machine vision like Multi-Object Tracking (MOT), which involves both the detection and tracking of objects of interest appearing in the video sequence. The goal in this case is not only to identify and locate the objects contained in each frame, but to also associate them across frames to keep track continuity and follow their dynamics over time. This task is usually solved by combining algorithms addressing object detection and data association, with architectures like ByteTrack or the SORT family (Simple Online and Real-time Tracking) including deepSORT, StrongSORT or OCT-Sort. Typical metrics to evaluate MOT solutions include detection, localization, and association over time, which metrics like MOTA, IDF1 or HOTA, to assess the object association and detection accuracy in complex scenarios such as sets of interacting objects trajectories.
This Special Issue is aimed at contributions focused on these topics, showing the capability of novel mathematical algorithms, architectures and methods to improve the object detection and recognition tasks, with the possibility of multi-object tracking, with an emphasis in new solutions and analysis of their performance in challenging conditions in relevant applications.
Prof. Dr. Jesús García-Herrero
Dr. Juan Pedro Llerena Caña
Guest Editors
Manuscript Submission Information
Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.
Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Mathematics is an international peer-reviewed open access semimonthly journal published by MDPI.
Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2600 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.
Keywords
- object detection and classification
- open-vocabulary object detection
- vision-language models
- foundation models
- multi-object tracking
- deep-learning architectures
- transformers for object detection and segmentation
- loss functions in learning
- class imbalance
- model generalization and domain shift
- evaluation metrics and datasets
- applications of object detection and object tracking
- aerial object identification
- edge ai for object detection
- multi-modal object detection
Benefits of Publishing in a Special Issue
- Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
- Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
- Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
- External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
- Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.
Further information on MDPI's Special Issue policies can be found here.

